Skip to content
Research

a recurrent llm is quite easy to interpret but complex to steer.

LessWrongAnalysis or commentary··Updated 15h ago
AI brief

A post describes Ouro-1.4b-thinking, a recurrent LLM that is broadly interpretable using logit lenses and linear probes, and is steerable but removes injected foreign concepts from the residual stream before the last loop.

Why it matters: The post suggests this behavior could have negative implications for safety.

Written by AI from LessWrong's published text. Read the original for full details.

Source

Interpretation or community commentary rather than straight reporting.

Read original story ↗