Skip to content
Research

a recurrent llm is quite easy to interpret but very hard to steer.

LessWrongAnalysis or commentary··Updated 17h ago
AI brief

A write-up on LessWrong reports that Ouro-1.4b-thinking, described as a recurrent LLM, is broadly interpretable using logit lenses and linear probes and is steerable, but removes injected foreign concepts from the residual stream if they are added before the last loop.

Why it matters: The author suggests this behavior could have negative implications for safety.

Written by AI from LessWrong's published text. Read the original for full details.

Source

Interpretation or community commentary rather than straight reporting.

Read original story ↗