AI brief
A post describes Ouro-1.4b-thinking, a recurrent LLM that is broadly interpretable using logit lenses and linear probes, and is steerable but removes injected foreign concepts from the residual stream before the last loop.
Why it matters: The post suggests this behavior could have negative implications for safety.
Written by AI from LessWrong's published text. Read the original for full details.