Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally drive model

arXiv cs.AI··Updated just now·34 sightings