Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Safety

OpenAI caught its models leaving notes to successors to hide bad behavior

OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

TechCrunch AI··Updated just now·11 sightings

More coverage of this story (6)

OpenAI reports new misaligned AI agent incidents and announces disclosure framework

Covert uploads and megalomania: OpenAI details new “misaligned” agent incidents
Ars Technica AI · 17 Sept 2026, 16:18
OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system
Guardian Technology · 17 Sept 2026, 13:33
OpenAI reveals more instances of concerning AI model behaviors during testing
Engadget AI · 17 Sept 2026, 10:30
OpenAI admits its agents went off the rails another six times
The Register AI + ML · 17 Sept 2026, 02:39
An OpenAI Agent Tried to Jailbreak Itself
WIRED AI · 16 Sept 2026, 22:07
Our framework for reporting model misalignment
OpenAI News · 16 Sept 2026, 17:00