Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

First Token Matters: Understanding Safety Collapse in Large Reasoning Models

arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to improving safety largely rely on additional

arXiv cs.AI··Updated just now·34 sightings