Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models
arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concealed within seemingly benign contexts. Ro
arXiv cs.AI··Updated just now·34 sightings