Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Agents

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

arXiv:2604.13954v2 Announce Type: replace-cross Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study this complementary but underexplored set

arXiv cs.AI··Updated just now·38 sightings