Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Agents

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

arXiv:2609.18909v1 Announce Type: cross Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model

arXiv cs.AI··Updated just now·33 sightings