RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions, performing redundant searches, or makin
arXiv cs.AI··Updated just now·33 sightings