Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We
arXiv cs.AI··Updated just now·38 sightings