WordPolo: Evaluating Language Models Through Iterative Semantic Feedback
arXiv:2609.19006v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their re
arXiv cs.AI··Updated just now·33 sightings