Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Agents

MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation se

arXiv cs.AI··Updated just now·38 sightings