MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a future trend. Current MCP evaluation se
arXiv cs.AI··Updated just now·38 sightings