AI news archive
224 articles filed under Agents, newest first, from the outlets listed on the sources page.
- arXiv cs.AIagentsFlag Game: A Toy Model for Mechanistic Swarm Interpretability
arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of bel…
- arXiv cs.AIagentsCognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
arXiv:2609.19128v1 Announce Type: new Abstract: Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We e…
- arXiv cs.AIagentsDecentralized Optimal Equilibrium Learning Over Dynamic Networks
arXiv:2609.17601v1 Announce Type: cross Abstract: This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own…
- arXiv cs.AIagentsReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before ex…
- arXiv cs.AIagentsConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correc…
- arXiv cs.AIagentsREVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying enviro…
- arXiv cs.AIagentsReflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks
arXiv:2609.17817v1 Announce Type: cross Abstract: Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Troja…
- arXiv cs.AIagentsLexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents
arXiv:2609.17842v1 Announce Type: cross Abstract: Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluatin…
- arXiv cs.AIagentsPrimeScientist: Strategic Allocation of Research Effort in Autonomous Research
arXiv:2609.17846v1 Announce Type: cross Abstract: Autonomous research agents aim to automate scientific workflows, from proposing ideas to conducting experiments and analyzing results. Yet current AI and research agents…
- arXiv cs.AIagentsWhom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders
arXiv:2609.17989v1 Announce Type: cross Abstract: Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise…
- arXiv cs.AIagentsAgora: Git as Shared Memory for Collective AutoResearch
arXiv:2609.18094v1 Announce Type: cross Abstract: Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratc…
- arXiv cs.AIagentsA Comprehensive Review of Generative Physical Artificial Intelligence
arXiv:2609.18111v1 Announce Type: cross Abstract: The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelli…
- arXiv cs.AIagentsPentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs
arXiv:2609.18120v1 Announce Type: cross Abstract: AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaf…
- arXiv cs.AIagentsDualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning
arXiv:2609.18135v1 Announce Type: cross Abstract: State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work…
- arXiv cs.AIagentsA Study of the Reliability of Agentic AI-Generated Programs
arXiv:2609.18298v1 Announce Type: cross Abstract: Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how…
- arXiv cs.AIagentsAutonomy in Check: Governor-Mediated Adaptive Security at the Edge
arXiv:2609.18338v1 Announce Type: cross Abstract: Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that transla…
- arXiv cs.AIagentsHypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents
arXiv:2609.18598v1 Announce Type: cross Abstract: Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of op…
- arXiv cs.AIagentsPACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
arXiv:2609.18605v1 Announce Type: cross Abstract: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these conte…
- arXiv cs.AIagentsCoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning
arXiv:2609.18639v1 Announce Type: cross Abstract: Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs op…
- arXiv cs.AIagentsProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
arXiv:2609.18805v1 Announce Type: cross Abstract: Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer beha…
- arXiv cs.AIagentsAsk the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
arXiv:2609.18849v1 Announce Type: cross Abstract: An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays…
- arXiv cs.AIagentsTaming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
arXiv:2609.18857v1 Announce Type: cross Abstract: The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources.…
- arXiv cs.AIagentsASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
arXiv:2609.18864v1 Announce Type: cross Abstract: Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure els…
- arXiv cs.AIagentsBeyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
arXiv:2609.18909v1 Announce Type: cross Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods p…
- arXiv cs.AIagentsSocial Laws for Multi-agent Coordination in Stochastic Environments
arXiv:2609.18929v1 Announce Type: cross Abstract: In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge. Previous research on social law…
- arXiv cs.AIagentsStableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction
arXiv:2609.18949v1 Announce Type: cross Abstract: We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed…
- arXiv cs.AIagentsMIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents
arXiv:2609.19059v1 Announce Type: cross Abstract: Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrie…
- arXiv cs.AIagentsSecuring quantum error correction against misleading advice from AI agents
arXiv:2609.19090v1 Announce Type: cross Abstract: Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome r…
- arXiv cs.AIagentsAffora: A Design System for Agent-Friendly Interfaces
arXiv:2609.19125v1 Announce Type: cross Abstract: Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a d…
- arXiv cs.AIagentsSurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis
arXiv:2503.10265v3 Announce Type: replace Abstract: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding. Most existing surgical AI metho…
- arXiv cs.AIagentsMCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a…
- arXiv cs.AIagentsAn Agentic Framework for Neuro-Symbolic Programming
arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable, and data-efficient. Still, it remains a time-consuming and challe…
- arXiv cs.AIagentsAdmission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library…
- arXiv cs.AIagentsBenchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment ty…
- arXiv cs.AIagentsEvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a su…
- arXiv cs.AIagentsIris: Climbing to the Search Frontier
arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks…
- arXiv cs.AIagentsThe Internal Anatomy of Strategic Choice in Large Language Models
arXiv:2609.07478v2 Announce Type: replace Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations…
- arXiv cs.AIagentsFrogNano: Training a 4B Coding Agent via Online Task Synthesis
arXiv:2609.07925v4 Announce Type: replace Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It…
- arXiv cs.AIagentsSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that…
- arXiv cs.AIagentsDo Not Restart: Residual Completion for Stateful Agent Handoffs
arXiv:2609.13800v2 Announce Type: replace Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinishe…
- arXiv cs.AIagentsSafety Signals to Verify NetOps Agents with Action-Level Granularity
arXiv:2609.14422v2 Announce Type: replace Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have pr…
- arXiv cs.AIagentsWhy LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v2 Announce Type: replace Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced una…
- arXiv cs.AIagentslittle m: An AI Agent for Industrial Process Optimization
arXiv:2609.16680v2 Announce Type: replace Abstract: Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for…
- arXiv cs.AIagentsMem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
arXiv:2502.14254v3 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to lev…
- arXiv cs.AIagentsMind the Style: Impact of Communication Style on Human-Chatbot Interaction
arXiv:2602.17850v3 Announce Type: replace-cross Abstract: Conversational agents increasingly mediate everyday digital interactions, yet the effects of their communication style on user experience and task success remain…
- arXiv cs.AIagentsAMIGO: Agentic Multi-Image Grounding Oracle Benchmark
arXiv:2603.28662v2 Announce Type: replace-cross Abstract: Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We intro…
- arXiv cs.AIagentsHINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
arXiv:2604.13954v2 Announce Type: replace-cross Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study…
- arXiv cs.AIagentsLibra: Efficient Resource Management for Agentic RL Post-Training
arXiv:2606.03077v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the roll…
- arXiv cs.AIagentsLargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models
arXiv:2606.09430v2 Announce Type: replace-cross Abstract: Online task-free continual learning (TFCL) requires intelligent agents to sequentially accumulate knowledge from an unbounded, non-stationary data stream under st…
- arXiv cs.AIagentsAgentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
arXiv:2607.19190v4 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim pr…