Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

224 articles filed under Agents, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. arXiv cs.AIagents
    Flag Game: A Toy Model for Mechanistic Swarm Interpretability

    arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of bel…

  2. arXiv cs.AIagents
    Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

    arXiv:2609.19128v1 Announce Type: new Abstract: Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We e…

  3. arXiv cs.AIagents
    Decentralized Optimal Equilibrium Learning Over Dynamic Networks

    arXiv:2609.17601v1 Announce Type: cross Abstract: This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own…

  4. arXiv cs.AIagents
    Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

    arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before ex…

  5. arXiv cs.AIagents
    Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correc…

  6. arXiv cs.AIagents
    REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

    arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying enviro…

  7. arXiv cs.AIagents
    Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

    arXiv:2609.17817v1 Announce Type: cross Abstract: Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Troja…

  8. arXiv cs.AIagents
    Lexara-RF: Reference-Free Metrics for Evaluating Conversational Visual Analytics Agents

    arXiv:2609.17842v1 Announce Type: cross Abstract: Conversational visual analytics (CVA) agents powered by large language models generate visualizations and natural-language explanations from open-ended queries. Evaluatin…

  9. arXiv cs.AIagents
    PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

    arXiv:2609.17846v1 Announce Type: cross Abstract: Autonomous research agents aim to automate scientific workflows, from proposing ideas to conducting experiments and analyzing results. Yet current AI and research agents…

  10. arXiv cs.AIagents
    Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders

    arXiv:2609.17989v1 Announce Type: cross Abstract: Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise…

  11. arXiv cs.AIagents
    Agora: Git as Shared Memory for Collective AutoResearch

    arXiv:2609.18094v1 Announce Type: cross Abstract: Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended. Run several of them and each session starts from scratc…

  12. arXiv cs.AIagents
    A Comprehensive Review of Generative Physical Artificial Intelligence

    arXiv:2609.18111v1 Announce Type: cross Abstract: The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelli…

  13. arXiv cs.AIagents
    PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs

    arXiv:2609.18120v1 Announce Type: cross Abstract: AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaf…

  14. arXiv cs.AIagents
    DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning

    arXiv:2609.18135v1 Announce Type: cross Abstract: State-of-the-art Text-to-SQL systems are typically multi-agent pipelines centered around two fundamental tasks: schema linking and SQL generation. However, existing work…

  15. arXiv cs.AIagents
    A Study of the Reliability of Agentic AI-Generated Programs

    arXiv:2609.18298v1 Announce Type: cross Abstract: Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how…

  16. arXiv cs.AIagents
    Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

    arXiv:2609.18338v1 Announce Type: cross Abstract: Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that transla…

  17. arXiv cs.AIagents
    Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents

    arXiv:2609.18598v1 Announce Type: cross Abstract: Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of op…

  18. arXiv cs.AIagents
    PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    arXiv:2609.18605v1 Announce Type: cross Abstract: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these conte…

  19. arXiv cs.AIagents
    CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning

    arXiv:2609.18639v1 Announce Type: cross Abstract: Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs op…

  20. arXiv cs.AIagents
    ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    arXiv:2609.18805v1 Announce Type: cross Abstract: Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer beha…

  21. arXiv cs.AIagents
    Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    arXiv:2609.18849v1 Announce Type: cross Abstract: An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays…

  22. arXiv cs.AIagents
    Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN

    arXiv:2609.18857v1 Announce Type: cross Abstract: The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources.…

  23. arXiv cs.AIagents
    ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

    arXiv:2609.18864v1 Announce Type: cross Abstract: Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure els…

  24. arXiv cs.AIagents
    Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    arXiv:2609.18909v1 Announce Type: cross Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods p…

  25. arXiv cs.AIagents
    Social Laws for Multi-agent Coordination in Stochastic Environments

    arXiv:2609.18929v1 Announce Type: cross Abstract: In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge. Previous research on social law…

  26. arXiv cs.AIagents
    StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction

    arXiv:2609.18949v1 Announce Type: cross Abstract: We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed…

  27. arXiv cs.AIagents
    MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

    arXiv:2609.19059v1 Announce Type: cross Abstract: Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrie…

  28. arXiv cs.AIagents
    Securing quantum error correction against misleading advice from AI agents

    arXiv:2609.19090v1 Announce Type: cross Abstract: Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome r…

  29. arXiv cs.AIagents
    Affora: A Design System for Agent-Friendly Interfaces

    arXiv:2609.19125v1 Announce Type: cross Abstract: Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a d…

  30. arXiv cs.AIagents
    SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis

    arXiv:2503.10265v3 Announce Type: replace Abstract: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding. Most existing surgical AI metho…

  31. arXiv cs.AIagents
    MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

    arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a…

  32. arXiv cs.AIagents
    An Agentic Framework for Neuro-Symbolic Programming

    arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable, and data-efficient. Still, it remains a time-consuming and challe…

  33. arXiv cs.AIagents
    Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

    arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library…

  34. arXiv cs.AIagents
    Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment ty…

  35. arXiv cs.AIagents
    EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

    arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a su…

  36. arXiv cs.AIagents
    Iris: Climbing to the Search Frontier

    arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks…

  37. arXiv cs.AIagents
    The Internal Anatomy of Strategic Choice in Large Language Models

    arXiv:2609.07478v2 Announce Type: replace Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations…

  38. arXiv cs.AIagents
    FrogNano: Training a 4B Coding Agent via Online Task Synthesis

    arXiv:2609.07925v4 Announce Type: replace Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It…

  39. arXiv cs.AIagents
    SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

    arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that…

  40. arXiv cs.AIagents
    Do Not Restart: Residual Completion for Stateful Agent Handoffs

    arXiv:2609.13800v2 Announce Type: replace Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinishe…

  41. arXiv cs.AIagents
    Safety Signals to Verify NetOps Agents with Action-Level Granularity

    arXiv:2609.14422v2 Announce Type: replace Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have pr…

  42. arXiv cs.AIagents
    Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

    arXiv:2609.15293v2 Announce Type: replace Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced una…

  43. arXiv cs.AIagents
    little m: An AI Agent for Industrial Process Optimization

    arXiv:2609.16680v2 Announce Type: replace Abstract: Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for…

  44. arXiv cs.AIagents
    Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

    arXiv:2502.14254v3 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to lev…

  45. arXiv cs.AIagents
    Mind the Style: Impact of Communication Style on Human-Chatbot Interaction

    arXiv:2602.17850v3 Announce Type: replace-cross Abstract: Conversational agents increasingly mediate everyday digital interactions, yet the effects of their communication style on user experience and task success remain…

  46. arXiv cs.AIagents
    AMIGO: Agentic Multi-Image Grounding Oracle Benchmark

    arXiv:2603.28662v2 Announce Type: replace-cross Abstract: Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We intro…

  47. arXiv cs.AIagents
    HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

    arXiv:2604.13954v2 Announce Type: replace-cross Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study…

  48. arXiv cs.AIagents
    Libra: Efficient Resource Management for Agentic RL Post-Training

    arXiv:2606.03077v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the roll…

  49. arXiv cs.AIagents
    LargeMonitor: Monitoring Online Task-Free Continual Learning via Large Pretrained Models

    arXiv:2606.09430v2 Announce Type: replace-cross Abstract: Online task-free continual learning (TFCL) requires intelligent agents to sequentially accumulate knowledge from an unbounded, non-stationary data stream under st…

  50. arXiv cs.AIagents
    Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

    arXiv:2607.19190v4 Announce Type: replace-cross Abstract: Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim pr…