AI news archive
995 articles, newest first, from the outlets listed on the sources page.
- arXiv cs.AIagentsMarket Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
arXiv:2609.18357v1 Announce Type: new Abstract: Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection…
- arXiv cs.AIagentsBad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
arXiv:2609.18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits pro…
- arXiv cs.AIresearchCultural Competence in Context: A Large Language Model Passes the Turing Test in Finland
arXiv:2609.18394v1 Announce Type: new Abstract: We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data,…
- arXiv cs.AIagentsHPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition
arXiv:2609.18431v1 Announce Type: new Abstract: More than 300 million people worldwide are affected by one of over 7,000 known rare diseases, yet diagnosis remains difficult because patients initially present with incomp…
- arXiv cs.AIagentsWetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories
arXiv:2609.18435v1 Announce Type: new Abstract: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-lab researchers to delegate robot tasks without performing tel…
- arXiv cs.AIresearchRisk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving
arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory. We introduce RiskWorld, a…
- arXiv cs.AIresearchThe Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models
arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence. We find th…
- arXiv cs.AIagentsCollective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control? We propose an epidemic explanation organized around accidental mutation, contag…
- arXiv cs.AIagentsDisentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence. While early flat retriev…
- arXiv cs.AIresearchFirst Token Matters: Understanding Safety Collapse in Large Reasoning Models
arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to impr…
- arXiv cs.AIresearchHyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs
arXiv:2609.18481v1 Announce Type: new Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, a…
- arXiv cs.AIresearchBeyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models
arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concea…
- arXiv cs.AIagentsAeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution
arXiv:2609.18520v1 Announce Type: new Abstract: Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinat…
- arXiv cs.AIresearchTRIPROBE: Probing Task Separability Beyond Classification for XAI
arXiv:2609.18525v1 Announce Type: new Abstract: Modern evaluation of learning pipelines often reduces to downstream accuracy, leaving open the question of why tasks succeed or fail. TriProbe addresses this gap with a mul…
- arXiv cs.AIagentsRecursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
arXiv:2609.18591v1 Announce Type: new Abstract: In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects…
- arXiv cs.AIresearchReasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection
arXiv:2609.18597v1 Announce Type: new Abstract: Propagation structures provide crucial evidence for fake news detection, yet existing approaches primarily rely on supervised GNN-based models, which require substantial la…
- arXiv cs.AIresearchThe Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses
arXiv:2609.18676v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance rema…
- arXiv cs.AIresearchBeyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. E…
- arXiv cs.AIresearchWhich LLM is Best for Translating Natural Language Goals to PDDL
arXiv:2609.18731v1 Announce Type: new Abstract: Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts access…
- arXiv cs.AIresearchClueing up LLMs with Tool-Augmented Deductive Reasoning
arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that requ…
- arXiv cs.AIresearchVersion- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale
arXiv:2609.18769v1 Announce Type: new Abstract: Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version curren…
- arXiv cs.AIagentsCERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent cap…
- arXiv cs.AIagentsCompositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
arXiv:2609.18820v1 Announce Type: new Abstract: Agentic workflows now make consequential decisions in regulated settings, and the governance placed around them is almost entirely step-scoped: input-output classifiers, pe…
- arXiv cs.AIresearchInfinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these law…
- arXiv cs.AIresearchSuppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on near-edit prompts but not how much of the original fact remains dec…
- arXiv cs.AIresearchFunction Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation
arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-d…
- arXiv cs.AIresearchLost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning
arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception and reasoning as a single measurable process. We introduce a five-ta…
- arXiv cs.AIagentsCompiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization
arXiv:2609.18996v1 Announce Type: new Abstract: LLM agents have repeatedly struggled to convert knowledge of a game into competent play, even when researchers build the agent around the model - supplying perception, memo…
- arXiv cs.AIresearchMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
arXiv:2609.19088v1 Announce Type: new Abstract: Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated.…
- arXiv cs.AIagentsFlag Game: A Toy Model for Mechanistic Swarm Interpretability
arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of bel…
- arXiv cs.AIagentsCognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
arXiv:2609.19128v1 Announce Type: new Abstract: Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We e…
- arXiv cs.AIresearchThe Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models
arXiv:2408.07702v2 Announce Type: cross Abstract: Schema linking is a crucial step in Text-to-SQL pipelines. Its goal is to retrieve the relevant tables and columns of a target database for a user's query while disregard…
- arXiv cs.AIresearchREQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration
arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware envir…
- arXiv cs.AIresearchWARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI
arXiv:2609.17556v1 Announce Type: cross Abstract: Edge-deployed AI operate under dynamically changing power budgets, reliability requirements, and input distributions, requiring continuous adaptation. Such conditions ari…
- arXiv cs.AIresearchPay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds
arXiv:2609.17560v1 Announce Type: cross Abstract: Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced. We formaliz…
- arXiv cs.AIresearchBLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection
arXiv:2609.17562v1 Announce Type: cross Abstract: Hybrid Spiking Neural Network (SNN)-Artificial Neural Network (ANN) architectures combine the energy efficiency of SNNs with the superior detection accuracy of ANNs for e…
- arXiv cs.AIresearchIndependence-System Realisations in Single-Source Unsplittable Flow
arXiv:2609.17568v1 Announce Type: cross Abstract: Additive-congestion constraints in single-source unsplittable flow can enforce stable-set structure. This note isolates and generalises that mechanism. We introduce a pat…
- arXiv cs.AIresearchWhere Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch
arXiv:2609.17571v1 Announce Type: cross Abstract: Where in a Transformer is the change from memorization to generalization functionally expressed? We introduce Transition Games--behavior-aligned exact activation games wi…
- arXiv cs.AIresearchEvolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory
arXiv:2609.17590v1 Announce Type: cross Abstract: Evolutionary Ensemble Search (EES) constructs machine-learning procedures through expert-guided program evolution. A role-specialized council turns task evidence and expe…
- arXiv cs.AIresearchStructure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
arXiv:2609.17599v1 Announce Type: cross Abstract: A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foun…
- arXiv cs.AIagentsDecentralized Optimal Equilibrium Learning Over Dynamic Networks
arXiv:2609.17601v1 Announce Type: cross Abstract: This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks. Each agent observes only its own…
- arXiv cs.AIresearchLecture notes on Physics Informed Neural Networks, Neural Operators, and their applications
arXiv:2609.17638v1 Announce Type: cross Abstract: This is the set of lecture notes for the PhD course \href{https://www.unibz.it/en/faculties/engineering/phd-computer-science/study-course-offering/2025/36967}{\textit{Phy…
- arXiv cs.AIresearchScaling Articulated Rationales for MLLM-based Recommendation
arXiv:2609.17639v1 Announce Type: cross Abstract: Modern recommendation systems largely infer user preferences from implicit behaviors such as clicks, watch time, and negative feedback, but these signals reveal what user…
- arXiv cs.AIresearchRethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
arXiv:2609.17644v1 Announce Type: cross Abstract: Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-spec…
- arXiv cs.AIagentsReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before ex…
- arXiv cs.AIresearchThe Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention
arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymano…
- arXiv cs.AIresearchAccelerating Diffusion Sampling via Speculative Draft Trees
arXiv:2609.17691v1 Announce Type: cross Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and correcting them under a coupling that preserves the target distri…
- arXiv cs.AIagentsConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correc…
- arXiv cs.AIresearchOne Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG
arXiv:2609.17709v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query comple…
- arXiv cs.AIagentsREVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying enviro…