Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

AI research papers

362 papers in Research Sources, newest first, from the research feeds listed on the sources page.

All categoriesResearch Sources · 362Science · 4Frontier Red Team · 3Alignment · 2Societal Impacts · 1Economics · 1
  1. arXivResearch Sources
    Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents

    arXiv:2609.18357v1 Announce Type: new Abstract: Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged.

  2. arXivResearch Sources
    Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

    arXiv:2609.18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits pro…

  3. arXivResearch Sources
    Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

    arXiv:2609.18394v1 Announce Type: new Abstract: We report the results of a Turing Test conducted in Finland in the Finnish language.

  4. arXivResearch Sources
    HPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition

    arXiv:2609.18431v1 Announce Type: new Abstract: More than 300 million people worldwide are affected by one of over 7,000 known rare diseases, yet diagnosis remains difficult because patients initially present with incomp…

  5. arXivResearch Sources
    WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories

    arXiv:2609.18435v1 Announce Type: new Abstract: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-lab researchers to delegate robot tasks without performing tel…

  6. arXivResearch Sources
    Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving

    arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory.

  7. arXivResearch Sources
    The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

    arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence.

  8. arXivResearch Sources
    Collective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery

    arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control?

  9. arXivResearch Sources
    Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning

    arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence.

  10. arXivResearch Sources
    First Token Matters: Understanding Safety Collapse in Large Reasoning Models

    arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries.

  11. arXivResearch Sources
    Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs

    arXiv:2609.18481v1 Announce Type: new Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, a…

  12. arXivResearch Sources
    Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

    arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions.

  13. arXivResearch Sources
    AeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution

    arXiv:2609.18520v1 Announce Type: new Abstract: Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinat…

  14. arXivResearch Sources
    TRIPROBE: Probing Task Separability Beyond Classification for XAI

    arXiv:2609.18525v1 Announce Type: new Abstract: Modern evaluation of learning pipelines often reduces to downstream accuracy, leaving open the question of why tasks succeed or fail.

  15. arXivResearch Sources
    Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making

    arXiv:2609.18591v1 Announce Type: new Abstract: In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects…

  16. arXivResearch Sources
    Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection

    arXiv:2609.18597v1 Announce Type: new Abstract: Propagation structures provide crucial evidence for fake news detection, yet existing approaches primarily rely on supervised GNN-based models, which require substantial la…

  17. arXivResearch Sources
    The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses

    arXiv:2609.18676v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance rema…

  18. arXivResearch Sources
    Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

    arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning.

  19. arXivResearch Sources
    Which LLM is Best for Translating Natural Language Goals to PDDL

    arXiv:2609.18731v1 Announce Type: new Abstract: Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts access…

  20. arXivResearch Sources
    Clueing up LLMs with Tool-Augmented Deductive Reasoning

    arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging.

  21. arXivResearch Sources
    Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale

    arXiv:2609.18769v1 Announce Type: new Abstract: Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version curren…

  22. arXivResearch Sources
    CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

    arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent cap…

  23. arXivResearch Sources
    Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

    arXiv:2609.18820v1 Announce Type: new Abstract: Agentic workflows now make consequential decisions in regulated settings, and the governance placed around them is almost entirely step-scoped: input-output classifiers, pe…

  24. arXivResearch Sources
    Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data

    arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these law…

  25. arXivResearch Sources
    Suppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing

    arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on near-edit prompts but not how much of the original fact remains dec…

  26. arXivResearch Sources
    Function Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation

    arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model's computation actually use?

  27. arXivResearch Sources
    Lost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning

    arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception and reasoning as a single measurable process.

  28. arXivResearch Sources
    Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

    arXiv:2609.18996v1 Announce Type: new Abstract: LLM agents have repeatedly struggled to convert knowledge of a game into competent play, even when researchers build the agent around the model - supplying perception, memo…

  29. arXivResearch Sources
    MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

    arXiv:2609.19088v1 Announce Type: new Abstract: Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated.

  30. arXivResearch Sources
    Flag Game: A Toy Model for Mechanistic Swarm Interpretability

    arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks.

  31. arXivResearch Sources
    Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

    arXiv:2609.19128v1 Announce Type: new Abstract: Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps.

  32. arXivResearch Sources
    The Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models

    arXiv:2408.07702v2 Announce Type: cross Abstract: Schema linking is a crucial step in Text-to-SQL pipelines.

  33. arXivResearch Sources
    REQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration

    arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware envir…

  34. arXivResearch Sources
    WARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI

    arXiv:2609.17556v1 Announce Type: cross Abstract: Edge-deployed AI operate under dynamically changing power budgets, reliability requirements, and input distributions, requiring continuous adaptation.

  35. arXivResearch Sources
    Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

    arXiv:2609.17560v1 Announce Type: cross Abstract: Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced.

  36. arXivResearch Sources
    BLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection

    arXiv:2609.17562v1 Announce Type: cross Abstract: Hybrid Spiking Neural Network (SNN)-Artificial Neural Network (ANN) architectures combine the energy efficiency of SNNs with the superior detection accuracy of ANNs for e…

  37. arXivResearch Sources
    Independence-System Realisations in Single-Source Unsplittable Flow

    arXiv:2609.17568v1 Announce Type: cross Abstract: Additive-congestion constraints in single-source unsplittable flow can enforce stable-set structure.

  38. arXivResearch Sources
    Where Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch

    arXiv:2609.17571v1 Announce Type: cross Abstract: Where in a Transformer is the change from memorization to generalization functionally expressed?

  39. arXivResearch Sources
    Evolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory

    arXiv:2609.17590v1 Announce Type: cross Abstract: Evolutionary Ensemble Search (EES) constructs machine-learning procedures through expert-guided program evolution.

  40. arXivResearch Sources
    Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models

    arXiv:2609.17599v1 Announce Type: cross Abstract: A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foun…

  41. arXivResearch Sources
    Decentralized Optimal Equilibrium Learning Over Dynamic Networks

    arXiv:2609.17601v1 Announce Type: cross Abstract: This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks.

  42. arXivResearch Sources
    Lecture notes on Physics Informed Neural Networks, Neural Operators, and their applications

    arXiv:2609.17638v1 Announce Type: cross Abstract: This is the set of lecture notes for the PhD course \href{https://www.unibz.it/en/faculties/engineering/phd-computer-science/study-course-offering/2025/36967}{\textit{Phy…

  43. arXivResearch Sources
    Scaling Articulated Rationales for MLLM-based Recommendation

    arXiv:2609.17639v1 Announce Type: cross Abstract: Modern recommendation systems largely infer user preferences from implicit behaviors such as clicks, watch time, and negative feedback, but these signals reveal what user…

  44. arXivResearch Sources
    Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

    arXiv:2609.17644v1 Announce Type: cross Abstract: Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-spec…

  45. arXivResearch Sources
    Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

    arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before ex…

  46. arXivResearch Sources
    The Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention

    arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems.

  47. arXivResearch Sources
    Accelerating Diffusion Sampling via Speculative Draft Trees

    arXiv:2609.17691v1 Announce Type: cross Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and correcting them under a coupling that preserves the target distri…

  48. arXivResearch Sources
    Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents

    arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correc…

  49. arXivResearch Sources
    One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

    arXiv:2609.17709v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query comple…

  50. arXivResearch Sources
    REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

    arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets.