AI research papers
362 papers in Research Sources, newest first, from the research feeds listed on the sources page.
- arXivResearch SourcesMarket Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents
arXiv:2609.18357v1 Announce Type: new Abstract: Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged.
- arXivResearch SourcesBad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts
arXiv:2609.18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits pro…
- arXivResearch SourcesCultural Competence in Context: A Large Language Model Passes the Turing Test in Finland
arXiv:2609.18394v1 Announce Type: new Abstract: We report the results of a Turing Test conducted in Finland in the Finnish language.
- arXivResearch SourcesHPOQuest: A Rare-Disease Diagnostic Agent Using Active Phenotype Acquisition
arXiv:2609.18431v1 Announce Type: new Abstract: More than 300 million people worldwide are affected by one of over 7,000 known rare diseases, yet diagnosis remains difficult because patients initially present with incomp…
- arXivResearch SourcesWetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories
arXiv:2609.18435v1 Announce Type: new Abstract: Automating biological research requires general-purpose, reproducible robot systems that allow individual wet-lab researchers to delegate robot tasks without performing tel…
- arXivResearch SourcesRisk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving
arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory.
- arXivResearch SourcesThe Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models
arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence.
- arXivResearch SourcesCollective Loss of Control in LLM Agent Systems: An Epidemic Account of Mutation, Contagion, and Recovery
arXiv:2609.18460v1 Announce Type: new Abstract: How does a multi-agent system evolve from a local deviation into collective loss of control?
- arXivResearch SourcesDisentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning
arXiv:2609.18461v1 Announce Type: new Abstract: Personalized agents are required to reason over long-term history interactions to infer both explicit preferences and implicit behavioral evidence.
- arXivResearch SourcesFirst Token Matters: Understanding Safety Collapse in Large Reasoning Models
arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries.
- arXivResearch SourcesHyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs
arXiv:2609.18481v1 Announce Type: new Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, a…
- arXivResearch SourcesBeyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models
arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions.
- arXivResearch SourcesAeroWeaver: An Embodied-Agent Harness for Weaving Aerial Skills into Distributed, Adaptive Swarm Execution
arXiv:2609.18520v1 Announce Type: new Abstract: Collective intelligence is a collaborative autonomy paradigm in which multiple agents pursue shared objectives through local perception, information exchange, and coordinat…
- arXivResearch SourcesTRIPROBE: Probing Task Separability Beyond Classification for XAI
arXiv:2609.18525v1 Announce Type: new Abstract: Modern evaluation of learning pipelines often reduces to downstream accuracy, leaving open the question of why tasks succeed or fail.
- arXivResearch SourcesRecursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making
arXiv:2609.18591v1 Announce Type: new Abstract: In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects…
- arXivResearch SourcesReasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection
arXiv:2609.18597v1 Announce Type: new Abstract: Propagation structures provide crucial evidence for fake news detection, yet existing approaches primarily rely on supervised GNN-based models, which require substantial la…
- arXivResearch SourcesThe Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses
arXiv:2609.18676v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance rema…
- arXivResearch SourcesBeyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning.
- arXivResearch SourcesWhich LLM is Best for Translating Natural Language Goals to PDDL
arXiv:2609.18731v1 Announce Type: new Abstract: Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts access…
- arXivResearch SourcesClueing up LLMs with Tool-Augmented Deductive Reasoning
arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging.
- arXivResearch SourcesVersion- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale
arXiv:2609.18769v1 Announce Type: new Abstract: Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version curren…
- arXivResearch SourcesCERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
arXiv:2609.18779v1 Announce Type: new Abstract: Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent cap…
- arXivResearch SourcesCompositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows
arXiv:2609.18820v1 Announce Type: new Abstract: Agentic workflows now make consequential decisions in regulated settings, and the governance placed around them is almost entirely step-scoped: input-output classifiers, pe…
- arXivResearch SourcesInfinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these law…
- arXivResearch SourcesSuppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on near-edit prompts but not how much of the original fact remains dec…
- arXivResearch SourcesFunction Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation
arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model's computation actually use?
- arXivResearch SourcesLost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning
arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception and reasoning as a single measurable process.
- arXivResearch SourcesCompiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization
arXiv:2609.18996v1 Announce Type: new Abstract: LLM agents have repeatedly struggled to convert knowledge of a game into competent play, even when researchers build the agent around the model - supplying perception, memo…
- arXivResearch SourcesMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
arXiv:2609.19088v1 Announce Type: new Abstract: Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated.
- arXivResearch SourcesFlag Game: A Toy Model for Mechanistic Swarm Interpretability
arXiv:2609.19124v1 Announce Type: new Abstract: Emergent coordinated behaviors of AI agents are starting to present critical safety risks.
- arXivResearch SourcesCognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments
arXiv:2609.19128v1 Announce Type: new Abstract: Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps.
- arXivResearch SourcesThe Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models
arXiv:2408.07702v2 Announce Type: cross Abstract: Schema linking is a crucial step in Text-to-SQL pipelines.
- arXivResearch SourcesREQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration
arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware envir…
- arXivResearch SourcesWARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI
arXiv:2609.17556v1 Announce Type: cross Abstract: Edge-deployed AI operate under dynamically changing power budgets, reliability requirements, and input distributions, requiring continuous adaptation.
- arXivResearch SourcesPay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds
arXiv:2609.17560v1 Announce Type: cross Abstract: Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced.
- arXivResearch SourcesBLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection
arXiv:2609.17562v1 Announce Type: cross Abstract: Hybrid Spiking Neural Network (SNN)-Artificial Neural Network (ANN) architectures combine the energy efficiency of SNNs with the superior detection accuracy of ANNs for e…
- arXivResearch SourcesIndependence-System Realisations in Single-Source Unsplittable Flow
arXiv:2609.17568v1 Announce Type: cross Abstract: Additive-congestion constraints in single-source unsplittable flow can enforce stable-set structure.
- arXivResearch SourcesWhere Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch
arXiv:2609.17571v1 Announce Type: cross Abstract: Where in a Transformer is the change from memorization to generalization functionally expressed?
- arXivResearch SourcesEvolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory
arXiv:2609.17590v1 Announce Type: cross Abstract: Evolutionary Ensemble Search (EES) constructs machine-learning procedures through expert-guided program evolution.
- arXivResearch SourcesStructure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
arXiv:2609.17599v1 Announce Type: cross Abstract: A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foun…
- arXivResearch SourcesDecentralized Optimal Equilibrium Learning Over Dynamic Networks
arXiv:2609.17601v1 Announce Type: cross Abstract: This paper studies decentralized learning of socially optimal equilibria in finite normal-form games over dynamic communication networks.
- arXivResearch SourcesLecture notes on Physics Informed Neural Networks, Neural Operators, and their applications
arXiv:2609.17638v1 Announce Type: cross Abstract: This is the set of lecture notes for the PhD course \href{https://www.unibz.it/en/faculties/engineering/phd-computer-science/study-course-offering/2025/36967}{\textit{Phy…
- arXivResearch SourcesScaling Articulated Rationales for MLLM-based Recommendation
arXiv:2609.17639v1 Announce Type: cross Abstract: Modern recommendation systems largely infer user preferences from implicit behaviors such as clicks, watch time, and negative feedback, but these signals reveal what user…
- arXivResearch SourcesRethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
arXiv:2609.17644v1 Announce Type: cross Abstract: Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-spec…
- arXivResearch SourcesReflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents
arXiv:2609.17653v1 Announce Type: cross Abstract: GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before ex…
- arXivResearch SourcesThe Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention
arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems.
- arXivResearch SourcesAccelerating Diffusion Sampling via Speculative Draft Trees
arXiv:2609.17691v1 Announce Type: cross Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and correcting them under a coupling that preserves the target distri…
- arXivResearch SourcesConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
arXiv:2609.17708v1 Announce Type: cross Abstract: Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correc…
- arXivResearch SourcesOne Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG
arXiv:2609.17709v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query comple…
- arXivResearch SourcesREVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff
arXiv:2609.17745v1 Announce Type: cross Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets.