Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

AI research papers

373 papers, newest first, from the research feeds listed on the sources page.

All categoriesResearch Sources · 362Science · 4Frontier Red Team · 3Alignment · 2Societal Impacts · 1Economics · 1
  1. MIT Technology Review AIResearch Sources
    The Download: mice with part-human brains and climate tech innovators

    This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

  2. MIT Technology Review AIResearch Sources
    Meet the innovators under 35 shaping climate tech

    Each year, the editorial team at MIT Technology Review puts together a list of 35 innovators under 35—a group of researchers, inventors, and other young minds worth following.

  3. arXivResearch Sources
    Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

    arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states.

  4. arXivResearch Sources
    EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

    arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use pol…

  5. arXivResearch Sources
    One Color Preprocessing Improves DSATUR

    arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-…

  6. arXivResearch Sources
    Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees

    arXiv:2609.17635v1 Announce Type: new Abstract: City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as gro…

  7. arXivResearch Sources
    What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

    arXiv:2609.17637v1 Announce Type: new Abstract: Restricting what a module can read may improve what a system learns to compute.

  8. arXivResearch Sources
    CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

    arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-conte…

  9. arXivResearch Sources
    GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

    arXiv:2609.17695v1 Announce Type: new Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence.

  10. arXivResearch Sources
    GVD: Governed Versioning and Deduplication for Document Repositories

    arXiv:2609.17696v1 Announce Type: new Abstract: Document repositories evolve continuously.

  11. arXivResearch Sources
    NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

    arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG).

  12. arXivResearch Sources
    A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products

    arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for environmental monitoring and land management, yet their performa…

  13. arXivResearch Sources
    Imitation Learning for Autonomous Driving in CARLA

    arXiv:2609.17757v1 Announce Type: new Abstract: Behavioral cloning trains a policy offline on expert demonstrations, but deployment is closed loop: each action affects the observations the policy receives next.

  14. arXivResearch Sources
    SAGE: Governed Artifact Generation from Enterprise Guidelines

    arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of…

  15. arXivResearch Sources
    FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

    arXiv:2609.17786v1 Announce Type: new Abstract: Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost.

  16. arXivResearch Sources
    A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it.

  17. arXivResearch Sources
    Learning Heterogeneous Preferences

    arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning.

  18. arXivResearch Sources
    SNOMED CT Concept Recommendation from Masked Clinical Context

    arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant conc…

  19. arXivResearch Sources
    The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

    arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine.

  20. arXivResearch Sources
    Do Frontier Models Seek Safety Evidence Before Acting?

    arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context.

  21. arXivResearch Sources
    ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

    arXiv:2609.17885v1 Announce Type: new Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks.

  22. arXivResearch Sources
    OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead.

  23. arXivResearch Sources
    Collaborative Memory for Multi-Agent VLM Systems

    arXiv:2609.17921v1 Announce Type: new Abstract: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks.

  24. arXivResearch Sources
    Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

    arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-e…

  25. arXivResearch Sources
    Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

    arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, ac…

  26. arXivResearch Sources
    When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI

    arXiv:2609.17977v1 Announce Type: new Abstract: Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service…

  27. arXivResearch Sources
    Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

    arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memor…

  28. arXivResearch Sources
    TuiML: Machine Learning for AI Agents

    arXiv:2609.17984v1 Announce Type: new Abstract: Machine-learning libraries such as Weka and scikit-learn were designed for human programmers.

  29. arXivResearch Sources
    RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

    arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task.

  30. arXivResearch Sources
    Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

    arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design meth…

  31. arXivResearch Sources
    Missing Bridges: Composition-Aware Active Imitation Learning

    arXiv:2609.18004v1 Announce Type: new Abstract: Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs.

  32. arXivResearch Sources
    Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

    arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs).

  33. arXivResearch Sources
    The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

    arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not t…

  34. arXivResearch Sources
    Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools

    arXiv:2609.18072v1 Announce Type: new Abstract: K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship.

  35. arXivResearch Sources
    Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features…

  36. arXivResearch Sources
    When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation

    arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from many documents.

  37. arXivResearch Sources
    AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

    arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustwor…

  38. arXivResearch Sources
    Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

    arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost.

  39. arXivResearch Sources
    Symbolic Temporal Supervision of LLM Agents Using Contracts

    arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting…

  40. arXivResearch Sources
    Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting

    arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they emerge.

  41. arXivResearch Sources
    WFM: Wiki Foundation Model for Complex Agentic Reasoning

    arXiv:2609.18182v1 Announce Type: new Abstract: Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation.

  42. arXivResearch Sources
    Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

    arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during user-recommender interactions.

  43. arXivResearch Sources
    REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

    arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora.

  44. arXivResearch Sources
    BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

    arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evid…

  45. arXivResearch Sources
    Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

    arXiv:2609.18272v1 Announce Type: new Abstract: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor.

  46. arXivResearch Sources
    Building Trust in Artificial Intelligence: A Necessity for Railway Applications

    arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries.

  47. arXivResearch Sources
    Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

    arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) exe…

  48. arXivResearch Sources
    What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

    arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence.

  49. arXivResearch Sources
    Visual Compliance via Executable Safety Rule Entailment

    arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward more contextual and semantic safety concerns.

  50. arXivResearch Sources
    Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

    arXiv:2609.18346v1 Announce Type: new Abstract: Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination.