AI research papers
373 papers, newest first, from the research feeds listed on the sources page.
- MIT Technology Review AIResearch SourcesThe Download: mice with part-human brains and climate tech innovators
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
- MIT Technology Review AIResearch SourcesMeet the innovators under 35 shaping climate tech
Each year, the editorial team at MIT Technology Review puts together a list of 35 innovators under 35—a group of researchers, inventors, and other young minds worth following.
- arXivResearch SourcesMaking AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records
arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states.
- arXivResearch SourcesEvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use pol…
- arXivResearch SourcesOne Color Preprocessing Improves DSATUR
arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-…
- arXivResearch SourcesPhysics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees
arXiv:2609.17635v1 Announce Type: new Abstract: City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as gro…
- arXivResearch SourcesWhat You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization
arXiv:2609.17637v1 Announce Type: new Abstract: Restricting what a module can read may improve what a system learns to compute.
- arXivResearch SourcesCapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-conte…
- arXivResearch SourcesGraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents
arXiv:2609.17695v1 Announce Type: new Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence.
- arXivResearch SourcesGVD: Governed Versioning and Deduplication for Document Repositories
arXiv:2609.17696v1 Announce Type: new Abstract: Document repositories evolve continuously.
- arXivResearch SourcesNeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation
arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG).
- arXivResearch SourcesA Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products
arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for environmental monitoring and land management, yet their performa…
- arXivResearch SourcesImitation Learning for Autonomous Driving in CARLA
arXiv:2609.17757v1 Announce Type: new Abstract: Behavioral cloning trains a policy offline on expert demonstrations, but deployment is closed loop: each action affects the observations the policy receives next.
- arXivResearch SourcesSAGE: Governed Artifact Generation from Enterprise Guidelines
arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of…
- arXivResearch SourcesFairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment
arXiv:2609.17786v1 Announce Type: new Abstract: Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost.
- arXivResearch SourcesA Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning
arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it.
- arXivResearch SourcesLearning Heterogeneous Preferences
arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning.
- arXivResearch SourcesSNOMED CT Concept Recommendation from Masked Clinical Context
arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant conc…
- arXivResearch SourcesThe Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?
arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine.
- arXivResearch SourcesDo Frontier Models Seek Safety Evidence Before Acting?
arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context.
- arXivResearch SourcesERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
arXiv:2609.17885v1 Announce Type: new Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks.
- arXivResearch SourcesOBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead.
- arXivResearch SourcesCollaborative Memory for Multi-Agent VLM Systems
arXiv:2609.17921v1 Announce Type: new Abstract: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks.
- arXivResearch SourcesMeasuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations
arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-e…
- arXivResearch SourcesMemory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI
arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, ac…
- arXivResearch SourcesWhen to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI
arXiv:2609.17977v1 Announce Type: new Abstract: Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service…
- arXivResearch SourcesContiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits
arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memor…
- arXivResearch SourcesTuiML: Machine Learning for AI Agents
arXiv:2609.17984v1 Announce Type: new Abstract: Machine-learning libraries such as Weka and scikit-learn were designed for human programmers.
- arXivResearch SourcesRideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents
arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task.
- arXivResearch SourcesMultimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation
arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design meth…
- arXivResearch SourcesMissing Bridges: Composition-Aware Active Imitation Learning
arXiv:2609.18004v1 Announce Type: new Abstract: Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs.
- arXivResearch SourcesAnchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs).
- arXivResearch SourcesThe Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not t…
- arXivResearch SourcesTeaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools
arXiv:2609.18072v1 Announce Type: new Abstract: K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship.
- arXivResearch SourcesDecodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features…
- arXivResearch SourcesWhen Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation
arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from many documents.
- arXivResearch SourcesAutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustwor…
- arXivResearch SourcesDesigning Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost.
- arXivResearch SourcesSymbolic Temporal Supervision of LLM Agents Using Contracts
arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting…
- arXivResearch SourcesTime-Aligned Evolving Concept Graphs for Scientific Relation Forecasting
arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they emerge.
- arXivResearch SourcesWFM: Wiki Foundation Model for Complex Agentic Reasoning
arXiv:2609.18182v1 Announce Type: new Abstract: Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation.
- arXivResearch SourcesRe2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment
arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during user-recommender interactions.
- arXivResearch SourcesREPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora.
- arXivResearch SourcesBENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evid…
- arXivResearch SourcesWho Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
arXiv:2609.18272v1 Announce Type: new Abstract: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor.
- arXivResearch SourcesBuilding Trust in Artificial Intelligence: A Necessity for Railway Applications
arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries.
- arXivResearch SourcesWhere Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum
arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) exe…
- arXivResearch SourcesWhat Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models
arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence.
- arXivResearch SourcesVisual Compliance via Executable Safety Rule Entailment
arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward more contextual and semantic safety concerns.
- arXivResearch SourcesFaithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition
arXiv:2609.18346v1 Announce Type: new Abstract: Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination.