AI research papers
373 papers, newest first, from the research feeds listed on the sources page.
- arXivResearch SourcesPlanning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs
arXiv:2609.15322v2 Announce Type: replace-cross Abstract: Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objecti…
- arXivResearch SourcesGeospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
arXiv:2609.16498v2 Announce Type: replace-cross Abstract: Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, an…
- arXivResearch SourcesVisual Cue Guided Video Planning for Generalizable Robot Navigation
arXiv:2609.16737v2 Announce Type: replace-cross Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans.
- arXivResearch SourcesA unified framework for global and local interpretability using adaptive derivative-ordered random explanation
arXiv:2609.17171v2 Announce Type: replace-cross Abstract: The interpretability of complex machine learning models is of paramount importance, especially in real-world high-stakes domains such as healthcare and finance.
- arXivResearch SourcesAfter the Party: Growth, Governance, and Security Scanning in the OpenClaw Agent Skill Ecosystem
arXiv:2609.17274v2 Announce Type: replace-cross Abstract: AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and proces…
- arXivResearch SourcesVroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
arXiv:2609.17327v2 Announce Type: replace-cross Abstract: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs a…
- Anthropic ResearchScienceHow Claude is uplifting biomolecular modeling
- OpenAI ResearchResearch SourcesOur framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
- MIT Technology Review AIResearch SourcesMeet a mouse whose brain cortex is made up of human cells
Multiple cameras tracked a mouse as it wandered around a small arena.
- MIT Technology Review AIResearch SourcesBuilding the materials foundation for AI
The AI boom is becoming a materials challenge.
- MIT Technology Review AIResearch SourcesThe Download: AI’s trillion-dollar gamble and OpenAI’s biology data bid
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
- Hugging Face PapersResearch SourcesCERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities.
- Hugging Face PapersResearch SourcesIn-Context Robot Learning with VLM Agents
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI.
- Hugging Face PapersResearch SourcesThe Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held.
- Hugging Face PapersResearch SourcesPANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded.
- Hugging Face PapersResearch SourcesRethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates.
- Hugging Face PapersResearch SourcesActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands recovered from predicted tokens.
- Hugging Face PapersResearch SourcesA Zeroth-Order Paradigm for LLM Preference Alignment
Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency.
- Hugging Face PapersResearch SourcesAgora: Git as Shared Memory for Collective AutoResearch
Autonomous research loops such as AutoResearch show that one coding agent can improve a training setup unattended.
- Hugging Face PapersResearch SourcesScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments
Scientific code repositories encode decades of human knowledge in executable models, methods, and tools.
- Hugging Face PapersResearch SourcesGaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX
In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction.
- Hugging Face PapersResearch SourcesProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Coding agents are typically evaluated with desired behavior specified through issues or instructions.
- MIT Technology Review AIResearch SourcesRoundtables: Could AI really kill us all?
Listen to the session or watch below Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity.
- MIT Technology Review AIResearch SourcesThe Download: AI doomers, whistleblowing agents, and de-aged livers
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.
- MIT Technology Review AIResearch SourcesAI models need more data about biology, and OpenAI is paying to create it
Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies.
- MIT Technology Review AIResearch SourcesWhat’s at stake in AI’s trillion-dollar gamble
When Jessica Wachter, a finance professor at the University of Pennsylvania’s Wharton School, wanted to assess AI’s impact on the economy over the next few years, she faced a long list of business and technical uncertain…
- Hugging Face PapersResearch SourcesFingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand
A walking robotic hand must use the same fingers to move its body, support its weight, and interact with the environment.
- Hugging Face PapersResearch SourcesFathom: Per-Query Read Depth for Sparse Decoding over Offloaded KV Caches
When agentic sessions run to a million tokens with many sessions resident at once, the KV cache and the index that ranks it live in host memory, and the scan that ranks all n keys for a top-k step becomes the traffic tha…
- Hugging Face PapersResearch SourcesLimiX-2: A Contextual Mechanism Network Towards General Structured-Data Intelligence
We introduce LimiX-2, a new model in the LimiX family, developed through model and data scaling guided by our previously established scaling laws.
- Hugging Face PapersResearch SourcesConfidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents
Reliable confidence estimation is increasingly central to the trustworthy deployment of language models: a calibrated estimate of the probability that an output is correct decides what to ship, what to escalate, and what…
- Hugging Face PapersResearch SourcesZing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
We introduce Zing-0.5, a 5B autoregressive world model designed for playability: users can explore generated worlds, influence unfolding events, and respond to the resulting feedback through joint keyboard and online tex…
- Hugging Face PapersResearch SourcesEvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents
Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment.
- Hugging Face PapersResearch SourcesEventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset
3D hand mesh reconstruction is a challenging yet essential task for downstream applications, including human-robot interaction and AR/VR.
- Hugging Face PapersResearch SourcesFLAT: Resampling Image and Text into 1D Flexible-Length Aligned Transmodal Tokens for Retrieval and Generation
Traditional multimodal representation learning and generation are two stages: a contrastive or self-supervised visual encoder is trained first, followed by a separate downstream generative model.
- Hugging Face PapersResearch SourcesImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading.
- MIT Technology Review AIResearch SourcesThe AI industry has taken a doomer turn. What now?
This story appeared in The Algorithm, our weekly newsletter on AI.
- Hugging Face PapersResearch SourcesAssessing nnU-Net Generalization across Brain Tumor Populations in BraTS-GoAT 2026
BraTS-GoAT evaluates tumor segmentation across heterogeneous populations.
- Hugging Face PapersResearch SourcesHypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses
Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations.
- Hugging Face PapersResearch SourcesVC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention
Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast.
- Hugging Face PapersResearch SourcesRegister Tokens for Bounded-State Reasoning in Diffusion Language Models
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked tokens with bidirectional attention.
- Hugging Face PapersResearch SourcesFlattening Every Memory Peak in Long-Context Mixture-of-Experts Training
Training a Mixture-of-Experts (MoE) model at long context or large batch size fails as soon as any one component's peak allocation exceeds device memory, so the target is every peak at once, not the average footprint.
- Hugging Face PapersResearch SourcesSpectralShift: Effective Context Window Extension of Gated DeltaNet via Spectral Reparameterization
Recently, linear attention layers have been increasingly adopted to replace softmax attention at scale for long-context modeling.
- Hugging Face PapersResearch SourcesOmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning
Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation.
- Hugging Face PapersResearch SourcesConvergent Emergence of In-Context Learning Across Modalities
Few-shot in-context learning (ICL), the capacity of a model to infer abstract patterns from input-output examples provided in its prompt and apply them to new inputs, has been extensively studied in large language models…
- Hugging Face PapersResearch SourcesRelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
Open-vocabulary detection accepts any class list at inference, and promptable segmentation returns regions without class names: the taxonomy has left the model and become an input.
- Hugging Face PapersResearch SourcesGeneralized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect?
- Anthropic ResearchFrontier Red TeamMeasuring tactical intelligence targeting and conventional weapons capabilities of AI models
- Anthropic ResearchAlignmentAn alignment assessment of recent cybersecurity incidents
- OpenAI ResearchResearch SourcesOn the Navier–Stokes Millennium Prize Problem
We’re sharing an AI-generated solution to the Navier–Stokes Millennium Prize Problem, including a writeup and a formal proof in Lean.
- OpenAI ResearchResearch SourcesResearch acceleration: The view inside OpenAI
Inside OpenAI, coding agents are reshaping AI research.