AI news archive
294 articles filed under Research, newest first, from the outlets listed on the sources page.
- arXiv cs.AIresearchHow to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
arXiv:2605.06850v2 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has emerged as a crucial paradigm for unlocking the advanced reasoning capabilities of Large Language Models (LLMs), encompassing fram…
- arXiv cs.AIresearchEfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control
arXiv:2605.16692v4 Announce Type: replace-cross Abstract: We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorithms. Central…
- arXiv cs.AIresearchDetect Before You Leap: Mirage Detection in Vision-Language Models
arXiv:2606.00435v4 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can produce confident answers without relevant visual evidence, a failure mode known as mirage reasoning (Asadi et al., 2026). To th…
- arXiv cs.AIresearchTime-Aware Diffusion based on Preference Disentanglement for Generative Recommendation
arXiv:2606.01670v3 Announce Type: replace-cross Abstract: Recently, Generative Recommenders (GRs) have emerged as a transformative recommendation paradigm by replacing traditional item IDs with semantic indices (SIDs). O…
- arXiv cs.AIresearchFLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement
arXiv:2606.03852v2 Announce Type: replace-cross Abstract: Large language models often generate code with bugs. Existing methods rely on feedback signals such as test failures and self-critiques to iteratively refine the…
- arXiv cs.AIresearchFrom 'May' to 'Is': Certainty Distortion in Language Model Rewriting
arXiv:2606.07951v2 Announce Type: replace-cross Abstract: Humans increasingly turn to Language Models (LMs) in ways that shape beliefs and drive decisions, including discussing, rewriting, and summarizing information fro…
- arXiv cs.AIresearchFollow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
arXiv:2606.16847v5 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable…
- arXiv cs.AIresearchSegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology
arXiv:2606.17702v3 Announce Type: replace-cross Abstract: Characterising the TME from routine H&E-stained histology images requires simultaneous cell segmentation, biological feature extraction, and interpretable clinica…
- arXiv cs.AIresearchSubjective Risk Decomposition: A New View for Uncertainty Quantification
arXiv:2607.15196v3 Announce Type: replace-cross Abstract: We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequence…
- arXiv cs.AIresearchDebiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
arXiv:2607.15740v3 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly important for fair and trustwo…
- arXiv cs.AIresearchRiemannian Deep Learning: Modules, Networks, and Geometries
arXiv:2607.19305v4 Announce Type: replace-cross Abstract: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Eucl…
- arXiv cs.AIresearchPost-Training in Time Series Foundation Models: A Unifying Framework
arXiv:2607.20002v3 Announce Type: replace-cross Abstract: Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insufficient for reliable do…
- arXiv cs.AIresearchLatency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization
arXiv:2608.00569v3 Announce Type: replace-cross Abstract: Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits from cloud GPUs, wherea…
- arXiv cs.AIresearchRanking Infrared-Visible Fusion the Way Humans Do: A Learned Pairwise Preference Measure
arXiv:2608.01301v4 Announce Type: replace-cross Abstract: Human pairwise comparison provides a direct basis for perceptual infrared-visible image fusion assessment, but dense annotation becomes costly as method pools gro…
- arXiv cs.AIresearchDeep Divide-and-Reduce in Symbolic Regression
arXiv:2608.02628v3 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover underlying mathematical expressions from data while preserving interpretability. Most existing learning-based SR methods…
- arXiv cs.AIresearchExplicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
arXiv:2608.04765v3 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control. However, existing VL…
- arXiv cs.AIresearchUnsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
arXiv:2608.16725v2 Announce Type: replace-cross Abstract: Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognition of dataset quality…
- arXiv cs.AIresearchPosition Matters: Feature Inversion Attacks in ViT Split Inference with Token Reduction and Shuffling
arXiv:2609.01232v2 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) are increasingly used in split-inference systems, where edge devices transmit intermediate token representations to a remote cloud. In…
- arXiv cs.AIresearchNot All Agreement Counts as Corroboration: Provenance-Conserving Multi-View Fusion for Typed Action Admission in Human-Robot Collaboration
arXiv:2609.01662v2 Announce Type: replace-cross Abstract: Better probability scores do not establish that evidence has been counted correctly. Repeated inference over one observation can improve predictions without addin…
- arXiv cs.AIresearchExploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing Data
arXiv:2609.03391v2 Announce Type: replace-cross Abstract: Contrastive language-image learning (CLIP) has become a key paradigm for remote sensing vision-language understanding. However, existing remote sensing contrastiv…
- arXiv cs.AIresearchCalendar-Structured Sparse Principal Component Analysis for Interpretable Multi-Periodic Electricity Consumption Profiles
arXiv:2609.06060v2 Announce Type: replace-cross Abstract: Long-term electricity-consumption profiles exhibit several simultaneous periodic structures, including daily, weekly, and annual cycles. This work introduces Cale…
- arXiv cs.AIresearchSteering Interference Reflects the Model's Defaults, Not the Behavior Directions
arXiv:2609.06951v2 Announce Type: replace-cross Abstract: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and ad…
- arXiv cs.AIresearchProprioception-Anchored Cross-Modal Pretraining for Zero-Shot Sim-to-Real Contact-Rich Assembly
arXiv:2609.07534v3 Announce Type: replace-cross Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Altho…
- arXiv cs.AIresearchAdaptive Anisotropic Attention for Axis-Structured Signals
arXiv:2609.08788v3 Announce Type: replace-cross Abstract: Dense self-attention treats all token pairs as equally plausible before learning, an interaction-isotropic prior that can be mismatched to structured signals. For…
- arXiv cs.AIresearchA Mathematical Theory of Pragmatic Information
arXiv:2609.10986v3 Announce Type: replace-cross Abstract: We propose a mathematical theory of pragmatic information that connects communication, control, and decision-making. Its central notion is the isoteleia mapping,…
- arXiv cs.AIresearchEvaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks
arXiv:2609.12839v3 Announce Type: replace-cross Abstract: The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing an escalating risk…
- arXiv cs.AIresearchCross-Block Conditioning in Deep Boltzmann Machines for Statistical Data Fusion
arXiv:2609.14934v2 Announce Type: replace-cross Abstract: Statistical data fusion combines two panels that share a block of covariates but observe disjoint outcome blocks, and in its traditional form no row observes both…
- arXiv cs.AIresearchPlanning in the Backbone: DiffAdapterVLA for Native Continuous Trajectory Generation with Driving VLMs
arXiv:2609.15322v2 Announce Type: replace-cross Abstract: Pretrained driving vision-language models (VLMs) integrate visual, route, language, and driving context into rich driving priors, yet their representation objecti…
- arXiv cs.AIresearchGeospatial Metadata Improves Discoverability by Connecting Datasets Across Scientific Disciplines
arXiv:2609.16498v2 Announce Type: replace-cross Abstract: Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, an…
- arXiv cs.AIresearchVisual Cue Guided Video Planning for Generalizable Robot Navigation
arXiv:2609.16737v2 Announce Type: replace-cross Abstract: Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition…
- arXiv cs.AIresearchA unified framework for global and local interpretability using adaptive derivative-ordered random explanation
arXiv:2609.17171v2 Announce Type: replace-cross Abstract: The interpretability of complex machine learning models is of paramount importance, especially in real-world high-stakes domains such as healthcare and finance. H…
- arXiv cs.AIresearchVroom-Vroom at SHROOM-Visions: A Multi-Judge Committee for Detecting Hallucinated Spans in Vision-Language Outputs
arXiv:2609.17327v2 Announce Type: replace-cross Abstract: This paper describes our submission to the SHROOM-Visions shared task on detecting and classifying hallucinated character spans in vision-language model outputs a…
- TechCrunch AIresearchAnthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regul…
- AI Stack ExchangeresearchCan in principle GPT language models learn physics?
Does anyone know of research involving the GPT models to learn not only regular texts, but also learn from physics books with the equations written in latex format? My intuition is that the model might learn the rules re…
- Guardian TechnologyresearchMirror publisher to cut 220 editorial jobs as readers turn to AI summaries
Reach, which also owns Express, makes decision because of ‘mammoth shift’ in how audiences seek out content The publisher of the Mirror and Express newspapers is to cut a further 220 editorial jobs as it adapts to a dram…
- Guardian Technologyresearch‘If you’re building Frankenstein, stop’: JD Vance dismisses calls for AI regulation
US vice-president’s comments come as former Anthropic researcher revisits recent claim AI could destroy humanity The US vice-president has dismissed calls for global regulation of AI safety risks, telling companies creat…
- OpenAI NewsresearchHow workers are unlocking new ways of working
New OpenAI Economic Research shows how workers use AI beyond traditional roles and which new activities become recurring parts of their work.
- LessWrongresearchIs METR A Meaningful Check On Anthropic?
Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR ), whose role is to verify adherence to safety practices and commitments, report inciden…
- LessWrongresearchCooperation with AIs seems to be a low-hanging fruit for better eval practices
Summary In his post , Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and an…
- Google ResearchresearchAsk a Scientist: How can researchers use AI to spot a wildfire?
Google Research is exploring how to use AI and satellites to scan the world every 20 minutes and catch wildfires the size of a car.
- Guardian TechnologyresearchWhy a decade of doomsday warnings failed to slow the AI race
From Stephen Hawking to Jacob Coxon’s viral Anthropic resignation, fears that AI could threaten humanity have shaken the industry without stopping its pursuit Before an Anthropic researcher resigned and declared human ex…
- Google AI BlogresearchWatch astronaut Christina Koch and Google’s James Manyika discuss space, technology, and discovery.
Christina Koch sits down with James Manyika, Google’s Senior Vice President of Research, Labs, Technology & Society.
- BBC TechnologyresearchAI staff 'genuinely frightened' for humanity's future, ex-Anthropic researcher tells BBC
It comes as the AI firm's boss has called for the technology's development to be slowed down, citing "serious" risks.
- Ars Technica AIresearchClaude users found ways around safeguards for bioweapons research
Some dangerous biology looks much like legitimate research, complicating AI safeguards.
- AI Stack ExchangeresearchHow does high entropy targets relate to less variance of the gradient between training cases?
I've been trying to understand the Distilling the Knowledge in a Neural Network paper by Hinton et al. But I cannot fully understand this: When the soft targets have high entropy, they provide much more information per t…
- Apple Machine Learning ResearchresearchPutting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from th…
- Engadget AIresearchAnthropic caught scientists using Claude to further biological weapon research
Anthropic produced an extensive collection of case studies covering the ways its current AI models have been misused.
- OpenAI NewsresearchHow a researcher uses Codex and ChatGPT to search for new antimicrobial molecules
César de la Fuente’s lab uses Codex and ChatGPT to search living and extinct genomes for antimicrobial candidates to fight drug-resistant infections.
- OpenAI NewsresearchIntroducing ChatGPT for Financial Services
Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.
- Ars Technica AIresearchAnthropic researcher quits with a warning: Self-improving AI could “kill us all”
“We really do earnestly believe AI could kill all humans!”