AI news archive
995 articles, newest first, from the outlets listed on the sources page.
- arXiv cs.AIresearchLook Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR
arXiv:2609.18333v1 Announce Type: cross Abstract: Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has bec…
- arXiv cs.AIagentsAutonomy in Check: Governor-Mediated Adaptive Security at the Edge
arXiv:2609.18338v1 Announce Type: cross Abstract: Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that transla…
- arXiv cs.AIresearchSemantic CSI Feedback for Beam Selection: When Task-Aware Embeddings from Sparse Pilots Outperform Full-Bandwidth Reconstruction
arXiv:2609.18368v1 Announce Type: cross Abstract: Classical CSI feedback in FDD massive MIMO transmits a compressed reconstruction of the channel, optimizing fidelity to the original signal regardless of the downstream t…
- arXiv cs.AIresearchGYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models
arXiv:2609.18384v1 Announce Type: cross Abstract: We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYR…
- arXiv cs.AIresearchReliable Virtual Sensing: A Multi-Domain Benchmark for Robustness Under Sensor Failures
arXiv:2609.18396v1 Announce Type: cross Abstract: Virtual sensing, the estimation of hard-to-measure quantities from available sensor measurements, is a critical enabler for control and monitoring in cyber-physical syste…
- arXiv cs.AIresearchA Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields
arXiv:2609.18399v1 Announce Type: cross Abstract: Identifying isolated points is important in image processing applications such as medical imaging, astronomy and quality control management. Other domains, such as cybers…
- arXiv cs.AIresearchTERN: A Delta-rule Memory with a Seasonal Reference and Online Adaptation for Epidemic Forecasting
arXiv:2609.18407v1 Announce Type: cross Abstract: Weekly influenza surveillance counts guide vaccine distribution and public-health alerts, yet they are hard to forecast. Each region offers only a few seasons, waves shif…
- arXiv cs.AIresearchMultitask Reinforcement Learning for Assisting Choice Model Specification
arXiv:2609.18441v1 Announce Type: cross Abstract: Discrete choice model specification is a time-consuming task in which modellers often specify and estimate multiple models while balancing goodness-of-fit, parsimony, and…
- arXiv cs.AIresearchCSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models
arXiv:2609.18462v1 Announce Type: cross Abstract: FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representa…
- arXiv cs.AIresearchActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
arXiv:2609.18487v1 Announce Type: cross Abstract: Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands…
- arXiv cs.AIresearchMiST: Mid-Training LLMs for Cybersecurity
arXiv:2609.18496v1 Announce Type: cross Abstract: Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs. We present MiST (Mid-trained Security…
- arXiv cs.AIresearchVoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval
arXiv:2609.18521v1 Announce Type: cross Abstract: Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos. Existing benchmarks and models hav…
- arXiv cs.AIresearchInterpretable Patch-Based Deep Learning for Wildfire Spread Prediction from Ensemble Simulations
arXiv:2609.18555v1 Announce Type: cross Abstract: Wildfire spread is traditionally predicted using physics-based simulators, which are physically interpretable but whose cost increases with each additional ensemble membe…
- arXiv cs.AIresearchOn-the-Fly Homographies Calibration for Multi-Camera Tracking
arXiv:2609.18582v1 Announce Type: cross Abstract: Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments. Pri…
- arXiv cs.AIresearchLabel-free steering: Compressing test-time reinforcement learning into bias-only subspaces
arXiv:2609.18587v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a l…
- arXiv cs.AIagentsHypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents
arXiv:2609.18598v1 Announce Type: cross Abstract: Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of op…
- arXiv cs.AIresearchOnline Robust Reinforcement Learning Through Monte-Carlo Planning
arXiv:2609.18599v1 Announce Type: cross Abstract: Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real…
- arXiv cs.AIagentsPACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
arXiv:2609.18605v1 Announce Type: cross Abstract: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these conte…
- arXiv cs.AIresearchGenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media
arXiv:2609.18634v1 Announce Type: cross Abstract: Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or int…
- arXiv cs.AIagentsCoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning
arXiv:2609.18639v1 Announce Type: cross Abstract: Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs op…
- arXiv cs.AIresearchBeyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification
arXiv:2609.18673v1 Announce Type: cross Abstract: Speaker de-identification (SDID) aims to preserve privacy by concealing speaker identity while maintaining speech utility. However, current evaluations often reduce priva…
- arXiv cs.AIresearchGeneralist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging
arXiv:2609.18688v1 Announce Type: cross Abstract: AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (…
- arXiv cs.AIresearchEcho: Learning-based Matching Decompilation using Trusted Back Translation
arXiv:2609.18706v1 Announce Type: cross Abstract: Neural decompilers can recover readable and recompilable source code from binaries, but their predictions remain difficult to trust. Matching decompilation addresses this…
- arXiv cs.AIresearchRethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening
arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy…
- arXiv cs.AIresearchA Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages
arXiv:2609.18739v1 Announce Type: cross Abstract: Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance. In this paper,…
- arXiv cs.AIagentsProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
arXiv:2609.18805v1 Announce Type: cross Abstract: Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer beha…
- arXiv cs.AIresearchUsing OCR Heads to Verbalize Image Semantics
arXiv:2609.18823v1 Announce Type: cross Abstract: How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Ac…
- arXiv cs.AIagentsAsk the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It
arXiv:2609.18849v1 Announce Type: cross Abstract: An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays…
- arXiv cs.AIresearchGrainSpeech: Less Context, More Detail for Compact Speech Synthesis
arXiv:2609.18856v1 Announce Type: cross Abstract: Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A recep…
- arXiv cs.AIagentsTaming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
arXiv:2609.18857v1 Announce Type: cross Abstract: The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources.…
- arXiv cs.AIresearchDecodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
arXiv:2609.18860v1 Announce Type: cross Abstract: When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its ou…
- arXiv cs.AIagentsASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions
arXiv:2609.18864v1 Announce Type: cross Abstract: Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report. These local proxies can miss unauthorized exposure els…
- arXiv cs.AIresearchNeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest
arXiv:2609.18891v1 Announce Type: cross Abstract: Neurological prognostication after cardiac arrest commonly relies on electroencephalography (EEG). However, EEG demands high clinical resources. Bedside electrocardiograp…
- arXiv cs.AIagentsBeyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking
arXiv:2609.18909v1 Announce Type: cross Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods p…
- arXiv cs.AIresearchHigher-order pruning of experts in mixture-of-experts language models
arXiv:2609.18916v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for…
- arXiv cs.AIagentsSocial Laws for Multi-agent Coordination in Stochastic Environments
arXiv:2609.18929v1 Announce Type: cross Abstract: In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge. Previous research on social law…
- arXiv cs.AIresearchDose-Aware Cold Diffusion with Physics Consistency for Generalizable Low-Dose CT Reconstruction
arXiv:2609.18943v1 Announce Type: cross Abstract: Reducing radiation dose in computed tomography significantly degrades image quality and poses challenges for accurate and clinically reliable reconstruction. While recent…
- arXiv cs.AIagentsStableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction
arXiv:2609.18949v1 Announce Type: cross Abstract: We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed…
- arXiv cs.AIresearchAutomated Dental Caries Segmentation in Panoramic Radiographs Using Dual-Stage Deep Learning
arXiv:2609.18952v1 Announce Type: cross Abstract: Early detection of dental caries remains challenging due to limitations in traditional diagnostic methods, particularly for proximal lesions in posterior teeth. Deep lear…
- arXiv cs.AIresearchTranscribe, Then Reason: Two-Pass Decomposition for Multimodal Review
arXiv:2609.18958v1 Announce Type: cross Abstract: The natural way to review a long recording or document with a multimodal model is to hand it the raw source and ask for a review in one call. We show that this quietly fa…
- arXiv cs.AIresearchBadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits
arXiv:2609.18965v1 Announce Type: cross Abstract: This paper presents BadQubits, an LLM-based framework for static pre-execution detection of structurally harmful OpenQASM 2.0 circuits. The framework targets physical-exe…
- arXiv cs.AIresearchWordPolo: Evaluating Language Models Through Iterative Semantic Feedback
arXiv:2609.19006v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into…
- arXiv cs.AIresearchTabular Deep Learning vs Classical Machine Learning for Urban Land Cover Classification
arXiv:2609.19010v1 Announce Type: cross Abstract: Urban Land Cover (ULC) classification plays a crucial role in urban planning, environmental monitoring, and sustainable development. We study this task using the ULC data…
- arXiv cs.AIresearchTalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse
arXiv:2609.19022v1 Announce Type: cross Abstract: Candidate-based decoding typically selects a completion for each prompt independently, but many applications require a collection of outputs that satisfies global, non-de…
- arXiv cs.AIagentsMIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents
arXiv:2609.19059v1 Announce Type: cross Abstract: Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks. Their utility depends on continuity: agents must retrie…
- arXiv cs.AIresearchRLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control
arXiv:2609.19074v1 Announce Type: cross Abstract: Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing. However, RL builds on rather complex interactions between differen…
- arXiv cs.AIresearchDouble descent is the principle of least action
arXiv:2609.19076v1 Announce Type: cross Abstract: The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double…
- arXiv cs.AIresearchProbabilistic Linear Explanations
arXiv:2609.19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions. However, abductive explanations often exceed human cognitive limits by i…
- arXiv cs.AIagentsSecuring quantum error correction against misleading advice from AI agents
arXiv:2609.19090v1 Announce Type: cross Abstract: Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update? We identify an ambiguity in passive syndrome r…
- arXiv cs.AIresearchReporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
arXiv:2609.19093v1 Announce Type: cross Abstract: Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose sup…