Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

AI research papers

373 papers, newest first, from the research feeds listed on the sources page.

All categoriesResearch Sources · 362Science · 4Frontier Red Team · 3Alignment · 2Societal Impacts · 1Economics · 1
  1. arXivResearch Sources
    Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

    arXiv:2609.18333v1 Announce Type: cross Abstract: Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word.

  2. arXivResearch Sources
    Autonomy in Check: Governor-Mediated Adaptive Security at the Edge

    arXiv:2609.18338v1 Announce Type: cross Abstract: Adaptive security at the network edge increasingly relies on automated planners, including rule-based controllers, learned policies, and LLM-assisted agents, that transla…

  3. arXivResearch Sources
    Semantic CSI Feedback for Beam Selection: When Task-Aware Embeddings from Sparse Pilots Outperform Full-Bandwidth Reconstruction

    arXiv:2609.18368v1 Announce Type: cross Abstract: Classical CSI feedback in FDD massive MIMO transmits a compressed reconstruction of the channel, optimizing fidelity to the original signal regardless of the downstream t…

  4. arXivResearch Sources
    GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models

    arXiv:2609.18384v1 Announce Type: cross Abstract: We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYR…

  5. arXivResearch Sources
    Reliable Virtual Sensing: A Multi-Domain Benchmark for Robustness Under Sensor Failures

    arXiv:2609.18396v1 Announce Type: cross Abstract: Virtual sensing, the estimation of hard-to-measure quantities from available sensor measurements, is a critical enabler for control and monitoring in cyber-physical syste…

  6. arXivResearch Sources
    A Non-Linear Neuron Based Detection of Isolated Pixels in Binary and Grayscale Images using Contrast Sensitive Receptive Fields

    arXiv:2609.18399v1 Announce Type: cross Abstract: Identifying isolated points is important in image processing applications such as medical imaging, astronomy and quality control management.

  7. arXivResearch Sources
    TERN: A Delta-rule Memory with a Seasonal Reference and Online Adaptation for Epidemic Forecasting

    arXiv:2609.18407v1 Announce Type: cross Abstract: Weekly influenza surveillance counts guide vaccine distribution and public-health alerts, yet they are hard to forecast.

  8. arXivResearch Sources
    Multitask Reinforcement Learning for Assisting Choice Model Specification

    arXiv:2609.18441v1 Announce Type: cross Abstract: Discrete choice model specification is a time-consuming task in which modellers often specify and estimate multiple models while balancing goodness-of-fit, parsimony, and…

  9. arXivResearch Sources
    CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

    arXiv:2609.18462v1 Announce Type: cross Abstract: FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts.

  10. arXivResearch Sources
    ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    arXiv:2609.18487v1 Announce Type: cross Abstract: Action tokenizers play a central role in autoregressive vision-language-action (VLA) models, determining both the targets for policy training and the executable commands…

  11. arXivResearch Sources
    MiST: Mid-Training LLMs for Cybersecurity

    arXiv:2609.18496v1 Announce Type: cross Abstract: Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs.

  12. arXivResearch Sources
    VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval

    arXiv:2609.18521v1 Announce Type: cross Abstract: Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos.

  13. arXivResearch Sources
    Interpretable Patch-Based Deep Learning for Wildfire Spread Prediction from Ensemble Simulations

    arXiv:2609.18555v1 Announce Type: cross Abstract: Wildfire spread is traditionally predicted using physics-based simulators, which are physically interpretable but whose cost increases with each additional ensemble membe…

  14. arXivResearch Sources
    On-the-Fly Homographies Calibration for Multi-Camera Tracking

    arXiv:2609.18582v1 Announce Type: cross Abstract: Precise multi-camera tracking traditionally relies on rigorous 3D site calibration, yet this requirement is often operationally impossible in large-scale deployments.

  15. arXivResearch Sources
    Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

    arXiv:2609.18587v1 Announce Type: cross Abstract: Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a l…

  16. arXivResearch Sources
    Hypothesis-Driven Autonomous Materials Synthesis with Multimodal LLM Agents

    arXiv:2609.18598v1 Announce Type: cross Abstract: Self-driving laboratories can explore synthesis conditions autonomously, but their decision-making layer is typically a black-box optimizer, and the output is a set of op…

  17. arXivResearch Sources
    Online Robust Reinforcement Learning Through Monte-Carlo Planning

    arXiv:2609.18599v1 Announce Type: cross Abstract: Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real…

  18. arXivResearch Sources
    PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

    arXiv:2609.18605v1 Announce Type: cross Abstract: As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance.

  19. arXivResearch Sources
    GenStream: Semantic Streaming Framework for Generative Reconstruction of Human-centric Media

    arXiv:2609.18634v1 Announce Type: cross Abstract: Video streaming dominates global internet traffic, yet conventional pipelines remain inefficient for structured, human-centric content such as sports, performance, or int…

  20. arXivResearch Sources
    CoRe-MARL: Cooperative Redistribution Under Unknown Dynamics Using Recurrent Multi-Agent Reinforcement Learning

    arXiv:2609.18639v1 Announce Type: cross Abstract: Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities.

  21. arXivResearch Sources
    Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification

    arXiv:2609.18673v1 Announce Type: cross Abstract: Speaker de-identification (SDID) aims to preserve privacy by concealing speaker identity while maintaining speech utility.

  22. arXivResearch Sources
    Generalist-Specialist Mixture-of-Experts for Rare Pathology Detection in Multimodal Imaging

    arXiv:2609.18688v1 Announce Type: cross Abstract: AI models for multimodal medical imaging must balance modality-specific specialization with cross-modal shared representations, a trade-off that pure Mixture-of-Experts (…

  23. arXivResearch Sources
    Echo: Learning-based Matching Decompilation using Trusted Back Translation

    arXiv:2609.18706v1 Announce Type: cross Abstract: Neural decompilers can recover readable and recompilable source code from binaries, but their predictions remain difficult to trust.

  24. arXivResearch Sources
    Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

    arXiv:2609.18708v1 Announce Type: cross Abstract: In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy…

  25. arXivResearch Sources
    A Scalable Framework for Automated NER Annotation Correction in Low-Resource Languages

    arXiv:2609.18739v1 Announce Type: cross Abstract: Poor quality or noisy annotations in Named Entity Recognition (NER), as in any other NLP task, make it challenging to achieve state-of-the-art performance.

  26. arXivResearch Sources
    ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    arXiv:2609.18805v1 Announce Type: cross Abstract: Coding agents are typically evaluated with desired behavior specified through issues or instructions.

  27. arXivResearch Sources
    Using OCR Heads to Verbalize Image Semantics

    arXiv:2609.18823v1 Announce Type: cross Abstract: How do VLMs map from pixels to semantics?

  28. arXivResearch Sources
    Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It

    arXiv:2609.18849v1 Announce Type: cross Abstract: An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time.

  29. arXivResearch Sources
    GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

    arXiv:2609.18856v1 Announce Type: cross Abstract: Compact acoustic models face a challenging quality-capacity trade-off.

  30. arXivResearch Sources
    Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN

    arXiv:2609.18857v1 Announce Type: cross Abstract: The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources.

  31. arXivResearch Sources
    Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

    arXiv:2609.18860v1 Announce Type: cross Abstract: When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its ou…

  32. arXivResearch Sources
    ASLEval: Measuring Privacy Exposure Displacement in LLM Agent Sessions

    arXiv:2609.18864v1 Announce Type: cross Abstract: Privacy evaluations of tool-using LLM agents often inspect a designated action, final response, or attacker report.

  33. arXivResearch Sources
    NeuroECG: ECGFounder-Based Deep ECG Representation for EEG-Free Neurological Prognostication After Cardiac Arrest

    arXiv:2609.18891v1 Announce Type: cross Abstract: Neurological prognostication after cardiac arrest commonly relies on electroencephalography (EEG).

  34. arXivResearch Sources
    Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

    arXiv:2609.18909v1 Announce Type: cross Abstract: Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks.

  35. arXivResearch Sources
    Higher-order pruning of experts in mixture-of-experts language models

    arXiv:2609.18916v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck.

  36. arXivResearch Sources
    Social Laws for Multi-agent Coordination in Stochastic Environments

    arXiv:2609.18929v1 Announce Type: cross Abstract: In multi-agent environments, coordinating agents to prevent interference and ensure robust individual performance is a critical challenge.

  37. arXivResearch Sources
    Dose-Aware Cold Diffusion with Physics Consistency for Generalizable Low-Dose CT Reconstruction

    arXiv:2609.18943v1 Announce Type: cross Abstract: Reducing radiation dose in computed tomography significantly degrades image quality and poses challenges for accurate and clinically reliable reconstruction.

  38. arXivResearch Sources
    StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction

    arXiv:2609.18949v1 Announce Type: cross Abstract: We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction.

  39. arXivResearch Sources
    Automated Dental Caries Segmentation in Panoramic Radiographs Using Dual-Stage Deep Learning

    arXiv:2609.18952v1 Announce Type: cross Abstract: Early detection of dental caries remains challenging due to limitations in traditional diagnostic methods, particularly for proximal lesions in posterior teeth.

  40. arXivResearch Sources
    Transcribe, Then Reason: Two-Pass Decomposition for Multimodal Review

    arXiv:2609.18958v1 Announce Type: cross Abstract: The natural way to review a long recording or document with a multimodal model is to hand it the raw source and ask for a review in one call.

  41. arXivResearch Sources
    BadQubits: An LLM-Based Framework for Static Pre-Execution Detection of Structurally Harmful Quantum Circuits

    arXiv:2609.18965v1 Announce Type: cross Abstract: This paper presents BadQubits, an LLM-based framework for static pre-execution detection of structurally harmful OpenQASM 2.0 circuits.

  42. arXivResearch Sources
    WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

    arXiv:2609.19006v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into…

  43. arXivResearch Sources
    Tabular Deep Learning vs Classical Machine Learning for Urban Land Cover Classification

    arXiv:2609.19010v1 Announce Type: cross Abstract: Urban Land Cover (ULC) classification plays a crucial role in urban planning, environmental monitoring, and sustainable development.

  44. arXivResearch Sources
    TalkMatrix: Generating Character Dialogue that is Both Consistent and Diverse

    arXiv:2609.19022v1 Announce Type: cross Abstract: Candidate-based decoding typically selects a completion for each prompt independently, but many applications require a collection of outputs that satisfies global, non-de…

  45. arXivResearch Sources
    MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

    arXiv:2609.19059v1 Announce Type: cross Abstract: Multimodal large language model (MLLM) agents are increasingly used as personal assistants for long-running tasks.

  46. arXivResearch Sources
    RLLBC-Lib: An Educational Code Library for Reinforcement Learning and Learning-Based Control

    arXiv:2609.19074v1 Announce Type: cross Abstract: Reinforcement learning (RL) is an exciting concept as well as a remarkable success story worth sharing.

  47. arXivResearch Sources
    Double descent is the principle of least action

    arXiv:2609.19076v1 Announce Type: cross Abstract: The test error of a model plotted against its number of parameters $d$ falls, peaks when the model can just fit the training data, and falls again, exhibiting the double…

  48. arXivResearch Sources
    Probabilistic Linear Explanations

    arXiv:2609.19077v1 Announce Type: cross Abstract: Formal explainability provides mathematically grounded justifications for individual predictions.

  49. arXivResearch Sources
    Securing quantum error correction against misleading advice from AI agents

    arXiv:2609.19090v1 Announce Type: cross Abstract: Can an attacker turn influence over an artificial intelligence (AI) adviser into a harmful quantum error-correction update?

  50. arXivResearch Sources
    Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

    arXiv:2609.19093v1 Announce Type: cross Abstract: Radiologists follow heterogeneous reporting practices.