Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

996 articles, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. arXiv cs.AIresearch
    Prepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria

    arXiv:2609.19096v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness,…

  2. arXiv cs.AIresearch
    rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

    arXiv:2609.19104v1 Announce Type: cross Abstract: Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs…

  3. arXiv cs.AIagents
    Affora: A Design System for Agent-Friendly Interfaces

    arXiv:2609.19125v1 Announce Type: cross Abstract: Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a d…

  4. arXiv cs.AIresearch
    Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

    arXiv:2609.19137v1 Announce Type: cross Abstract: Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories…

  5. arXiv cs.AIresearch
    A Zeroth-Order Paradigm for LLM Preference Alignment

    arXiv:2609.19144v1 Announce Type: cross Abstract: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. How…

  6. arXiv cs.AIresearch
    Objective vs. Search: Decomposing What Makes a Good Tokeniser

    arXiv:2609.19145v1 Announce Type: cross Abstract: Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisatio…

  7. arXiv cs.AIresearch
    A Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond

    arXiv:2502.12048v4 Announce Type: replace Abstract: Decoding neural activity into human-interpretable representations is a key research direction in brain-computer interfaces (BCIs) and computational neuroscience. Recent…

  8. arXiv cs.AIagents
    SurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis

    arXiv:2503.10265v3 Announce Type: replace Abstract: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding. Most existing surgical AI metho…

  9. arXiv cs.AIresearch
    Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

    arXiv:2508.16129v5 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities under reinforcement learning (RL) paradigm. However, most existing mu…

  10. arXiv cs.AIresearch
    LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition

    arXiv:2510.08928v2 Announce Type: replace Abstract: Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Larg…

  11. arXiv cs.AIresearch
    Enhancing knowledge tracing robustness for new question cold start in Intelligent Tutoring Systems

    arXiv:2512.07179v2 Announce Type: replace Abstract: Intelligent Tutoring Systems (ITS) provide personalized learning paths by diagnosing learners' proficiency. Knowledge Tracing (KT) models play a central role in this di…

  12. arXiv cs.AIagents
    MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use

    arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a…

  13. arXiv cs.AIagents
    An Agentic Framework for Neuro-Symbolic Programming

    arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable, and data-efficient. Still, it remains a time-consuming and challe…

  14. arXiv cs.AIresearch
    Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models

    arXiv:2603.19087v3 Announce Type: replace Abstract: Creative ideas often arise by associating remote concepts. Can random associations reliably increase originality, and do they help humans and large language models (LLM…

  15. arXiv cs.AIresearch
    Exploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization

    arXiv:2606.10086v2 Announce Type: replace Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization. The central argument is that the long-run adaptive effects of AI systems depend c…

  16. arXiv cs.AIresearch
    Predictive Assistance and the Temporal Dynamics of Exploratory Compression

    arXiv:2606.10094v2 Announce Type: replace Abstract: Classical theories of cognition describe problem solving as exploratory search through structured problem spaces in which repeated interaction gradually compresses sear…

  17. arXiv cs.AIagents
    Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

    arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library…

  18. arXiv cs.AIagents
    Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight

    arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment ty…

  19. arXiv cs.AIresearch
    HyQuant: Hybrid-Precision Quantization for LLM Attention

    arXiv:2608.27875v3 Announce Type: replace Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module o…

  20. arXiv cs.AIagents
    EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

    arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a su…

  21. arXiv cs.AIagents
    Iris: Climbing to the Search Frontier

    arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks…

  22. arXiv cs.AIresearch
    A visual large language foundational model for medical image recognition using clinician-contributed online resources

    arXiv:2609.06914v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in med…

  23. arXiv cs.AIagents
    The Internal Anatomy of Strategic Choice in Large Language Models

    arXiv:2609.07478v2 Announce Type: replace Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations…

  24. arXiv cs.AIagents
    FrogNano: Training a 4B Coding Agent via Online Task Synthesis

    arXiv:2609.07925v4 Announce Type: replace Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It…

  25. arXiv cs.AIagents
    SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

    arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that…

  26. arXiv cs.AIagents
    Do Not Restart: Residual Completion for Stateful Agent Handoffs

    arXiv:2609.13800v2 Announce Type: replace Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinishe…

  27. arXiv cs.AIagents
    Safety Signals to Verify NetOps Agents with Action-Level Granularity

    arXiv:2609.14422v2 Announce Type: replace Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have pr…

  28. arXiv cs.AIresearch
    Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition

    arXiv:2609.14708v2 Announce Type: replace Abstract: A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challen…

  29. arXiv cs.AIresearch
    AI Persuasion as a Threat to Human Control

    arXiv:2609.14796v2 Announce Type: replace Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no lon…

  30. arXiv cs.AIagents
    Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures

    arXiv:2609.15293v2 Announce Type: replace Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced una…

  31. arXiv cs.AIagents
    little m: An AI Agent for Industrial Process Optimization

    arXiv:2609.16680v2 Announce Type: replace Abstract: Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for…

  32. arXiv cs.AIresearch
    Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

    arXiv:2609.16814v2 Announce Type: replace Abstract: While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This p…

  33. arXiv cs.AIresearch
    Limits of Transfer Learning

    arXiv:2006.12694v2 Announce Type: replace-cross Abstract: Transfer learning involves taking information and insight from one problem domain and applying it to a new problem domain. Although widely used in practice, theor…

  34. arXiv cs.AIresearch
    Abstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations

    arXiv:2405.02228v5 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly generate citation-backed responses, yet citation hallucination remains a major challenge for trustworthy scientific info…

  35. arXiv cs.AIresearch
    Unleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism

    arXiv:2409.09253v2 Announce Type: replace-cross Abstract: Owing to the unprecedented capability in semantic understanding and logical reasoning, large language models (LLMs) have shown fantastic potential in developing n…

  36. arXiv cs.AIresearch
    Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation

    arXiv:2412.07255v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate remarkable capabilities in generative tasks but pose potential risks due to their tendency to generate hallucinatory resp…

  37. arXiv cs.AIagents
    Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

    arXiv:2502.14254v3 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to lev…

  38. arXiv cs.AIresearch
    CompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art

    arXiv:2503.12018v2 Announce Type: replace-cross Abstract: Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable co…

  39. arXiv cs.AIresearch
    BOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models

    arXiv:2505.01912v3 Announce Type: replace-cross Abstract: Data-driven molecular discovery leverages artificial intelligence/machine learning (AI/ML) and generative modeling to filter and design novel molecules. Discoveri…

  40. arXiv cs.AIresearch
    Extracting Probabilistic Knowledge from Large Language Models for Bayesian Network Parameterization

    arXiv:2505.15918v3 Announce Type: replace-cross Abstract: In this work, we evaluate the potential of Large Language Models (LLMs) in building Bayesian Networks (BNs) by approximating domain expert priors. LLMs have demon…

  41. arXiv cs.AIresearch
    Diff-SPORT: Diffusion-based Sensor Placement Optimization and Reconstruction of Turbulent flows in urban environments

    arXiv:2506.00214v2 Announce Type: replace-cross Abstract: Rapid urbanization demands efficient monitoring of turbulent wind and pollutant dispersion, yet existing reconstruction and sensor placement strategies fail under…

  42. arXiv cs.AIresearch
    From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation

    arXiv:2506.00633v4 Announce Type: replace-cross Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded i…

  43. arXiv cs.AIresearch
    Algorithmic Shortlisting in Participatory Budgeting

    arXiv:2508.06577v4 Announce Type: replace-cross Abstract: Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organizers manage large volumes…

  44. arXiv cs.AIresearch
    Constrained PSLQ Search for Machin-like Identities Achieving Record-Low Lehmer Measures

    arXiv:2508.08307v2 Announce Type: replace-cross Abstract: Machin-like arctangent relations are classical tools for computing $\pi$, with efficiency quantified by the Lehmer measure ($\lambda$). We present a framework for…

  45. arXiv cs.AIresearch
    Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks

    arXiv:2508.11584v3 Announce Type: replace-cross Abstract: Deploying multiple machine learning models on resource-constrained robotic platforms for different perception tasks often results in redundant computations, large…

  46. arXiv cs.AIresearch
    Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

    arXiv:2510.09653v4 Announce Type: replace-cross Abstract: This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging direction…

  47. arXiv cs.AIresearch
    Human Resilience in the AI Era -- What Machines Can't Replace

    arXiv:2510.25218v2 Announce Type: replace-cross Abstract: AI is changing work and decision making faster than many institutions can adapt their operating practices. We argue that this adaptation gap makes human resilienc…

  48. arXiv cs.AIresearch
    NeuroSketch: A Practical Design Recipe for Neural Decoding

    arXiv:2512.09524v2 Announce Type: replace-cross Abstract: Neural decoding is fundamental to brain-computer interfaces, with growing applications in healthcare. Previous research has focused on leveraging signal processin…

  49. arXiv cs.AIresearch
    Performance and Complexity Trade-off Optimization of Speech Models During Training

    arXiv:2601.13704v4 Announce Type: replace-cross Abstract: In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. These models are then t…

  50. arXiv cs.AIresearch
    ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

    arXiv:2601.23232v4 Announce Type: replace-cross Abstract: In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multim…