AI news archive
996 articles, newest first, from the outlets listed on the sources page.
- arXiv cs.AIresearchPrepared Or Unprepared? Evaluating Healthcare Workforce Readiness for Clinical Adoption of Artificial Intelligence in Nigeria
arXiv:2609.19096v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into healthcare systems worldwide, yet its successful clinical adoption depends critically on workforce readiness,…
- arXiv cs.AIresearchrMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference
arXiv:2609.19104v1 Announce Type: cross Abstract: Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs…
- arXiv cs.AIagentsAffora: A Design System for Agent-Friendly Interfaces
arXiv:2609.19125v1 Announce Type: cross Abstract: Computer-use agents increasingly operate software designed for people, but interfaces often leave actions or task state unclear to machine readers. We present Affora, a d…
- arXiv cs.AIresearchDreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation
arXiv:2609.19137v1 Announce Type: cross Abstract: Recent advances in video generation allow robots to learn manipulation trajectories from generated videos. However, these approaches produce purely kinematic trajectories…
- arXiv cs.AIresearchA Zeroth-Order Paradigm for LLM Preference Alignment
arXiv:2609.19144v1 Announce Type: cross Abstract: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. How…
- arXiv cs.AIresearchObjective vs. Search: Decomposing What Makes a Good Tokeniser
arXiv:2609.19145v1 Announce Type: cross Abstract: Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisatio…
- arXiv cs.AIresearchA Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond
arXiv:2502.12048v4 Announce Type: replace Abstract: Decoding neural activity into human-interpretable representations is a key research direction in brain-computer interfaces (BCIs) and computational neuroscience. Recent…
- arXiv cs.AIagentsSurgRAW: Multi-Agent Workflow with Chain of Thought Reasoning for Robotic Surgical Video Analysis
arXiv:2503.10265v3 Announce Type: replace Abstract: Robotic-assisted surgery (RAS) is central to modern surgery, driving the need for intelligent systems with accurate scene understanding. Most existing surgical AI metho…
- arXiv cs.AIresearchBridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
arXiv:2508.16129v5 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities under reinforcement learning (RL) paradigm. However, most existing mu…
- arXiv cs.AIresearchLM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
arXiv:2510.08928v2 Announce Type: replace Abstract: Existing benchmarks for large multimodal models (LMMs) often fail to capture their performance in real-time, adversarial environments. We introduce LM Fight Arena (Larg…
- arXiv cs.AIresearchEnhancing knowledge tracing robustness for new question cold start in Intelligent Tutoring Systems
arXiv:2512.07179v2 Announce Type: replace Abstract: Intelligent Tutoring Systems (ITS) provide personalized learning paths by diagnosing learners' proficiency. Knowledge Tracing (KT) models play a central role in this di…
- arXiv cs.AIagentsMCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
arXiv:2512.24565v5 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly serving as autonomous agents, and their utilization of external tools via the Model Context Protocol (MCP) is considered a…
- arXiv cs.AIagentsAn Agentic Framework for Neuro-Symbolic Programming
arXiv:2601.00743v2 Announce Type: replace Abstract: Integrating symbolic constraints into deep learning models could make them more robust, interpretable, and data-efficient. Still, it remains a time-consuming and challe…
- arXiv cs.AIresearchAssessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
arXiv:2603.19087v3 Announce Type: replace Abstract: Creative ideas often arise by associating remote concepts. Can random associations reliably increase originality, and do they help humans and large language models (LLM…
- arXiv cs.AIresearchExploratory Responsiveness and Adaptive Rigidity under AI-Assisted Optimization
arXiv:2606.10086v2 Announce Type: replace Abstract: This paper develops a theory of exploratory adaptation under AI-assisted optimization. The central argument is that the long-run adaptive effects of AI systems depend c…
- arXiv cs.AIresearchPredictive Assistance and the Temporal Dynamics of Exploratory Compression
arXiv:2606.10094v2 Announce Type: replace Abstract: Classical theories of cognition describe problem solving as exploratory search through structured problem spaces in which repeated interaction gradually compresses sear…
- arXiv cs.AIagentsAdmission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
arXiv:2608.15565v4 Announce Type: replace Abstract: Agents that learn from experience improve at optimization modeling by storing solved trajectories and reusing them as skills. A wrong trajectory that enters the library…
- arXiv cs.AIagentsBenchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
arXiv:2608.24314v2 Announce Type: replace Abstract: Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment ty…
- arXiv cs.AIresearchHyQuant: Hybrid-Precision Quantization for LLM Attention
arXiv:2608.27875v3 Announce Type: replace Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module o…
- arXiv cs.AIagentsEvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
arXiv:2608.28363v2 Announce Type: replace Abstract: LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a su…
- arXiv cs.AIagentsIris: Climbing to the Search Frontier
arXiv:2609.04304v2 Announce Type: replace Abstract: We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks…
- arXiv cs.AIresearchA visual large language foundational model for medical image recognition using clinician-contributed online resources
arXiv:2609.06914v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in med…
- arXiv cs.AIagentsThe Internal Anatomy of Strategic Choice in Large Language Models
arXiv:2609.07478v2 Announce Type: replace Abstract: Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations…
- arXiv cs.AIagentsFrogNano: Training a 4B Coding Agent via Online Task Synthesis
arXiv:2609.07925v4 Announce Type: replace Abstract: We present FrogNano, a 4B coding agent designed to tackle software engineering (SWE) tasks efficiently and effectively, even under resource-constrained environments. It…
- arXiv cs.AIagentsSWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that…
- arXiv cs.AIagentsDo Not Restart: Residual Completion for Stateful Agent Handoffs
arXiv:2609.13800v2 Announce Type: replace Abstract: Routing and cascades reduce tool-agent cost by transferring control across models, but stateful handoffs must preserve accepted choices, realized effects, and unfinishe…
- arXiv cs.AIagentsSafety Signals to Verify NetOps Agents with Action-Level Granularity
arXiv:2609.14422v2 Announce Type: replace Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. While agents have pr…
- arXiv cs.AIresearchLightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
arXiv:2609.14708v2 Announce Type: replace Abstract: A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challen…
- arXiv cs.AIresearchAI Persuasion as a Threat to Human Control
arXiv:2609.14796v2 Announce Type: replace Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no lon…
- arXiv cs.AIagentsWhy LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures
arXiv:2609.15293v2 Announce Type: replace Abstract: When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced una…
- arXiv cs.AIagentslittle m: An AI Agent for Industrial Process Optimization
arXiv:2609.16680v2 Announce Type: replace Abstract: Manufacturing consumes one third of global energy and still has significant room for improvement in terms of energy efficiency. Optimal process control is essential for…
- arXiv cs.AIresearchCan We Do Interpretable NLI with Graphs Based on Atomic Propositions?
arXiv:2609.16814v2 Announce Type: replace Abstract: While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This p…
- arXiv cs.AIresearchLimits of Transfer Learning
arXiv:2006.12694v2 Announce Type: replace-cross Abstract: Transfer learning involves taking information and insight from one problem domain and applying it to a new problem domain. Although widely used in practice, theor…
- arXiv cs.AIresearchAbstention vs. Hallucination: Benchmarking LLM Source Attribution for Scientific Citations
arXiv:2405.02228v5 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly generate citation-backed responses, yet citation hallucination remains a major challenge for trustworthy scientific info…
- arXiv cs.AIresearchUnleash LLMs Potential for Sequential Recommendation by Coordinating Dual Dynamic Index Mechanism
arXiv:2409.09253v2 Announce Type: replace-cross Abstract: Owing to the unprecedented capability in semantic understanding and logical reasoning, large language models (LLMs) have shown fantastic potential in developing n…
- arXiv cs.AIresearchLabel-Confidence-Aware Uncertainty Estimation in Natural Language Generation
arXiv:2412.07255v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate remarkable capabilities in generative tasks but pose potential risks due to their tendency to generate hallucinatory resp…
- arXiv cs.AIagentsMem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
arXiv:2502.14254v3 Announce Type: replace-cross Abstract: Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to lev…
- arXiv cs.AIresearchCompArt: Operationalizing Aesthetic Alignment in Text-to-Image Generation via Principles of Art
arXiv:2503.12018v2 Announce Type: replace-cross Abstract: Text-to-Image (T2I) diffusion models have made rapid progress on semantic alignment (generating what is described in the prompt), yet users still lack reliable co…
- arXiv cs.AIresearchBOOM: Benchmarking Out-Of-distribution Molecular Property Predictions of Machine Learning Models
arXiv:2505.01912v3 Announce Type: replace-cross Abstract: Data-driven molecular discovery leverages artificial intelligence/machine learning (AI/ML) and generative modeling to filter and design novel molecules. Discoveri…
- arXiv cs.AIresearchExtracting Probabilistic Knowledge from Large Language Models for Bayesian Network Parameterization
arXiv:2505.15918v3 Announce Type: replace-cross Abstract: In this work, we evaluate the potential of Large Language Models (LLMs) in building Bayesian Networks (BNs) by approximating domain expert priors. LLMs have demon…
- arXiv cs.AIresearchDiff-SPORT: Diffusion-based Sensor Placement Optimization and Reconstruction of Turbulent flows in urban environments
arXiv:2506.00214v2 Announce Type: replace-cross Abstract: Rapid urbanization demands efficient monitoring of turbulent wind and pollutant dispersion, yet existing reconstruction and sensor placement strategies fail under…
- arXiv cs.AIresearchFrom Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation
arXiv:2506.00633v4 Announce Type: replace-cross Abstract: Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded i…
- arXiv cs.AIresearchAlgorithmic Shortlisting in Participatory Budgeting
arXiv:2508.06577v4 Announce Type: replace-cross Abstract: Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organizers manage large volumes…
- arXiv cs.AIresearchConstrained PSLQ Search for Machin-like Identities Achieving Record-Low Lehmer Measures
arXiv:2508.08307v2 Announce Type: replace-cross Abstract: Machin-like arctangent relations are classical tools for computing $\pi$, with efficiency quantified by the Lehmer measure ($\lambda$). We present a framework for…
- arXiv cs.AIresearchVisual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks
arXiv:2508.11584v3 Announce Type: replace-cross Abstract: Deploying multiple machine learning models on resource-constrained robotic platforms for different perception tasks often results in redundant computations, large…
- arXiv cs.AIresearchUltralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
arXiv:2510.09653v4 Announce Type: replace-cross Abstract: This paper presents a comprehensive overview of the Ultralytics YOLO family, emphasizing architectural evolution, benchmarking, deployment, and emerging direction…
- arXiv cs.AIresearchHuman Resilience in the AI Era -- What Machines Can't Replace
arXiv:2510.25218v2 Announce Type: replace-cross Abstract: AI is changing work and decision making faster than many institutions can adapt their operating practices. We argue that this adaptation gap makes human resilienc…
- arXiv cs.AIresearchNeuroSketch: A Practical Design Recipe for Neural Decoding
arXiv:2512.09524v2 Announce Type: replace-cross Abstract: Neural decoding is fundamental to brain-computer interfaces, with growing applications in healthcare. Previous research has focused on leveraging signal processin…
- arXiv cs.AIresearchPerformance and Complexity Trade-off Optimization of Speech Models During Training
arXiv:2601.13704v4 Announce Type: replace-cross Abstract: In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. These models are then t…
- arXiv cs.AIresearchShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
arXiv:2601.23232v4 Announce Type: replace-cross Abstract: In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multim…