Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

995 articles, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. Ai2 Blogresearch
    What a crowdsourced game revealed about steering Olmo 3

    A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those te…

  2. TechCrunch AImodels
    Iceland-based Treble raises $18 million for its voice simulation platform

    Treble's voice simulation platform is used by voice AI model developers and AI wearable and robotics companies,

  3. arXiv cs.AIresearch
    Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

    arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states. Provenance, attest…

  4. arXiv cs.AIagents
    EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents

    arXiv:2609.17632v1 Announce Type: new Abstract: Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use pol…

  5. arXiv cs.AIresearch
    One Color Preprocessing Improves DSATUR

    arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-…

  6. arXiv cs.AIresearch
    Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees

    arXiv:2609.17635v1 Announce Type: new Abstract: City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as gro…

  7. arXiv cs.AIresearch
    What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

    arXiv:2609.17637v1 Announce Type: new Abstract: Restricting what a module can read may improve what a system learns to compute. We test this in a preregistered confirmation with sixty four-cell systems sharing a frozen l…

  8. arXiv cs.AIagents
    CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

    arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-conte…

  9. arXiv cs.AIagents
    GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

    arXiv:2609.17695v1 Announce Type: new Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence. GraphEcho tests whether agents mistake these repeated encounters…

  10. arXiv cs.AIresearch
    GVD: Governed Versioning and Deduplication for Document Repositories

    arXiv:2609.17696v1 Announce Type: new Abstract: Document repositories evolve continuously. Guidelines and policies are revised, superseded, and re-uploaded, so the same content recurs in different wording and newer versi…

  11. arXiv cs.AIresearch
    NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

    arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG). Designed to be intuitive to use, NDD provide…

  12. arXiv cs.AIresearch
    A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products

    arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for environmental monitoring and land management, yet their performa…

  13. arXiv cs.AIagents
    Imitation Learning for Autonomous Driving in CARLA

    arXiv:2609.17757v1 Announce Type: new Abstract: Behavioral cloning trains a policy offline on expert demonstrations, but deployment is closed loop: each action affects the observations the policy receives next. We study…

  14. arXiv cs.AIresearch
    SAGE: Governed Artifact Generation from Enterprise Guidelines

    arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of…

  15. arXiv cs.AIagents
    FairCompressAgent: An Agentic Framework for Fairness-Aware Model Compression for FPGA Deployment

    arXiv:2609.17786v1 Announce Type: new Abstract: Fairness-aware model compression requires selecting methods and configurations that balance accuracy, fairness, and deployment cost. These decisions become more difficult w…

  16. arXiv cs.AIresearch
    A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these…

  17. arXiv cs.AIresearch
    Learning Heterogeneous Preferences

    arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Exist…

  18. arXiv cs.AIresearch
    SNOMED CT Concept Recommendation from Masked Clinical Context

    arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant conc…

  19. arXiv cs.AIresearch
    The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

    arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and l…

  20. arXiv cs.AIresearch
    Do Frontier Models Seek Safety Evidence Before Acting?

    arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to ac…

  21. arXiv cs.AIagents
    ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software

    arXiv:2609.17885v1 Announce Type: new Abstract: Computer-use agents that operate through screenshots and simulated actions are advancing rapidly, yet their evaluation remains anchored to general desktop and web tasks. En…

  22. arXiv cs.AIresearch
    OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effecti…

  23. arXiv cs.AIagents
    Collaborative Memory for Multi-Agent VLM Systems

    arXiv:2609.17921v1 Announce Type: new Abstract: Vision-language model (VLM) agents combine specialized perception, tools, and reasoning to address complex visual tasks. In multi-agent settings, different agents inspect d…

  24. arXiv cs.AIresearch
    Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

    arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-e…

  25. arXiv cs.AIresearch
    Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

    arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, ac…

  26. arXiv cs.AIagents
    When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AI

    arXiv:2609.17977v1 Announce Type: new Abstract: Emotion recognition in conversation (ERC) is a production capability behind agent-assist prompts, escalation routing, and post-call analytics in contact-center-as-a-service…

  27. arXiv cs.AIagents
    Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

    arXiv:2609.17983v1 Announce Type: new Abstract: KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memor…

  28. arXiv cs.AIagents
    TuiML: Machine Learning for AI Agents

    arXiv:2609.17984v1 Announce Type: new Abstract: Machine-learning libraries such as Weka and scikit-learn were designed for human programmers. Language-model agents now use these same libraries by recalling APIs from memo…

  29. arXiv cs.AIagents
    RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

    arXiv:2609.17985v1 Announce Type: new Abstract: AI agents are usually evaluated by whether they complete a task. In interactive service settings, a successful agent can still frustrate users by asking repeated questions,…

  30. arXiv cs.AIresearch
    Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

    arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design meth…

  31. arXiv cs.AIresearch
    Missing Bridges: Composition-Aware Active Imitation Learning

    arXiv:2609.18004v1 Announce Type: new Abstract: Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs. Existing methods typically select these requests for their exp…

  32. arXiv cs.AIresearch
    Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

    arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-po…

  33. arXiv cs.AIresearch
    The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

    arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not t…

  34. arXiv cs.AIresearch
    Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools

    arXiv:2609.18072v1 Announce Type: new Abstract: K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship. Programs like FIRST provide competition pathw…

  35. arXiv cs.AIresearch
    Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features…

  36. arXiv cs.AIresearch
    When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation

    arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from many documents. However, building a graph often requires many langu…

  37. arXiv cs.AIagents
    AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

    arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustwor…

  38. arXiv cs.AIagents
    Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

    arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment…

  39. arXiv cs.AIagents
    Symbolic Temporal Supervision of LLM Agents Using Contracts

    arXiv:2609.18128v1 Announce Type: new Abstract: Large language model (LLM) agents augmented by tools can automate complex, multi-step tasks, such as web navigation, code generation, and workflow orchestration, by acting…

  40. arXiv cs.AIresearch
    Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting

    arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they emerge. Existing approaches often model concept semantics and graph st…

  41. arXiv cs.AIagents
    WFM: Wiki Foundation Model for Complex Agentic Reasoning

    arXiv:2609.18182v1 Announce Type: new Abstract: Real-world agents fundamentally require persistent non-parametric knowledge for dynamic reasoning, i.e., long-term memory and retrieval-augmented generation. While graphs h…

  42. arXiv cs.AIresearch
    Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

    arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during user-recommender interactions. This motivates situated conversational recom…

  43. arXiv cs.AIresearch
    REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

    arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora. These challenges often lim…

  44. arXiv cs.AIresearch
    BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

    arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evid…

  45. arXiv cs.AIagents
    Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI

    arXiv:2609.18272v1 Announce Type: new Abstract: Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation…

  46. arXiv cs.AIresearch
    Building Trust in Artificial Intelligence: A Necessity for Railway Applications

    arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries. We propose to…

  47. arXiv cs.AIagents
    Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

    arXiv:2609.18283v1 Announce Type: new Abstract: As telecommunication networks evolve toward autonomous 5G-Advanced and 6G operations, agentic artificial intelligence (AI) workflows, where large language models (LLMs) exe…

  48. arXiv cs.AIresearch
    What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

    arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has rene…

  49. arXiv cs.AIresearch
    Visual Compliance via Executable Safety Rule Entailment

    arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward more contextual and semantic safety concerns. However, as risk pat…

  50. arXiv cs.AIagents
    Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition

    arXiv:2609.18346v1 Announce Type: new Abstract: Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination. We develop a causal graph divergence frame…