Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

293 articles filed under Research, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. LessWrongresearch
    A Defense of Gradual Disempowerment

    (Or: Why Bentham's Bulldog and John Halstead are wrong in their critique of Kulveit et al. ) Gradual Disempowerment is a 2025 paper (with a nice, dedicated website ) proposing a form of existential risk from AI that goes…

  2. LessWrongresearch
    Good and bad ways to evaluate a definition

    Sometimes conversations involve people using the same word differently. In the best case scenario, participants notice and choose a favorable provisional definition. Outside of conversations, people advocate for specific…

  3. AWS Machine Learning Blogresearch
    Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent

    Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify s…

  4. TechCrunch AIresearch
    Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire

    Base Labs, the research group Baseten spun up earlier this year, will develop and publish methods for training and monitoring open models.

  5. TechRadar AIresearch
    Beyond EAA compliance: Accessibility becomes the benchmark for digital quality

    Mainstream AI news from a large tech publication.

  6. LessWrongresearch
    Pacing the Frontier: A Framework & Research Agenda

    Below is the executive summary from our new paper at pacing.tech . The full paper is available on the site and as a PDF. The full author list is Raymond Douglas, Charles Dillon, Nikola Moore, Gavin Leech, Shahar Avin, Ma…

  7. LessWrongresearch
    Astra uses some of its no-CoT capability in practice

    Astra scores significantly higher than previous models on no-CoT benchmarks, as for example shown in Neel Nanda's post last week. This raises the question of whether, and to what degree, Astra uses this no-CoT capability…

  8. The Verge AIresearch
    The sexy AI-powered dating app scams are here

    Security researcher Matthew "Zigula" Gore-Kormanik was analyzing a fraudulent dating app called Dora when he got a pop-up message saying he was receiving a call from Jennifer. According to her bio, she's a 41-year-old Sa…

  9. The Verge AIresearch
    AI is feared globally as the destroyer of jobs

    Pew Research has published a new global survey that sheds light on how people view AI, including its impact on jobs, life in general, and income inequality. The survey questioned 42,151 people across 37 countries from Fe…

  10. The Verge AIresearch
    Inside the suddenly explosive world of AI safety

    On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building. They had come together for a "war room" to dissect the high-profile cybersecurit…

  11. Ai2 Blogresearch
    What a crowdsourced game revealed about steering Olmo 3

    A crowdsourced game built on Olmo 3 showed how people can exploit unexpected model behaviors to stress-test prosocial AI evaluations—and how open access to a model’s internals can help researchers understand why those te…

  12. arXiv cs.AIresearch
    Making AI-Assisted Claims Independently Challengeable: Publication Authority and a Protocol for Falsifiable Publication Records

    arXiv:2609.17631v1 Announce Type: new Abstract: AI-assisted claims can appear authoritative when evidence, analysis, human authorization, presentation, and correction history refer to different states. Provenance, attest…

  13. arXiv cs.AIresearch
    One Color Preprocessing Improves DSATUR

    arXiv:2609.17633v1 Announce Type: new Abstract: The Graph Coloring Problem (GCP) is NP-hard and DSATUR stands as one of the fastest heuristics for it despite producing colorings that typically use more colors than state-…

  14. arXiv cs.AIresearch
    Physics-Constrained Digital Twins for Sensor Integrity in Urban Pedestrian Flow: Detecting Stealthy False Data Injection with Conformal Guarantees

    arXiv:2609.17635v1 Announce Type: new Abstract: City pedestrian counting systems now feed economic indicators, planning decisions and safety operations, yet the twins built on top of them treat the incoming stream as gro…

  15. arXiv cs.AIresearch
    What You Can't See Is Still What You Learn: A Preregistered Sixty-Society Confirmation That Evidence Masking Drives Compositional Generalization

    arXiv:2609.17637v1 Announce Type: new Abstract: Restricting what a module can read may improve what a system learns to compute. We test this in a preregistered confirmation with sixty four-cell systems sharing a frozen l…

  16. arXiv cs.AIresearch
    GVD: Governed Versioning and Deduplication for Document Repositories

    arXiv:2609.17696v1 Announce Type: new Abstract: Document repositories evolve continuously. Guidelines and policies are revised, superseded, and re-uploaded, so the same content recurs in different wording and newer versi…

  17. arXiv cs.AIresearch
    NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation

    arXiv:2609.17699v1 Announce Type: new Abstract: We present NeMo Data Designer (NDD), an open-source, general-purpose framework for multi-modal synthetic data generation (SDG). Designed to be intuitive to use, NDD provide…

  18. arXiv cs.AIresearch
    A Systematic Evaluation of the COTQ Provincial Land Cover Product: Structural Consistency, Spectral Separability, and Relative Positioning Against ESA, ESRI, and Google Products

    arXiv:2609.17731v1 Announce Type: new Abstract: High-resolution land use and land cover (LULC) products derived from Sentinel-2 imagery are widely used for environmental monitoring and land management, yet their performa…

  19. arXiv cs.AIresearch
    SAGE: Governed Artifact Generation from Enterprise Guidelines

    arXiv:2609.17775v1 Announce Type: new Abstract: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of…

  20. arXiv cs.AIresearch
    A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning

    arXiv:2609.17804v1 Announce Type: new Abstract: Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these…

  21. arXiv cs.AIresearch
    Learning Heterogeneous Preferences

    arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Exist…

  22. arXiv cs.AIresearch
    SNOMED CT Concept Recommendation from Masked Clinical Context

    arXiv:2609.17855v1 Announce Type: new Abstract: Standardizing clinical language to SNOMED CT supports interoperability, analytics, and reusable phenotyping, but concept recommendation remains difficult when relevant conc…

  23. arXiv cs.AIresearch
    The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?

    arXiv:2609.17863v1 Announce Type: new Abstract: LLM inference optimizations report speedups on different models, GPUs, prompts, and quality metrics, making them hard to compare or combine. We build a cost, quality, and l…

  24. arXiv cs.AIresearch
    Do Frontier Models Seek Safety Evidence Before Acting?

    arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to ac…

  25. arXiv cs.AIresearch
    OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning

    arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effecti…

  26. arXiv cs.AIresearch
    Measuring AI Leadership: Development and Validation of a Multidimensional Measure for AI-Native Organizations

    arXiv:2609.17965v1 Announce Type: new Abstract: AI is changing what leaders must judge, explain, learn, and coordinate, yet existing measures do not capture these behaviors at the level needed to study leadership in AI-e…

  27. arXiv cs.AIresearch
    Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AI

    arXiv:2609.17969v1 Announce Type: new Abstract: Long-term memory is becoming a core substrate for personalized AI, yet most systems still represent personalization as discrete records in a largely static latent space, ac…

  28. arXiv cs.AIresearch
    Multimodal Conditioning of Fine-Tuned Stable Diffusion XL for Controllable and Culturally Faithful Ulos Motif Generation

    arXiv:2609.17987v1 Announce Type: new Abstract: The traditional Batak Ulos weaving industry faces growing challenges in producing diverse, innovative motifs due to limitations in conventional, manually driven design meth…

  29. arXiv cs.AIresearch
    Missing Bridges: Composition-Aware Active Imitation Learning

    arXiv:2609.18004v1 Announce Type: new Abstract: Active imitation learning reduces expert effort by allowing a learner to request the demonstrations it needs. Existing methods typically select these requests for their exp…

  30. arXiv cs.AIresearch
    Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning

    arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-po…

  31. arXiv cs.AIresearch
    The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

    arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not t…

  32. arXiv cs.AIresearch
    Teaching AI, Robotics, & Community: A Hubs-Based K-12 Education Framework for Reaching Rural Schools

    arXiv:2609.18072v1 Announce Type: new Abstract: K-12 robotics and AI education remains difficult to scale, especially in rural regions lacking sustained technical mentorship. Programs like FIRST provide competition pathw…

  33. arXiv cs.AIresearch
    Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

    arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features…

  34. arXiv cs.AIresearch
    When Is Graph Structure Worth Its Cost? The Case for Structure Pricing in Retrieval-Augmented Generation

    arXiv:2609.18099v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (RAG) can help answer questions that require information from many documents. However, building a graph often requires many langu…

  35. arXiv cs.AIresearch
    Time-Aligned Evolving Concept Graphs for Scientific Relation Forecasting

    arXiv:2609.18163v1 Announce Type: new Abstract: Forecasting scientific relations can guide discovery by identifying promising connections before they emerge. Existing approaches often model concept semantics and graph st…

  36. arXiv cs.AIresearch
    Re2A: Situated Conversational Recommendation via Rubric-based Preference Reasoning and Alignment

    arXiv:2609.18249v1 Announce Type: new Abstract: Real-world recommendation scenarios are commonly grounded in shared physical environments during user-recommender interactions. This motivates situated conversational recom…

  37. arXiv cs.AIresearch
    REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

    arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora. These challenges often lim…

  38. arXiv cs.AIresearch
    BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

    arXiv:2609.18270v1 Announce Type: new Abstract: Payment operations are a critical financial infrastructure, but the value of large language models in this domain remains unclear because payment rules change quickly, evid…

  39. arXiv cs.AIresearch
    Building Trust in Artificial Intelligence: A Necessity for Railway Applications

    arXiv:2609.18278v1 Announce Type: new Abstract: Artificial Intelligence (AI) is currently only applied to non-safety critical applications due to the strict standards and regulations for railway industries. We propose to…

  40. arXiv cs.AIresearch
    What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

    arXiv:2609.18286v1 Announce Type: new Abstract: Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has rene…

  41. arXiv cs.AIresearch
    Visual Compliance via Executable Safety Rule Entailment

    arXiv:2609.18328v1 Announce Type: new Abstract: Recent advances in LLMs and VLMs have enabled safety systems to reason beyond simple risk patterns toward more contextual and semantic safety concerns. However, as risk pat…

  42. arXiv cs.AIresearch
    Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland

    arXiv:2609.18394v1 Announce Type: new Abstract: We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data,…

  43. arXiv cs.AIresearch
    Risk-Aware World Modeling with Flow-Guided Occupancy Evolution for Selective Trajectory Planning in Automated Driving

    arXiv:2609.18442v1 Announce Type: new Abstract: Safe motion planning in automated driving requires anticipating evolving traffic risks and deciding when to revise the current planned trajectory. We introduce RiskWorld, a…

  44. arXiv cs.AIresearch
    The Mirage of Calibrated Confidence: Trajectory-Independence of Verbalized Confidence in Vision-Language Models

    arXiv:2609.18453v1 Announce Type: new Abstract: A calibrated Vision-Language Model (VLM) can repeatedly self-correct, say "Wait, I should recheck," arrive at the wrong answer, and still report high confidence. We find th…

  45. arXiv cs.AIresearch
    First Token Matters: Understanding Safety Collapse in Large Reasoning Models

    arXiv:2609.18471v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) exhibit strong problem-solving abilities, yet their safety alignment often degrades when handling harmful queries. Existing approaches to impr…

  46. arXiv cs.AIresearch
    Hyperbolic Graph Representation Learning for Differential Diagnosis on Biomedical Knowledge Graphs

    arXiv:2609.18481v1 Announce Type: new Abstract: Biomedical knowledge graphs combine ontology-derived hierarchies with transversal associations among heterogeneous entities such as phenotypes, diseases, genes, proteins, a…

  47. arXiv cs.AIresearch
    Beyond Routine Compliance: Cunning Data Cultivates Safety Vigilance in Large Language Models

    arXiv:2609.18515v1 Announce Type: new Abstract: Safety alignment teaches large language models (LLMs) to recognize harmful requests and reject risky instructions. Yet aligned models can fail when harmful intent is concea…

  48. arXiv cs.AIresearch
    TRIPROBE: Probing Task Separability Beyond Classification for XAI

    arXiv:2609.18525v1 Announce Type: new Abstract: Modern evaluation of learning pipelines often reduces to downstream accuracy, leaving open the question of why tasks succeed or fail. TriProbe addresses this gap with a mul…

  49. arXiv cs.AIresearch
    Reasoning through Evolution: Automatic Meta-path Discovery for LLM-based Fake News Detection

    arXiv:2609.18597v1 Announce Type: new Abstract: Propagation structures provide crucial evidence for fake news detection, yet existing approaches primarily rely on supervised GNN-based models, which require substantial la…

  50. arXiv cs.AIresearch
    The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses

    arXiv:2609.18676v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance rema…