AI news archive
294 articles filed under Research, newest first, from the outlets listed on the sources page.
- arXiv cs.AIresearchThe Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses
arXiv:2609.18676v1 Announce Type: new Abstract: Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance rema…
- arXiv cs.AIresearchBeyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. E…
- arXiv cs.AIresearchWhich LLM is Best for Translating Natural Language Goals to PDDL
arXiv:2609.18731v1 Announce Type: new Abstract: Bridging the gap between human intent and machine execution remains a challenge in automated planning, where expressing goals in formal languages like PDDL restricts access…
- arXiv cs.AIresearchClueing up LLMs with Tool-Augmented Deductive Reasoning
arXiv:2609.18736v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), performing logically consistent deductive reasoning over extended interactions remains challenging. Tasks that requ…
- arXiv cs.AIresearchVersion- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale
arXiv:2609.18769v1 Announce Type: new Abstract: Correctly answering a question grounded in normative documents often depends on information outside any single passage: whether the retrieved document is the version curren…
- arXiv cs.AIresearchInfinite-Parameter LLMs: Generating and Adapting Weights from Live Data
arXiv:2609.18842v1 Announce Type: new Abstract: The scaling laws hold that a language model grows more capable with more parameters and more training data, and Mixture-of-Experts (MoE) architectures have ridden these law…
- arXiv cs.AIresearchSuppressed, Not Erased: A Representational Trace of Edited Facts Survives Even Weight-Free Knowledge Editing
arXiv:2609.18985v1 Announce Type: new Abstract: Knowledge-editing benchmarks certify local correctness, whether an edited model produces the new fact on near-edit prompts but not how much of the original fact remains dec…
- arXiv cs.AIresearchFunction Lives Where Variance Doesn't: Task-Weighted Charts of a Language Model's Computation
arXiv:2609.18989v1 Announce Type: new Abstract: How many dimensions does a language model's computation actually use? The question is ill-posed until one names a functional. Task-weighted charts make it well-posed: low-d…
- arXiv cs.AIresearchLost in Perception: Isolating Perceptual and Reasoning Failures in Multimodal Physics and Geometry Reasoning
arXiv:2609.18991v1 Announce Type: new Abstract: Multimodal LLMs report strong performance on scientific reasoning benchmarks, yet most treat perception and reasoning as a single measurable process. We introduce a five-ta…
- arXiv cs.AIresearchMUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
arXiv:2609.19088v1 Announce Type: new Abstract: Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated.…
- arXiv cs.AIresearchThe Death of Schema Linking? Text-to-SQL in the Age of Well-Reasoned Language Models
arXiv:2408.07702v2 Announce Type: cross Abstract: Schema linking is a crucial step in Text-to-SQL pipelines. Its goal is to retrieve the relevant tables and columns of a target database for a user's query while disregard…
- arXiv cs.AIresearchREQAP: Resilient Weight Packing and Quantization for Edge DNN Acceleration
arXiv:2609.17555v1 Announce Type: cross Abstract: Efficient deployment of Deep Neural Networks (DNNs) on edge accelerators requires aggressive model compression while maintaining reliability in fault-prone hardware envir…
- arXiv cs.AIresearchWARD: Runtime Workload-Adaptive Vision TRansformer Framework for Dependable Edge AI
arXiv:2609.17556v1 Announce Type: cross Abstract: Edge-deployed AI operate under dynamically changing power budgets, reliability requirements, and input distributions, requiring continuous adaptation. Such conditions ari…
- arXiv cs.AIresearchPay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds
arXiv:2609.17560v1 Announce Type: cross Abstract: Every production model is updated, by retraining, fine-tuning, quantization, or a silent vendor swap, and each update risks being worse than what it replaced. We formaliz…
- arXiv cs.AIresearchBLADE: ReliaBle Dynamic Hardware-Aware SNN-ANN Boundary SeLection for Event-BAseD Object DEtection
arXiv:2609.17562v1 Announce Type: cross Abstract: Hybrid Spiking Neural Network (SNN)-Artificial Neural Network (ANN) architectures combine the energy efficiency of SNNs with the superior detection accuracy of ANNs for e…
- arXiv cs.AIresearchIndependence-System Realisations in Single-Source Unsplittable Flow
arXiv:2609.17568v1 Announce Type: cross Abstract: Additive-congestion constraints in single-source unsplittable flow can enforce stable-set structure. This note isolates and generalises that mechanism. We introduce a pat…
- arXiv cs.AIresearchWhere Grokking Happens: Distributed Utility and Fourier Recoding Without a Module Switch
arXiv:2609.17571v1 Announce Type: cross Abstract: Where in a Transformer is the change from memorization to generalization functionally expressed? We introduce Transition Games--behavior-aligned exact activation games wi…
- arXiv cs.AIresearchEvolutionary Ensemble Search: Council-Guided Program Evolution with Persistent Memory
arXiv:2609.17590v1 Announce Type: cross Abstract: Evolutionary Ensemble Search (EES) constructs machine-learning procedures through expert-guided program evolution. A role-specialized council turns task evidence and expe…
- arXiv cs.AIresearchStructure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
arXiv:2609.17599v1 Announce Type: cross Abstract: A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foun…
- arXiv cs.AIresearchLecture notes on Physics Informed Neural Networks, Neural Operators, and their applications
arXiv:2609.17638v1 Announce Type: cross Abstract: This is the set of lecture notes for the PhD course \href{https://www.unibz.it/en/faculties/engineering/phd-computer-science/study-course-offering/2025/36967}{\textit{Phy…
- arXiv cs.AIresearchScaling Articulated Rationales for MLLM-based Recommendation
arXiv:2609.17639v1 Announce Type: cross Abstract: Modern recommendation systems largely infer user preferences from implicit behaviors such as clicks, watch time, and negative feedback, but these signals reveal what user…
- arXiv cs.AIresearchRethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models
arXiv:2609.17644v1 Announce Type: cross Abstract: Domain-specialized language models are widely used for scientific question answering, but stronger general-purpose systems raise a sharper question: when does domain-spec…
- arXiv cs.AIresearchThe Missing "I Don't Know": Why Three Reasoning-Reliability Findings Converge on Calibrated Abstention
arXiv:2609.17686v1 Announce Type: cross Abstract: Three recent results describe what look like unrelated LLM reliability problems. Yin et al. (2026) show reasoning RL collapses tool-reliability representations. Suleymano…
- arXiv cs.AIresearchAccelerating Diffusion Sampling via Speculative Draft Trees
arXiv:2609.17691v1 Announce Type: cross Abstract: Speculative sampling accelerates diffusion model generation by drafting inexpensive candidate states and correcting them under a coupling that preserves the target distri…
- arXiv cs.AIresearchOne Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG
arXiv:2609.17709v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems typically employ fixed retriever and generator configurations across queries, despite substantial differences in query comple…
- arXiv cs.AIresearchEvolution of US Oral Political Language
arXiv:2609.17755v1 Announce Type: cross Abstract: The analysis of US political language is usually based on the written form (e.g. presidential addresses) or posts broadcasted on various social networks. Oral production,…
- arXiv cs.AIresearchCALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors
arXiv:2609.17758v1 Announce Type: cross Abstract: Deep Reinforcement Learning has demonstrated remarkable capability in quadrotor control, yet learned policies offer no guarantee of respecting safety constraints during t…
- arXiv cs.AIresearchIs Luke the Author of a Gospel and the Acts of the Apostles?
arXiv:2609.17762v1 Announce Type: cross Abstract: According to Christian tradition, Luke is credited with authoring a Gospel and the Acts of the Apostles, even if his name does not appear in either book, both originally…
- arXiv cs.AIresearchHINT-Plan: Human Intention-Aware Robot Task Planning in Context-Rich Environments using Vision Language Models
arXiv:2609.17771v1 Announce Type: cross Abstract: Approaches to incorporating human awareness into mobile robot decision-making mainly focus on collision avoidance in low-level motion planning, often overlooking the chal…
- arXiv cs.AIresearchWhen AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments
arXiv:2609.17772v1 Announce Type: cross Abstract: AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature…
- arXiv cs.AIresearchInformation Set Emulation: Causal Certificates for AI Derived EHR Features
arXiv:2609.17777v1 Announce Type: cross Abstract: AI and large language models can recover clinically meaningful features from electronic health records (EHRs), but predictive usefulness does not establish admissibility…
- arXiv cs.AIresearchAI and Human Approaches to Mathematical Problem Solving
arXiv:2609.17779v1 Announce Type: cross Abstract: AI systems have begun to report solutions, disproofs, and substantive advances on long-standing mathematical problems, raising questions about whether they approach resea…
- arXiv cs.AIresearchSAiFE-gym: Model-based Environments for Automated Market Making with Concentrated Liquidity
arXiv:2609.17788v1 Announce Type: cross Abstract: We present SAiFE_gym, a Python module that provides a collection of simulation environments for studying trading problems in Constant Product Markets (CPMs) with Concentr…
- arXiv cs.AIresearchQiT: Quantum-Inspired Transformer for Visual Recognition Task
arXiv:2609.17789v1 Announce Type: cross Abstract: Quantum machine learning offers a compelling representational perspective: angle-encoded states inhabit Hilbert spaces in which periodic similarities and interactions can…
- arXiv cs.AIresearchPrincipled Koopman Representations with Kalman Inference for Efficient Time-Series Prediction
arXiv:2609.17815v1 Announce Type: cross Abstract: The Koopman operator has been widely used for time-series prediction in dynamical systems. However, prior work that learns latent ``Koopman spaces'' using neural networks…
- arXiv cs.AIresearchThe Free Inference Dimension: Complexity Measure for Zero-Collision Navigation under Hypothesis Mixtures
arXiv:2609.17816v1 Announce Type: cross Abstract: Solomonoff induction frames prediction as a mixture over computable hypotheses, typically leading to identification of the true environment. In a finite meta-reinforcemen…
- arXiv cs.AIresearchLearning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control
arXiv:2609.17824v1 Announce Type: cross Abstract: We study cooperative multi-humanoid pickup and transport of objects with varying size, weight, and geometry, requiring robot teams of different sizes. Our approach uses d…
- arXiv cs.AIresearchProcedural Pretraining for Molecular Property Prediction
arXiv:2609.17831v1 Announce Type: cross Abstract: Molecular property prediction is often limited by the small size of labeled downstream datasets, motivating pretraining on large corpora of unlabeled molecules. In this w…
- arXiv cs.AIresearchAdaptive hybrid coupling with operator inference, the overlapping Schwarz alternating method and reinforcement learning
arXiv:2609.17837v1 Announce Type: cross Abstract: Hybrid domain decomposition methods provide a flexible framework for coupling full order models (FOMs) and reduced order models (ROMs), but typically assume the model ass…
- arXiv cs.AIresearchLearning Nuclear Structure with AI: Radii and Collectivity
arXiv:2609.17838v1 Announce Type: cross Abstract: Low-energy nuclear structure is encoded in a broad body of experimental information across the chart of nuclides. Learning how this information is organized across observ…
- arXiv cs.AIresearchRoboVAD: A Large Cross-Domain Evaluation Benchmark for Anomaly Detection in Robotic Arm Manipulation Videos
arXiv:2609.17843v1 Announce Type: cross Abstract: Video anomaly detection (VAD) is an actively studied task, having wide applications in typical scenarios such as public surveillance and road traffic safety. The task is…
- arXiv cs.AIresearchAfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content
arXiv:2609.17853v1 Announce Type: cross Abstract: AfriSyCo studies answer switching around African-language factual content with two complementary layers: native-language follow-ups and a controlled cross-language factor…
- arXiv cs.AIresearchWho Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels
arXiv:2609.17857v1 Announce Type: cross Abstract: Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families…
- arXiv cs.AIresearchDoes AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming
arXiv:2609.17883v1 Announce Type: cross Abstract: The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccura…
- arXiv cs.AIresearchWalking the Score Manifold: Continuous-time Generative Dynamics on Learned Data Manifolds
arXiv:2609.17901v1 Announce Type: cross Abstract: Generative modeling of time-dependent data is typically formulated on a discrete temporal grid, restricting supervision to the observed timestamps in the training data. W…
- arXiv cs.AIresearchEDCT-Bench: Uncovering Faithfulness Gaps in VLMs via Explanation-Driven Counterfactual Testing
arXiv:2609.17953v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) can produce Natural Language Explanations (NLEs) that sound plausible yet remain inconsistent with the visual evidence they cite. We present…
- arXiv cs.AIresearchThe Attention Within: Consensus Dynamics in Selective State Space Models
arXiv:2609.17997v1 Announce Type: cross Abstract: Selective state space models (SSMs) have recently emerged as a compelling alternative to transformers, combining competitive performance with substantially improved infer…
- arXiv cs.AIresearchNewer Is Not Fairer: Gender Stereotyping in Text-to-Image AI Across Model Generations
arXiv:2609.18007v1 Announce Type: cross Abstract: Text-to-image generative models are widely used in professional and creative settings, yet how they represent gender across occupations -- and whether newer models are fa…
- arXiv cs.AIresearchPhysics-Informed Neural Networks for Fast Multilayer Spectral Inversion of H{\alpha} 6562.8 A and Ca II 8542.1 A Spectra
arXiv:2609.18025v1 Announce Type: cross Abstract: Strong chromospheric absorption lines such as H$\alpha$ 6562.8 A and Ca II 8542.1 A provide vital diagnostics of plasma dynamics and thermal structure in the solar chromo…
- arXiv cs.AIresearchAn Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks
arXiv:2609.18052v1 Announce Type: cross Abstract: This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (…