AI news archive
90 articles filed under Chips and compute, newest first, from the outlets listed on the sources page.
- NVIDIA Developer BlogchipsGiga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules
The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,...
- NVIDIA BlogchipsGeForce NOW Gives Gamers More Ways to Play at Gamescom 2026
NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology contr…
- NVIDIA Developer BlogchipsNVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure
AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,...
- OpenAI NewschipsThe full stack behind abundant intelligence
OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.
- OpenAI NewschipsJalapeño’s first results show industry-leading speed and efficiency in AI inference
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
- NVIDIA Developer BlogchipsCUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access
For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and...
- NVIDIA Developer BlogchipsGPU-Accelerated Clustering for Financial Instruments at Scale
Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor...
- NVIDIA Developer BlogchipsAdvancing Semiconductor Innovation Across Materials Engineering and Manufacturing
As AI workloads increase, explosive compute demand is pushing the semiconductor industry to meet unprecedented performance targets. Even small delays can have...
- NVIDIA Developer BlogchipsNVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they...
- NVIDIA Developer BlogchipsNVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure
Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We...
- NVIDIA Developer BlogchipsRun High-Performance Core Math at Scale with NVIDIA nvmath-python
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users...
- NVIDIA Developer BlogchipsNVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek
The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,...
- NVIDIA Developer BlogchipsHow to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure
Running a dedicated Kubernetes cluster per team often results in more isolation than an organization requires. While one cluster can be successfully shared...
- NVIDIA Developer BlogchipsHow to Choose Full-Stack Observability for NVIDIA AI Factories
AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades, identifying the...
- NVIDIA Developer BlogchipsRun Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy
Uniform Manifold Approximation and Projection (UMAP) is a dimensionality reduction technique widely used for visualization and feature extraction. Applications...
- NVIDIA Developer BlogchipsDeveloping Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer
Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models, developers can find...
- NVIDIA Developer BlogchipsPost-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control
Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for...
- Hugging Face BlogchipsUp to 3.2x Faster Inference with LFM2.5-DSpark
Open-source model releases, tooling, and community updates.
- Hugging Face BlogchipsSame Cluster, 33 Points More Utilization: What Changed Was the Order
Open-source model releases, tooling, and community updates.
- NVIDIA Developer BlogchipsNVIDIA NVLink: The Scale-Up Network for AI Factories
The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute...
- Mistral AI NewschipsIn-region inference, open models, and new European infrastructure for sovereign AI.
Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, and setting a roadmap for the world.
- NVIDIA Developer BlogchipsMaximize Spectral Efficiency with AI-Native RAN and NVIDIA AI Aerial
Spectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to...
- NVIDIA Developer BlogchipsRunning Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72
Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance...
- NVIDIA Developer BlogchipsAI Model Co-Design: Hardware-Friendly LLM Design
AI performance comes down to three dimensions: Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a...
- NVIDIA Developer BlogchipsReducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading
Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,...
- NVIDIA Developer BlogchipsKernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead
There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead,...
- NVIDIA Developer BlogchipsLessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when...
- NVIDIA Developer BlogchipsBuilding Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3
For over fifteen years, x86 CPUs have shipped with a dedicated hardware instruction for carryless multiplication. It’s a small but stubborn primitive that...
- NVIDIA Developer BlogchipsBuild a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills
Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking...
- NVIDIA Developer BlogchipsIntegrate NVIDIA Omniverse RTX Sensor Simulation Into Existing Apps
Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and...
- NVIDIA Developer BlogchipsMake Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++
A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can...
- NVIDIA Developer BlogchipsSetting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...
- NVIDIA Developer BlogchipsDebugging Ray Tracing Applications Using NVIDIA OptiX Toolkit
NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways...
- NVIDIA Developer BlogchipsStart Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes
Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a...
- Hugging Face BlogchipsBaseten on Hugging Face Inference Providers 🔥
Open-source model releases, tooling, and community updates.
- OpenAI NewschipsHow we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
- OpenAI NewschipsTen advances in mathematics and theoretical computer science
OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.
- Hugging Face BlogchipsGPU Management: Why Idle GPUs Are the New Grounded Aircraft
Open-source model releases, tooling, and community updates.
- Hugging Face BlogchipsNVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics
Open-source model releases, tooling, and community updates.
- Hugging Face BlogchipsBringing Nunchaku 4-bit Diffusion Inference to Diffusers
Open-source model releases, tooling, and community updates.