Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Archive

AI news archive

90 articles filed under Chips and compute, newest first, from the outlets listed on the sources page.

All topicsModelsAgentsResearchChips and computeOpen sourceSafetyPolicy and regulationStartups and fundingCommunity
  1. NVIDIA Developer Blogchips
    Giga-Scale AI and the Ethernet Evolution: How Spectrum-X Ethernet Rewrites the Rules

    The massive growth of generative AI has fundamentally altered data center design. As distributed model training scales to span hundreds of thousands of GPUs,...

  2. NVIDIA Blogchips
    GeForce NOW Gives Gamers More Ways to Play at Gamescom 2026

    NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology contr…

  3. NVIDIA Developer Blogchips
    NVIDIA NVLink Fusion Brings NVHBM to Next-Generation AI Infrastructure

    AI factories must support increasingly large models and more complex reasoning workloads. To keep up with the insatiable compute demands of AI workloads,...

  4. OpenAI Newschips
    The full stack behind abundant intelligence

    OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.

  5. OpenAI Newschips
    Jalapeño’s first results show industry-leading speed and efficiency in AI inference

    Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

  6. NVIDIA Developer Blogchips
    CUDA Python 1.0: Stable APIs, One Foundation, Full Platform Access

    For years, a Python developer who needed a GPU had two realistic choices: Learn NVIDIA CUDA C++ well enough to write an extension, set up a build toolchain, and...

  7. NVIDIA Developer Blogchips
    GPU-Accelerated Clustering for Financial Instruments at Scale

    Use AdaptGrow, a GPU-accelerated matrix factorization algorithm, to turn rolling correlation and tail-dependence matrices into hard clusters, soft factor...

  8. NVIDIA Developer Blogchips
    Advancing Semiconductor Innovation Across Materials Engineering and Manufacturing

    As AI workloads increase, explosive compute demand is pushing the semiconductor industry to meet unprecedented performance targets. Even small delays can have...

  9. NVIDIA Developer Blogchips
    NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

    NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they...

  10. NVIDIA Developer Blogchips
    NVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure

    Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We...

  11. NVIDIA Developer Blogchips
    Run High-Performance Core Math at Scale with NVIDIA nvmath-python

    NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users...

  12. NVIDIA Developer Blogchips
    NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek

    The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,...

  13. NVIDIA Developer Blogchips
    How to Run Isolated Tenant Kubernetes Clusters on Shared GPU Infrastructure

    Running a dedicated Kubernetes cluster per team often results in more isolation than an organization requires. While one cluster can be successfully shared...

  14. NVIDIA Developer Blogchips
    How to Choose Full-Stack Observability for NVIDIA AI Factories

    AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades, identifying the...

  15. NVIDIA Developer Blogchips
    Run Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy

    Uniform Manifold Approximation and Projection (UMAP) is a dimensionality reduction technique widely used for visualization and feature extraction. Applications...

  16. NVIDIA Developer Blogchips
    Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer

    Teams customize their models to hit their targets for latency, speed, memory, and compute. With the open NVIDIA Nemotron family of models, developers can find...

  17. NVIDIA Developer Blogchips
    Post-Train NVIDIA Cosmos 3 Edge for On-Device Robot Control

    Robots need policies that can adapt to their sensors, environments, and tasks while running on onboard computing hardware. World models offer a foundation for...

  18. Hugging Face Blogchips
    Up to 3.2x Faster Inference with LFM2.5-DSpark

    Open-source model releases, tooling, and community updates.

  19. Hugging Face Blogchips
    Same Cluster, 33 Points More Utilization: What Changed Was the Order

    Open-source model releases, tooling, and community updates.

  20. NVIDIA Developer Blogchips
    NVIDIA NVLink: The Scale-Up Network for AI Factories

    The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute...

  21. Mistral AI Newschips
    In-region inference, open models, and new European infrastructure for sovereign AI.

    Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, and setting a roadmap for the world.

  22. NVIDIA Developer Blogchips
    Maximize Spectral Efficiency with AI-Native RAN and NVIDIA AI Aerial

    Spectrum is one of the most valuable assets in wireless communications. Over the last 30 years, telecom operators in the US have spent more than $240B to...

  23. NVIDIA Developer Blogchips
    Running Low-Latency Analytical Workloads with GPU-Accelerated Presto on NVIDIA GB200 NVL72

    Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance...

  24. NVIDIA Developer Blogchips
    AI Model Co-Design: Hardware-Friendly LLM Design

    AI performance comes down to three dimensions: Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a...

  25. NVIDIA Developer Blogchips
    Reducing High-Bandwidth Memory Bottlenecks in JAX-Based LLM Training with Host Offloading

    Large language model (LLM) training workloads increasingly run into GPU memory limits before compute is fully used. Model weights, gradients, optimizer states,...

  26. NVIDIA Developer Blogchips
    Kernel Fusion in NVIDIA CUDA: Optimizing Memory Traffic and Launch Overhead

    There are many ways to optimize code for GPUs. In this post, you’ll learn how kernel fusion can improve memory bandwidth and reduce kernel launch overhead,...

  27. NVIDIA Developer Blogchips
    Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning

    The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when...

  28. NVIDIA Developer Blogchips
    Building Faster Cryptography with Carryless Multiplication in NVIDIA CUDA 13.3

    For over fifteen years, x86 CPUs have shipped with a dedicated hardware instruction for carryless multiplication. It’s a small but stubborn primitive that...

  29. NVIDIA Developer Blogchips
    Build a Multi-Camera 3D Tracking Application with NVIDIA DeepStream 9.1 Skills

    Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking...

  30. NVIDIA Developer Blogchips
    Integrate NVIDIA Omniverse RTX Sensor Simulation Into Existing Apps

    Developers building 3D, design, simulation, robotics, and industrial digital twin applications need ways to bring physical AI capabilities into the tools and...

  31. NVIDIA Developer Blogchips
    Make Long-Running NVIDIA TensorRT Engine Builds Observable and Cancelable in Python or C++

    A TensorRT engine build can take seconds to many minutes. Large strongly typed models, deep tactic search, and a cold timing cache on a brand-new GPU SKU can...

  32. NVIDIA Developer Blogchips
    Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

    Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token...

  33. NVIDIA Developer Blogchips
    Debugging Ray Tracing Applications Using NVIDIA OptiX Toolkit

    NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways...

  34. NVIDIA Developer Blogchips
    Start Customizing NVIDIA Nemotron 3 Nano with Prime Intellect Lab in Minutes

    Customization is what enables developers to take a general model and tailor it to use cases, domains, languages, and more. However, customization comes with a...

  35. Hugging Face Blogchips
    Baseten on Hugging Face Inference Providers 🔥

    Open-source model releases, tooling, and community updates.

  36. OpenAI Newschips
    How we built a realtime system for responsive voice AI in six months

    GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.

  37. OpenAI Newschips
    Ten advances in mathematics and theoretical computer science

    OpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.

  38. Hugging Face Blogchips
    GPU Management: Why Idle GPUs Are the New Grounded Aircraft

    Open-source model releases, tooling, and community updates.

  39. Hugging Face Blogchips
    NVIDIA Cosmos-H-Dreams: Bringing Real-Time Generative Simulation to Surgical Robotics

    Open-source model releases, tooling, and community updates.

  40. Hugging Face Blogchips
    Bringing Nunchaku 4-bit Diffusion Inference to Diffusers

    Open-source model releases, tooling, and community updates.