AI brief
NVIDIA's developer blog explains how a 30B-parameter model can activate only 3B parameters per token while still using the capacity of the larger model, using Nemotron 3.5 Lightning as an example.
Why it matters: It lays out the tradeoffs between dense and mixture-of-experts models around active parameters and throughput to help developers choose between them.
Written by AI from NVIDIA Developer Blog's published text. Read the original for full details.