The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
arXiv:2609.18063v1 Announce Type: new Abstract: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloa
arXiv cs.AI··Updated just now·34 sightings