Higher-order pruning of experts in mixture-of-experts language models
arXiv:2609.18916v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet exist
arXiv cs.AI··Updated just now·34 sightings