Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
arXiv:2609.18723v1 Announce Type: new Abstract: We introduce Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning. Existing selection strategies rely predom
arXiv cs.AI··Updated just now·33 sightings