OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data u
arXiv cs.AI··Updated just now·34 sightings