AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines
arXiv:2609.18123v1 Announce Type: new Abstract: Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustworthy. We characterize four failure modes
arXiv cs.AI··Updated just now·34 sightings