HyQuant: Hybrid-Precision Quantization for LLM Attention
arXiv:2608.27875v3 Announce Type: replace Abstract: Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introduces large errors at very low
arXiv cs.AI··Updated just now·38 sightings