A Zeroth-Order Paradigm for LLM Preference Alignment
arXiv:2609.19144v1 Announce Type: cross Abstract: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates
arXiv cs.AI··Updated just now·38 sightings