Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

A Zeroth-Order Paradigm for LLM Preference Alignment

arXiv:2609.19144v1 Announce Type: cross Abstract: Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates

arXiv cs.AI··Updated just now·38 sightings