Learning Heterogeneous Preferences
arXiv:2609.17847v1 Announce Type: new Abstract: Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \emph{uni
arXiv cs.AI··Updated just now·33 sightings