AI brief
Researchers propose DACA-GRPO, a reinforcement learning method for diffusion language models that assigns credit to denoising steps rather than treating them all as equally important. The method targets biased, high-variance likelihood estimates in existing RL approaches for diffusion models.
Written by AI from Apple Machine Learning Research's published text. Read the original for full details.