Skip to content
THE AI WIREINTELLIGENCE THAT MATTERS
Research

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Apple Machine Learning Research··Updated just now·34 sightings
AI brief

Researchers propose DACA-GRPO, a reinforcement learning method for diffusion language models that assigns credit to denoising steps rather than treating them all as equally important. The method targets biased, high-variance likelihood estimates in existing RL approaches for diffusion models.

Written by AI from Apple Machine Learning Research's published text. Read the original for full details.