Anchoring What Matters: A Dual-Level Learning Framework for Visually-Grounded Multimodal Reasoning
arXiv:2609.18057v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capabilities of large vision-language models (LVLMs). However, standard on-policy RLVR algorithms face a critical opt
arXiv cs.AI··Updated just now·34 sightings