AI brief
Researchers at Apple Machine Learning Research propose evaluating video caption quality through multiple-choice question answering rather than matching generated text against ground-truth references.
Why it matters: The reference-matching paradigm is limited by the one-to-many nature of video description.
Written by AI from Apple Machine Learning Research's published text. Read the original for full details.