Using OCR Heads to Verbalize Image Semantics
arXiv:2609.18823v1 Announce Type: cross Abstract: How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention
arXiv cs.AI··Updated just now·33 sightings