CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these pract
arXiv cs.AI··Updated just now·29 sightings