CapMem: a benchmark for caption-based episodic memory in egocentric video
Read the original at arxiv.org→arXiv:2609.17688v1 Announce Type: new Abstract: Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and...
Original headline: "CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video"
Coverage timeline
- Sep 17, 04:00 UTC arXiv cs.AI lead source CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video