Reasoning with memory: a temporal granularity-adaptive framework for training-free long video understanding
Read the original at arxiv.org→arXiv:2607.24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video...
Original headline: "Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding"
Coverage timeline
- Jul 29, 04:00 UTC arXiv cs.AI lead source Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding