Learning to predict middle-layer attention in MLLMs for visual token pruning
Read the original at arxiv.org→arXiv:2608.06411v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing...
Original headline: "Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.AI lead source Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin