TokenSwap benchmarks and reduces the modality gap in multimodal LLMs
Read the original at arxiv.org→arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic...
Original headline: "TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs"
Coverage timeline
- Aug 3, 04:00 UTC arXiv cs.CL lead source TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs