Do multimodal LLMs see before they read? Diagnosing contextual sycophancy in multimodal reasoning
Read the original at arxiv.org→arXiv:2609.00067v1 Announce Type: new Abstract: External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case...
Original headline: "Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.CL lead source Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy