Text, pixels, or both? Evaluating input representations for multimodal document QA
Read the original at arxiv.org→arXiv:2609.22628v1 Announce Type: new Abstract: Every document QA system begins with a choice that is rarely studied on its own: whether to feed the model page images, extracted text, or both. We isolate this...
Original headline: "Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA"
Coverage timeline
- Sep 22, 04:00 UTC arXiv cs.AI lead source Text, Pixels, or Both? Evaluating Input Representations for Multimodal Document QA