CulturalMenuBench probes the knowledge-application gap in multimodal culinary reasoning across 10 languages and 18 regions with 4,870 items and 10 tasks
Read the original at arxiv.org→arXiv:2609.03526v1 Announce Type: new Abstract: Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural...
Original headline: "CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning"
Coverage timeline
- Sep 4, 04:00 UTC arXiv cs.AI lead source CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning