ConflictVLA-Bench benchmarks behavioral responses of vision-language-action models to premise conflicts
Read the original at arxiv.org→arXiv:2609.31792v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations...
Original headline: "ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts"
Coverage timeline
- Sep 29, 04:00 UTC arXiv cs.AI lead source ConflictVLA-Bench: Benchmarking Behavioral Responses of Vision-Language-Action Models to Premise Conflicts