Energy Vision–Language–Action: a controlled multimodal benchmark for intent-conditioned residential energy management
Read the original at arxiv.org→arXiv:2609.31648v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are studied mainly in robotics, where visual observations and language instructions are mapped to physical actions. This paper...
Original headline: "Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management"
Coverage timeline
- Sep 29, 04:00 UTC arXiv cs.LG lead source Energy Vision--Language--Action: A Controlled Multimodal Benchmark for Intent-Conditioned Residential Energy Management