460M VLM achieves first-token latency of 0.3s on iPhone using only 64 visual tokens
Read the original at old.reddit.com→A new 460M vision model called VisionPsy-Nano-460M-Flash is taking a slightly different approach to on-device VLM speed. Instead of making the language model much smaller, it reduces how much visual information...
Original headline: "A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens"
Coverage timeline
- Aug 5, 14:01 UTC r/LocalLLaMA lead source A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens