LiquidAI/LFM2.5-VL-3B is a multimodal on-device variant of LFM2.5 that processes text and images, uses LFM2.5-2.6B as backbone with SigLIP2 NaFlex encoder, and offers improved grounding, object detection with natural language queries, and full-page OCR with layout annotation.
Read the original at old.reddit.com→LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training. LFM2.5-VL-3B can process both text and images,...
Original headline: "LiquidAI/LFM2.5-VL-3B · Hugging Face"
Coverage timeline
- Aug 12, 14:37 UTC r/LocalLLaMA lead source LiquidAI/LFM2.5-VL-3B · Hugging Face
- Aug 12, 16:50 UTC r/LocalLLaMA CohereLabs/North-Micro-Vision-Instruct · Hugging Face