Mage-VL is a codec-native streaming multimodal foundation model with a 4B visual encoder trained from scratch for image and video understanding, addressing real-time streaming perception efficiency
Read the original at old.reddit.com→Mage-VL is a codec-native, proactive-streaming multimodal foundation model for image and video understanding, whose visual encoder is trained entirely from scratch at a compact 4B scale. It targets a modern Moravec's...
Original headline: "microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model"
Coverage timeline
- Jul 28, 18:47 UTC r/LocalLLaMA lead source microsoft/Mage-VL · Hugging Face - An Efficient Codec-Native Streaming Multimodal Foundation Model