DeepSeek V4 Flash gains basic vision by training a 40.1M connector on 100K image-text examples
Read the original at old.reddit.com→I wanted to find out whether a huge text-only MoE could be given basic vision without retraining the language model itself. The short answer is yes. I froze DeepSeek V4 Flash and a 417M-parameter MoonViT image...
Original headline: "I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples"
Coverage timeline
- Aug 11, 03:45 UTC r/LocalLLaMA lead source I gave DeepSeek V4 Flash basic vision by training a 40M connector on 100K examples