SenseNova-Vision is a 7B open model that performs segmentation, depth, detection, OCR, and 3D reconstruction with no task-specific heads
Read the original at old.reddit.com→Stumbled across this new vision model, SenseNova-Vision. It's a 7B MoT model, Apache 2.0 license, which is cool. The main idea is it treats pretty much all computer vision stuff as just one generation problem. Like,...
Original headline: "SenseNova-Vision: a 7B open model that does segmentation, depth, detection, OCR, and 3D reconstruction with no task-specific heads"
Coverage timeline
- Aug 13, 15:07 UTC r/LocalLLaMA lead source SenseNova-Vision: a 7B open model that does segmentation, depth, detection, OCR, and 3D reconstruction with no task-specific heads