Native long video understanding models locally?
Read the original at old.reddit.com→I've been building a personal project and wanted to check with the community on multi-modal inputs since I can't find a lot of material around this online. Ultimately I'm trying to build something that can ingest...
Original headline: "Native Long Video Understanding Models locally?"
Coverage timeline
- Aug 10, 12:06 UTC r/LocalLLaMA lead source Native Long Video Understanding Models locally?