1M-context with a 17 GB model on 24 GB VRAM achieved on RTX 3090; user manu69x reports stable 1M context with Qwen 3.5 35B A3B and extraction of 7 needles from text
Read the original at old.reddit.com→https://preview.redd.it/xxjh11f38jih1.png?width=1852&format=png&auto=webp&s=76850ed51e29a8bc86c2ca718d4320075eed4363 Just wanted to share a user report that I found to be very interesting. Some person with an...
Original headline: "1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text""
Coverage timeline
- Aug 10, 11:38 UTC r/LocalLLaMA lead source 1M context with 17 GB model in 24 GB VRAM: "for the first time I was able to load a context of almost 1M tokens and extract 7 needles from various parts of the text"