Nifer achieves 700 t/s with Qwen 3.6 35B in purpose-built RTX5090 setup with full 250k context; No thinking mode demonstrated on Windows deployment
Read the original at old.reddit.com→I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agents in parallel....
Original headline: "Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too."
Coverage timeline
- Jul 27, 19:17 UTC r/LocalLLaMA lead source Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.