Ninfer Stahp runs qwen3.8-27b-nvfp4 on Blackwell MaxQ, delivering 130–180 t/s and strong parallel decoding performance vs llama-server, with Windows 10 setup underway
Read the original at www.reddit.com→GPU: Blackwell MaxQ Model: qwen3.8-27b-nvfp4 (Orcarouter) - Dflash2 - MTP-7 Running this on my Blackwell MaxQ. I can't believe how fucking fast it is compared to llama-server. It easily reaches speeds between...
Original headline: "Ninfer Stahp"
Coverage timeline
- Oct 2, 04:06 UTC r/LocalLLaMA lead source Ninfer Stahp