Ornith-397B runs at Q4 on a single RTX PRO 6000; Krasis enables streaming MoE inference with 2,354 tok/s prefill and ~20–24 tok/s decode on one GPU
Read the original at old.reddit.com→I've been building Krasis, an MoE-focused runtime for streaming big models through limited VRAM on NVIDIA consumer/workstation GPUs, and I think this is the most interesting result so far: Ornith-1.0-397B running...
Original headline: "Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode"
Coverage timeline
- Jul 27, 13:59 UTC r/LocalLLaMA lead source Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode
- Jul 30, 06:46 UTC r/LocalLLaMA 4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s