Ornith-397B runs at Q4 on a single RTX PRO 6000; Krasis enables streaming MoE inference with 2,354 tok/s prefill and ~20–24 tok/s decode on one GPU
Read the original at old.reddit.com→I've been building Krasis, an MoE-focused runtime for streaming big models through limited VRAM on NVIDIA consumer/workstation GPUs, and I think this is the most interesting result so far: Ornith-1.0-397B running...
Original headline: "Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode"
Coverage timeline
- Jul 27, 13:59 UTC r/LocalLLaMA lead source Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode