Qwen3.8 achieves 50 t/s on Strata with 12GB VRAM and 64GB RAM laptop; Strata inference engine reports 1500 t/s PP on 64GB RAM featud?
Read the original at www.reddit.com→I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling...
Original headline: "Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine"
Coverage timeline
- Sep 30, 04:01 UTC r/LocalLLaMA lead source Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine