GLM5.3 Flash gains Project Maya acceleration, enabling private 321-billion-parameter GLM-5.3 on a single GPU to 30 tok/s
Read the original at www.reddit.com→I stumbled across this as I was currently having glm5.3 flash run only around 10tok/s basically unusable. With Maya i am running 30 tok/s now. GLM5.3 feels even with this low quant like a much more enjoyable model...
Original headline: ""Strata" for GLM5.3 Flash is here for some! Project Maya"
Coverage timeline
- Oct 11, 18:42 UTC r/LocalLLaMA lead source "Strata" for GLM5.3 Flash is here for some! Project Maya