Strata demonstrates 5-10x prefill and 3-4x decode speed improvements in llama.cpp with MoE caching; llama.cpp previously showed little progress, arXiv papers cited but no implementation
Read the original at www.reddit.com→So, I've been begging llama.cpp to do MoE caching for about a year, and watching them d ck around with 1% here and 2% improvements there instead... Until Strata (https://github.com/Niko1221/Strata) clowned llama with...
Original headline: "Imma just say it, Strata absolutely clowned llama.cpp"
Coverage timeline
- Oct 3, 14:05 UTC r/LocalLLaMA lead source Imma just say it, Strata absolutely clowned llama.cpp