CUDA: fuse shared experts into MMVQ in ggml-org/llama.cpp; MoE speedup shown for select architectures like Qwen 35B A3B
Read the original at www.reddit.com→MoE speedup, but only for some MoE architectures (like Qwen 35B A3B) submitted by /u/jacek2023 [link] [comments]
Original headline: "CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp"
Coverage timeline
- Oct 2, 18:47 UTC r/LocalLLaMA lead source CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp
- Oct 3, 06:18 UTC r/LocalLLaMA qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp