llama.cpp adds --n-cpu-ffn option for Dense models (building on --n-cpu-moe / --cpu-moe for MOE models) via pull request 26622
Read the original at old.reddit.com→PR by u/Stainless-Bacon 👍 It would be handy & awesome to have options --n-cpu-ffn / --cpu-ffn for Dense models like how we have --n-cpu-moe / --cpu-moe for MOE models. Also check his threads: On PR : llama.cpp CPU...
Original headline: "[Open PR] llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp"
Coverage timeline
- Aug 19, 11:40 UTC r/LocalLLaMA lead source [Open PR] llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp