Llama.cpp updates include chunked SSD matmul acceleration for Mamba-2 prefill and a ggml-metal FWHT kernel for the metal backend
Read the original at old.reddit.com→ggml-cuda: add chunked SSD matmul for Mamba-2 prefill acceleration- #22675 Nemotron-Nano-9B-v2 ub base (scan) branch (SSD) speedup 128 5,404 5,351 −1% (both scan) 256 6,180 7,110 +15% 512 6,627 7,778 +17% ...
Original headline: "It's the small things that matter the most. - llama.cpp - Bunch of updates(Boost & Fixes)"
Coverage timeline
- Jul 28, 14:16 UTC r/LocalLLaMA lead source It's the small things that matter the most. - llama.cpp - Bunch of updates(Boost & Fixes)