Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE run in a browser tab on a 24 GB Mac; experts streamed from disk, output matches llama.cpp
Read the original at www.reddit.com→This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab. LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no...
Original headline: "Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp"
Coverage timeline
- Oct 6, 06:41 UTC r/LocalLLaMA lead source Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp