POCKET-35B agentic model runs on CPU at 59 t/s using GGUF on stock llama.cpp
Read the original at old.reddit.com→https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF A 35B model that runs on your PC with no GPU — and on your phone. Just stock llama.cpp. No fork, no CUDA, no cloud. The POCKET lineup — pick by your device Repo...
Original headline: "POCKET-35B agentic model on cpu 59 t/s"
Coverage timeline
- Jul 26, 10:10 UTC r/LocalLLaMA lead source POCKET-35B agentic model on cpu 59 t/s