Gemma 4 e2b q4 with 32k context and 16GB VRAM used to control a wifi-connected robot car; user discusses throughput and Ollama integration
Read the original at www.reddit.com→Im thinking of using Gemma 4 e2b q4, running on 32k context with 16gb vram and 900GB/s bandwidth. I would be running it using a call to ollama(idrc about optimize, ill have hundreds of tok/s no matter what) and have...
Original headline: "Any advice?"
Coverage timeline
- Oct 8, 08:53 UTC r/LocalLLaMA lead source Any advice?