Is there a better option than llama.cpp for 4GB VRAM to increase tokens per second without PyTorch dependencies?
Read the original at www.reddit.com→I love the idea of running local models on consumer hardware. I currently use `llama.cpp`, but I’ve been looking for a "better" alternative for a while now. I want something that utilizes my system resources more...
Original headline: "Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?"
Coverage timeline
- Oct 10, 13:14 UTC r/LocalLLaMA lead source Is there a better option than llama.cpp for 4GB VRAM for Higher tokens/sec?