Improve token per second (tps) to 40 without changing quant settings; user reports 30tps on RX 6700 XT with llama.cpp Vulkan, seeking performance tweaks while keeping q8 and q4_k_xl unchanged
Read the original at www.reddit.com→Spent the past month tweaking and experimenting with many different numbers to achieve 30tps. Hardware: -Rx6700xt 12gb vram (AMD) -2x16 ddr4 3200 ram -r5 5600x -llama.cpp vulkan sdk -window11 (no wsl switching since...
Original headline: "Improve token per second without touching quant"
Coverage timeline
- Oct 10, 12:05 UTC r/LocalLLaMA lead source Improve token per second without touching quant