Two 5070 Ti GPUs run Qwen 27B full config with cu129-nightly KV cache improvements; 2 concurrent threads maintain speed, decode at 94-87 tps, 0-120k context, prefill 4.6k-2.4k, 170k GPU KV plus 246k with 8GB RAM
Read the original at old.reddit.com→Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-nightly. Which...
Original headline: "2 x 5070ti Qwen 27B full config / stats"
Coverage timeline
- Aug 6, 16:11 UTC r/LocalLLaMA lead source 2 x 5070ti Qwen 27B full config / stats