Asus ESC4000 G3 with 8× T4 GPUs: fastest way to deploy Qwen3.5-9B for 10–15 concurrent users (current setup uses llama.cpp)
Read the original at www.reddit.com→Hi All: I have an Asus ESC 4000 G3 with 128 GB DDR4 RAM - I tried putting in my V620s but couldn’t put more than 2. Sadly, pivoted to T4s, these are 72 watt passive cards and I thought I could use them like how I use...
Original headline: "I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp"
Coverage timeline
- Oct 10, 03:35 UTC r/LocalLLaMA lead source I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp