Dual Radeon AI PRO R9700 server slower than RTX 5090 for LLM inference; Ollama bottleneck, vLLM/llama.cpp and other recommendations?
Read the original at old.reddit.com→Hey All, First sorry for long post, and yes i dictated to AI and got it to fix grammer so its not painful for you all to read. Im trying to work out the optimal inference stack for a dedicated local AI server and...
Original headline: "Dual Radeon AI PRO R9700 server much slower than RTX 5090 for LLM inference, Ollama bottleneck? vLLM / llama.cpp / other recommendations?"
Coverage timeline
- Aug 9, 16:05 UTC r/LocalLLaMA lead source Dual Radeon AI PRO R9700 server much slower than RTX 5090 for LLM inference, Ollama bottleneck? vLLM / llama.cpp / other recommendations?