Best current R9700 inference engine?
Read the original at www.reddit.com→There are way too many forks to keep track of, so I've gotten lost. As far as I can tell, Radiance VLLM is best for models that fit in GPUs while some form of llama.cpp is probably best for MOE RAM-spill? My...
Coverage timeline
- Oct 8, 10:16 UTC r/LocalLLaMA lead source Best current R9700 inference engine?