Open-source LLMs for a 5060 Ti 16GB GPU: evaluating model options and throughput on Qwen2.5-14B with 32768 context
Read the original at old.reddit.com→Getting the above usage rate from running Qwen2.5-14B with the commands below ./llama-cli -m /home/XXXX/huggfacemodels/Qwen2.5-14B-Instruct-Q4_K_M.gguf -ngl 99 -c 32768 [ Prompt: 667.8 t/s | Generation: 44.0 t/s ] I...
Original headline: "Local LLM open-source model options (5060TI 16GB)"
Coverage timeline
- Aug 13, 08:05 UTC r/LocalLLaMA lead source Local LLM open-source model options (5060TI 16GB)