Show HN: I run a 30B model at 22 tokens/s and 109 tokens/s, not novel; 6GB/16GB RAM usage with llama.cpp
Read the original at github.com→Original headline: "Show HN: I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp"
Coverage timeline
- Jul 29, 03:47 UTC Hacker News (AI) lead source Show HN: I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp