GLM 5.2 speeds; user reports 3-bit version on 4x 8880 v4 CPUs, 1 TB 32-channel DDR3 RAM, and 2x RTX 3060 12GB achieving 44k token ingestion and 8k token generation on 20k context at $900/€ per system
Read the original at old.reddit.com→Tell me, to get on 20k context and ingestion 44tks, generation 8tks is good numbers for 4x 8880 v4 cpus, 1tb 32channels ddr3 ram and 2x 3060 12gb. ? im running 3bit version whole system cost 900$/€ who can beat me...
Original headline: "GLM 5.2 speeds"
Coverage timeline
- Jul 26, 17:54 UTC r/LocalLLaMA lead source GLM 5.2 speeds