GLM 5.2 on a 4-socket Xeon box with 1TB RAM and RTX 3060; ik_llama.cpp runs on CPU with 24 attention layers on GPU, but generation crashes with NaN logits beyond 32–64k context
Read the original at old.reddit.com→Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on the GPU. Works...
Original headline: "GLM 5.2 and ik_llama.ccp"
Coverage timeline
- Jul 26, 05:18 UTC r/LocalLLaMA lead source GLM 5.2 and ik_llama.ccp