Quantizing DeepSeek V4 0731 and benchmarking against popular quants on 8× RTX 5090; issues found with --no-lazy and FP8 downconversion noted
Read the original at old.reddit.com→We converted the model from the original safetensors and found two issues. The first one made our quantization fail several times, the second one does not fail at all, it just quietly ruins the base 1) You must use...
Original headline: "We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090"
Coverage timeline
- Aug 11, 21:34 UTC r/LocalLLaMA lead source We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090