WinterMix 59 GiB Qwen3.5-122B-A10B build with 20k+ context achieves improved perplexity via new reasoning trace annealing; MLX on Apple Silicon faster than llama.cpp on same hardware, Apache 2.0 weights on HF
Read the original at old.reddit.com→TL;DR: I spent another 8 days following my last post making major improvements to the WinterMix method for MLX models. At 20k+ context this 59 GiB build posts a better perplexity than even UnSloth's Q3_K_XL GGUF*...
Original headline: "[Release] WinterMix — 3 Bit WinterMix of Qwen3.5-122B-A10B in native MLX: a 59 GiB build with best-in-class Long Context coherence"
Coverage timeline
- Aug 10, 01:50 UTC r/LocalLLaMA lead source [Release] WinterMix — 3 Bit WinterMix of Qwen3.5-122B-A10B in native MLX: a 59 GiB build with best-in-class Long Context coherence