Staging quantized weights to FP8 instead of FP16 yields a 2x faster FP8 path on Apple Silicon M6; +40% MLX prefill on M5/M6 with int8 on M5
Read the original at www.reddit.com→Regular MLX QMM (quantized matmul) stages operands (dequant) to FP16 before matrix multiplication. But if you modify MLX to stage to FP8 instead, you can use the 2x faster FP8 matrix path on the new Apple Silicon M6...
Original headline: "Staging quantized weights to FP8 instead of fp16: 2× M6 matrix path, +40% MLX prefill (+ int8 on M5)"
Coverage timeline
- Oct 8, 21:31 UTC r/LocalLLaMA lead source Staging quantized weights to FP8 instead of fp16: 2× M6 matrix path, +40% MLX prefill (+ int8 on M5)