Format-aware fusion for fast FP4 pretraining improves efficiency by co-designing quantization producers with their scale domains and consumer layouts for native MXFP, global NVFP, and cooperative-thread-array-local NVFP
Read the original at arxiv.org→arXiv:2610.00053v1 Announce Type: new Abstract: Four-bit floating-point (FP4) Tensor Cores accelerate matrix multiplication, but scale computation, operand packing, layout construction, and saved backward state can...
Original headline: "Format-Aware Fusion for Fast FP4 Pretraining"
Coverage timeline
- Oct 2, 04:00 UTC arXiv cs.LG lead source Format-Aware Fusion for Fast FP4 Pretraining