Localizing cross-kernel divergence in INT8-quantized LLM inference across CUTLASS and Triton kernels
Read the original at arxiv.org→arXiv:2608.13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable. We test that assumption: holding the checkpoint, prompts,...
Original headline: "The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference"
Coverage timeline
- Aug 17, 04:00 UTC arXiv cs.LG lead source The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference