Deterministic LLM inference across GPU kernels; a tolerance-based conformance suite detects only minimal output differences in an INT8 Qwen3-1.7B pipeline with injected faults
Read the original at arxiv.org→arXiv:2609.00363v1 Announce Type: new Abstract: Conformance suites for quantized GEMM kernels ask whether two implementations agree within a tolerance. We measure what such a suite can detect. Injecting nine faults...
Original headline: "Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.LG lead source Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance