Multilingual quantization tax: edge SLMs suffer performance degradation from 4-bit weight quantization across languages, study finds across Gemma 4 and Qwen 3.5 using MMLU ProX Lite and GlobalPIQA
Read the original at arxiv.org→arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation-the...
Original headline: "The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs"
Coverage timeline
- Aug 12, 04:00 UTC arXiv cs.CL lead source The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs