QTEA: ternary LLMs with sparse residual salient weight and by-column optimization
Read the original at arxiv.org→arXiv:2609.00224v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existing PTQ methods...
Original headline: "QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization"
Coverage timeline
- Sep 2, 04:00 UTC arXiv cs.LG lead source QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization