Beyond single-dimensional compression: exploring compound sparsity to delay performance degradation in large language models
Read the original at arxiv.org→arXiv:2607.18280v1 Announce Type: new Abstract: Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid...
Original headline: "Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models"