Sparsity-aware pruning preserves differences between outputs in sparsity-sensitive neurons, not just activations, for large language model pruning
Read the original at arxiv.org→arXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that...
Original headline: "The Sparsity Whisperer"
Coverage timeline
- Aug 10, 04:00 UTC arXiv cs.LG lead source The Sparsity Whisperer