CausalGate: causal importance distillation for transformer module pruning
Read the original at arxiv.org→arXiv:2607.22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop...
Original headline: "CausalGate: Causal Importance Distillation for Transformer Module Pruning"
Coverage timeline
- Jul 28, 04:00 UTC arXiv cs.LG lead source CausalGate: Causal Importance Distillation for Transformer Module Pruning