Reasoning-aware compression benchmarks per-module quantization across five reasoning benchmarks to protect vulnerable circuits in energy-efficient LLM deployment
Read the original at arxiv.org→arXiv:2609.05512v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) impose substantial energy costs during deployment, yet current compression methods apply uniform quantization across all components,...
Original headline: "Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment"
Coverage timeline
- Sep 9, 04:00 UTC arXiv cs.AI lead source Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for Energy-Efficient LLM Deployment