LayerRoute: adaptive layer-skipping with LoRA-preserved quality for efficient LLM inference
Read the original at arxiv.org→arXiv:2609.13682v1 Announce Type: new Abstract: We introduce LayerRoute, a parameter-efficient method for adaptive transformer layer-skipping that combines per-layer hard-gated routing (trained via a...
Original headline: "LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference"
Coverage timeline
- Sep 15, 04:00 UTC arXiv cs.CL lead source LayerRoute: Adaptive Layer-Skipping with LoRA-Preserved Quality for Efficient LLM Inference