AQLoRA: Adaptive-Quantization LoRA speeds up quantized fine-tuning by a single-pass, no-search method that ranks layers by NF4 reconstruction error and keeps top-K in fp16 under memory budget.
Read the original at arxiv.org→arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA...
Original headline: "AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning"
Coverage timeline
- Aug 26, 04:00 UTC arXiv cs.LG lead source AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning