We swap AdamW optimizer states for an FFT to cut VRAM use in half; developers explore non-quantization methods to reduce memory during fine-tuning.
Read the original at www.reddit.com→Hey everyone, Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck. We didn't want to...
Original headline: "We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?"
Coverage timeline
- Oct 5, 19:54 UTC r/LocalLLaMA lead source We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?