LoRA over GGUF enables training DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM; 19 s/iteration on Strix Halo with vibe-coded Triton kernels for sliding attention, CSA, and HCA
Read the original at old.reddit.com→https://github.com/woct0rdho/transformers5-qwen3.5-recipe An update on my progress with low-VRAM LoRA training over GGUF base model: Now we can train DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, with no CPU...
Original headline: "LoRA over GGUF: Train DeepSeek-V4-Flash in 90G VRAM"
Coverage timeline
- Jul 28, 16:49 UTC r/LocalLLaMA lead source LoRA over GGUF: Train DeepSeek-V4-Flash in 90G VRAM