Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters front-ends Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy with under 2 GB extra memory
Read the original at www.reddit.com→A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO...
Original headline: "Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory"
Coverage timeline
- Oct 1, 13:58 UTC r/LocalLLaMA lead source Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory