A 4B model’s classification accuracy swung 60% to 82% depending on harness design, with identical weights, data, and scorer across runs.
Read the original at old.reddit.com→I ran a pre-registered ablation on a classification task (Kubernetes issue → SIG triage) using a 4B model on a 6GB laptop GPU. Same frozen weights, same 250-issue gold corpus, same scorer across every run. The...
Original headline: "60-82% accuracy swing on 4B model classification task: the only variable was harness design"
Coverage timeline
- Jul 31, 21:47 UTC r/LocalLLaMA lead source 60-82% accuracy swing on 4B model classification task: the only variable was harness design