Ling 3.0 Flash on Strix Halo; vLLM ROCm/HiP, int4 tensors, Qwen-122b on rocmFP4 compared for speed (not a fair comparison)
Read the original at old.reddit.com→vLLM ROCm/HiP, 4 bit compressed-tensors (int4) Not a fair comparison, but Qwen-122b on the most optimized format possible I have run (rocmFP4) does not touch Ling in speed....
Original headline: "Ling 3.0 Flash on Strix Halo"
Coverage timeline
- Aug 10, 22:15 UTC r/LocalLLaMA lead source Ling 3.0 Flash on Strix Halo