glm-5.2 outperforms ling-3.0-flash in decision-making tasks, with ling-3.0-flash often producing confident but incorrect actions in a shared executor setup
Read the original at old.reddit.com→Same harness, same task set, same agent scaffold, the only thing I swapped was the executor. Not a proper benchmark, no clean tok/s numbers, this is a workflow read not a leaderboard. glm-5.2 is the better model and...
Original headline: "I've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict"
Coverage timeline
- Aug 1, 17:33 UTC r/LocalLLaMA lead source I've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict