A benchmark shows LLMs playing Civilization V; GLM-5.3 leads Opus-5.5, with Qwen-3.8-27B performing surprisingly well.
Read the original at www.reddit.com→A while ago, I posted here getting OSS-120B and GLM-4.6 playing full games of Civilization V. Since then, models have moved pretty far, and we wanted a better understanding about models' capabilities playing the...
Original headline: "A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well."
Coverage timeline
- Oct 5, 23:16 UTC r/LocalLLaMA lead source A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.