How is qwen 27b so good compared with GPT-4o; questions arise about training data and technologies underlying its performance
Read the original at www.reddit.com→Picture from a post in r/amodei . People were praising qwen and I'm just wondering, what kind of new technologies are at play here? Does qwen just have "better" pre training data? That's more high quality? ...
Original headline: "How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?"
Coverage timeline
- Oct 5, 17:20 UTC r/LocalLLaMA lead source How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse?