Qwen 35B A3B tokenizes HTML/JS input differently from Gemma 26B A4B, helping explain why Qwen excels at coding while Gemma performs better on language tasks
Read the original at old.reddit.com→Pasted the same HTML/JS code (330 lines) into Qwen 35B A3B and Gemma 26B A4B. Qwen: tokenized the input to 1609 tokens Gemma: tokenized the input to 4258 tokens. Damn. I've never noticed this before and I haven't...
Original headline: "No wonder Qwen and Gemma are so different"
Coverage timeline
- Aug 9, 00:04 UTC r/LocalLLaMA lead source No wonder Qwen and Gemma are so different