Hemingway-1 achieves up to 32.0 tok/s for local inference on M5 Max, according to llm-bench.io
Read the original at www.reddit.com→I ran Altworld's Hemmingway-1 on my M5 Max this week. It's a 27B fine-tune of Qwen3.8-27B built specifically for "human-like" writing, which might be useful for everyday messages, emails, notes, your social accounts,...
Original headline: "Hemmingway-1-oQ8e-mtp: up to 32.0 tok/s for local inference — llm-bench.io"
Coverage timeline
- Oct 1, 19:28 UTC r/LocalLLaMA lead source Hemmingway-1-oQ8e-mtp: up to 32.0 tok/s for local inference — llm-bench.io