LFM2.5-2.6B runs on OnePlus 13 at ~17 tokens per second on pure CPU; Q4_K_M GGUF used with custom inference engine and device probe suite on Android
Read the original at old.reddit.com→As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine built from scratch....
Original headline: "LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU"
Coverage timeline
- Aug 5, 14:20 UTC r/LocalLLaMA lead source LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU