A 5 KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
Read the original at www.reddit.com→Hi everyone, Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM. Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for...
Original headline: "[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)"
Coverage timeline
- Oct 4, 03:48 UTC r/LocalLLaMA lead source [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)