Would extremely high decode tok/s be useful for inference?
Read the original at old.reddit.com→If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases? Let’s assume that this is for actually useful models...
Original headline: "Would extremely high decode tok/s even be useful?"
Coverage timeline
- Jul 30, 18:51 UTC r/LocalLLaMA lead source Would extremely high decode tok/s even be useful?