Local: asks for minimum useful tokens per second for prompt processing and text generation across models
Read the original at old.reddit.com→Hello guys, hoping you're doing fine. Lately with all the new models, and how popular is offloading, what are your min good or usable t/s for both PP and TG? Speaking on my case, I think PP about 300-350t/s for min,...
Original headline: "For Local, what are your minimum good or usable tokens per second, for both promp processing and text generation?"
Coverage timeline
- Aug 1, 23:31 UTC r/LocalLLaMA lead source For Local, what are your minimum good or usable tokens per second, for both promp processing and text generation?