Looking for load testing tools for inference servers to measure sustained load and storage needs; user asks about HTTP load-testers for LLM servers
Read the original at www.reddit.com→Basically, think of HTTP load-testers, but for inference. I want to throw generations (parallel) at my inference server and see how it behaves under sustained load and meassure token generation and prefill and figure...
Original headline: "Looking for "load testing" tools"
Coverage timeline
- Oct 11, 00:27 UTC r/LocalLLaMA lead source Looking for "load testing" tools