Llama-CPP: parallel agents decode well, but one agent’s prefill stalls all others during web searches
Read the original at old.reddit.com→Testing with 3-5 agents. Decode performance is superb, however if one performs a web search and needs to process a few thousand tokens, ALL other agents will grind to a halt: I've tried tuning a little bit, but no...
Original headline: "Llama-CPP Parallel Agents --> fine for decode, but one agent's prefill will grind all other agents to a halt"
Coverage timeline
- Aug 11, 01:19 UTC r/LocalLLaMA lead source Llama-CPP Parallel Agents --> fine for decode, but one agent's prefill will grind all other agents to a halt