tp=6 can work on vLLM when padding model architecture numbers to be divisible by six
Read the original at www.reddit.com→vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is...
Original headline: "tp=6 can work on vLLM, with padding"
Coverage timeline
- Oct 1, 00:39 UTC r/LocalLLaMA lead source tp=6 can work on vLLM, with padding