Choosing a model for Hermes on a 4x 3090, 96 GB VRAM setup; llama.cpp or vllm, for a personal AI playground with a few family users
Read the original at old.reddit.com→3 year lurker, now i finally got my server up and running. dont know which model to choose. llama.cpp or vllm, what makes more sense? mainly single user with maybe 2-3 more additional users in family, if everything...
Original headline: "4x 3090, 96gb vram what Model to drive Hermes?"
Coverage timeline
- Jul 25, 22:20 UTC r/LocalLLaMA lead source 4x 3090, 96gb vram what Model to drive Hermes?