Budget inference: GPU-focused dense models with 24GB VRAM vs RAM-focused MoE models with 64GB RAM for CPU; user seeks benchmarks
Read the original at old.reddit.com→Hi all, I’m building a budget inference machine primarily for personal use (chat/assistant tasks, possibly some RAG). I'm torn between two hardware paths and would love input from anyone who has actually benchmarked...
Original headline: "Budget Inference: A GPU for dense models vs. More RAM for MoE models?"
Coverage timeline
- Jul 29, 22:04 UTC r/LocalLLaMA lead source Budget Inference: A GPU for dense models vs. More RAM for MoE models?