Auto-fit beats tuned MoE offload for Qwen3.6-35B-A3B on RTX 3090, increasing prompt throughput from 564 to 1,330 tok/s with unchanged decode speed
Read the original at old.reddit.com→TL;DR: On a Qwen3.6-35B-A3B Q6 setup sized for 64K context on a 24GB RTX 3090, spilling eight MoE expert layers to CPU freed enough VRAM to increase -b from 512 to 1024 and -ub from 128 to 512. Prompt processing...
Original headline: "Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)"
Coverage timeline
- Aug 6, 11:52 UTC r/LocalLLaMA lead source Auto-fit vs tuned MoE offload: 564 → 1330 pp tok/s, unchanged decode (Qwen3.6-35B-A3B Q6 / RTX 3090)