Llama.cpp update adds probabilistic MTP with improvements in prose generation; draft-n-max and draft-p-min align with greedy sampling, per merge of PR 27694
Read the original at www.reddit.com→https://github.com/ggml-org/llama.cpp/pull/27694 Now merged. Update your llama if you haven't done so yet. Optimal draft-n-max / draft-p-min seem to be in line with greedy sampling. Main gain seems to be on prose...
Original headline: "Reminder: try probabilistic MTP if you missed it. Decode +14% on prose"
Coverage timeline
- Oct 11, 09:20 UTC r/LocalLLaMA lead source Reminder: try probabilistic MTP if you missed it. Decode +14% on prose