Triton backend for Falcon3-10B-1.58bit achieves 97.5 tok/s decode on RTX 5070 with packed ternary weights and DP4A decode path
Read the original at old.reddit.com→Hi r/LocalLLaMA — I’m sharing an experimental GPU-only inference backend and looking for independent reproductions, not just stars. Model: tiiuae/Falcon3-10B-Instruct-1.58bit GPU: NVIDIA RTX 5070 Batch: 1 Measured...
Original headline: "I built a Triton backend for Falcon3-10B-1.58bit: 97.5 tok/s decode on an RTX 5070"
Coverage timeline
- Jul 26, 17:04 UTC r/LocalLLaMA lead source I built a Triton backend for Falcon3-10B-1.58bit: 97.5 tok/s decode on an RTX 5070