Weight-aware streaming tensor engine runs Kimi K3 using 29 GB of RAM at 0.50 tok/s
Read the original at old.reddit.com→submitted by /u/galapag0 [link] [comments]
Original headline: "Weight-Aware Streaming Tensor Engine: run Kimi K3 using 29 GB of RAM at 0.50 tok/s"
Coverage timeline
- Aug 1, 08:09 UTC r/LocalLLaMA lead source Weight-Aware Streaming Tensor Engine: run Kimi K3 using 29 GB of RAM at 0.50 tok/s
- Aug 2, 04:26 UTC r/LocalLLaMA I pushed Kimi K3 onto one CPU with 8 GB of RAM
- Aug 3, 00:16 UTC r/LocalLLaMA GitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
- Aug 3, 18:25 UTC r/LocalLLaMA Qwen3.8-Max matches Kimi K3 and DeepSeek V4 Flash