Tech & AI News
Hacker News

Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs

The Deltafin fork runs the full 2.8-trillion-parameter Kimi K3 on Apple Silicon without pruning, using 16 experts and a 1-M-token context window. It streams the model from four SSDs, achieving 0.2901 token/s (3.447 s/token) on a MacBook Pro, a 1.9% throughput increase over the previous update. The setup requires a 1.7 TB local copy or a 215 GB streamed cache, and the code is built with Cargo and released under MIT.