Hacker News
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
The author created slotstream, a macOS-native tool built with MLX and Swift that runs the 125-billion-parameter Qwen-3.8-Flash-Next model in 4-bit mode on machines with as little as 16 GB RAM, using expert offloading and SSD streaming to keep memory under 48 GB. The system includes an auto-mode that balances memory consumption and throughput, achieving roughly 12 tokens per second, and the developer plans to add a speculative decoding MTP module.