Tech & AI News
Hacker News

My local model setup on an M4 Pro Mac Mini

The author runs a local LLM server on an M4 Pro Mac mini with 48 GB RAM, using Qwen3.6-35B-A3B-OptiQ (4-bit) for deep reasoning and Gemma-4-E4B-it-OptiQ (4-bit) for lightweight tasks, served via oMLX and accessed through Tailscale from a MacBook, iPhone, and Telegram. This setup eliminates cloud-API costs, rate limits, latency, and privacy risks while providing offline, predictable-cost inference for daily workflows such as Hermes agent backend, Apollo chat, Pi coding assistant, and Raycast AI.