Hacker News
Nine coding harnesses vs. your laptop
Swapping the harness’s API calls for a local Qwen 3.8 27B model on an M4 MacBook Pro yields ~22 seconds pre-fill for a 2,008-token system prompt and ~226 seconds for an 18,046-token prompt. After the prompt, only 44% of a 32 k token window remains for Opencode, versus 94% with the optimized “chad” harness. Frequent side requests overload the GPU, causing utilization to exceed 100% in some instances.