Tech & AI News
Hacker News

Small Models Have Arrived

The author reports that the gpt-5.6-luna model delivers roughly 100 tokens per second with strong coding and knowledge-base performance, while keeping average API costs around $0.10 per request, a stark contrast to earlier models that cost about $1 per query. This reduction in token cost makes “fast/cheap/good-enough” AI viable for consumer apps and routine business tasks, opening a market for high-volume, low-cost inference alongside the continued need for frontier-level models.