Tech & AI News
Hacker News

The AI Race Just Got Awkward

Chinese AI labs have released KV-cache optimizations that reduce memory usage by up to 437×, enabling long-context models to run with far lower VRAM costs. Western companies, including Anthropic and OpenAI, have adopted these techniques, cutting cache-read pricing by 60–80% in new models like Claude Opus 5.5 and GPT-6.1 Sol. The adoption improves inference margins and competitiveness for Western labs.