Tech & AI News
TechCrunch

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

OpenAI’s Jalapeño chip, co-designed with Broadcom, achieved higher token-per-user counts and greater throughput per kilowatt than Nvidia’s Blackwell system on the SemiAnalysis InferenceX benchmark, delivering faster, lower-latency responses. The architecture minimizes prefill and communication delays by keeping KV-cache data local and dynamically orchestrating compute, memory, and networking resources.