OpenAI has unveiled Jalapeño, a custom inference processor developed in partnership with Broadcom, that has achieved benchmark performance exceeding Nvidia's Blackwell architecture on SemiAnalysis's InferenceX suite — outperforming on both tokens-per-user and throughput-per-kilowatt metrics. The chip's design philosophy prioritises data locality, minimising the movement of model state between processing units that has been a primary source of latency and power draw in transformer-based inference at scale. Small-volume deployment is planned before the close of 2026, with broader commercial availability expected in 2027. The strategic significance extends beyond the performance metrics. OpenAI has been among the largest customers of Nvidia's inference infrastructure; a custom chip that matches or exceeds that infrastructure on its own workloads fundamentally alters the dependency relationship — mirroring the vertical integration strategies pursued by Google with TPUs and Amazon with Trainium, and suggesting that the major AI labs have concluded that inference economics are too central to their cost structures to remain entirely dependent on third-party silicon.
LLMs
Claude Fable 5.1 Cuts Cache Pricing by 75% — and the Real Story Is What Anthropic Is Building Toward
Anthropic's Fable 5.1 and Mythos 5.1 carry the same underlying model but two different deployment re…