Dispatch
OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users   •   OpenAI Announces $200B Valuation Round   •   EU AI Act Compliance Deadline Extended to 2027   •   Google DeepMind Releases Gemini Ultra 3.0   •   Y Combinator S26 Batch: 60% of Startups Are AI-Native   •   MarTech Consolidation: Salesforce Acquires MadTech Pioneer   •   LLM Token Costs Drop 80% Year-Over-Year   •   Meta Llama 4 Released Under Permissive Commercial Licence   •   Anthropic's Claude Achieves New Benchmarks on Reasoning Tasks   •   Venture Capital Flows to AI Infrastructure Exceed $4B in Q2   •   Adobe GenStudio Reaches 500,000 Enterprise Users
Est. MMXXV — Independent Digital PressWednesday, 2 September 2026Vol. I — No. 195
MarTech • Startups • LLMs • Digital Strategyterekhindigital.comMorning Edition

Terekhin Digital Media

Rigorous Journalism at the Frontier of Digital Commerce & Machine Intelligence

Wednesday, 2 September 2026Issue No. 195
LLMs

OpenAI's Jalapeño Chip Outperforms Nvidia Blackwell on Inference Benchmarks

Co-developed with Broadcom, the custom processor beats Blackwell on tokens-per-user and throughput-per-kilowatt — signalling OpenAI's intent to own the full stack from silicon to model, and to compete with the supplier it has long depended upon.

OpenAI has unveiled Jalapeño, a custom inference processor developed in partnership with Broadcom, that has achieved benchmark performance exceeding Nvidia's Blackwell architecture on SemiAnalysis's InferenceX suite — outperforming on both tokens-per-user and throughput-per-kilowatt metrics. The chip's design philosophy prioritises data locality, minimising the movement of model state between processing units that has been a primary source of latency and power draw in transformer-based inference at scale. Small-volume deployment is planned before the close of 2026, with broader commercial availability expected in 2027. The strategic significance extends beyond the performance metrics. OpenAI has been among the largest customers of Nvidia's inference infrastructure; a custom chip that matches or exceeds that infrastructure on its own workloads fundamentally alters the dependency relationship — mirroring the vertical integration strategies pursued by Google with TPUs and Amazon with Trainium, and suggesting that the major AI labs have concluded that inference economics are too central to their cost structures to remain entirely dependent on third-party silicon.

OpenAIJalapeñoinference chipNvidiaBroadcomAI hardware
← Return to Front Page
Related Articles
© MMXXVI Terekhin Digital Media — All Rights Reserved — An Independent Digital Publication