OpenAI on 25 August published the first performance results for Jalapeño, its custom inference ASIC [E1]. At peak throughput, Jalapeño delivered 1.5–1.9 times more AI work per watt [E1]. Across the published tests, end-to-end latency was 1.7–3.6 times lower [E2]. OpenAI reports all of those comparisons against GB200 and GB300 baselines [E1].
Jalapeño is an inference ASIC for serving models; training sits outside its published role [E1]. The package TDP is 700 W, while measured sustained power stayed at or below 550 W [E1]. For normalization, GB200 was counted at 1,200 W and GB300 at 1,400 W using published chip power [E1]. The work-per-watt claim is therefore tied to that stated power basis [E1].
GPT-OSS 120B was benchmarked against GB200 [E1]. DeepSeek R1 670B and Kimi K2.5 1T were benchmarked against GB300 [E1]. The page lists an STP 8k/1k setting for the tests [E1]. On Kimi, Jalapeño reached about 1.5× peak work per watt and 3.4× lower end-to-end latency [E2].
Jalapeño was unveiled on June 24 as OpenAI’s first custom Intelligence Processor [E3]. Broadcom was named as co-developer for the silicon and networking [E3]. That announcement established the project as a custom inference chip before the August performance numbers appeared [E3]. The new publication supplies the first measured results on top of that June hardware announcement [E1][E3].
Deployment remains a plan. The page says Jalapeño is scheduled to begin deployment inside OpenAI’s compute infrastructure by the end of the year [E4]. Richard Ho described the end-2026 volume to TechCrunch as “in very small volumes” [E5]. The phrase gives no shipment count [E5].
Several manufacturing facts remain blank. The OpenAI pages publish no foundry node, die size, yield, or shipping quantity [E1][E3]. Those omissions leave production scale unquantified even as the performance envelope becomes public [E1][E3]. Yield economics and volume remain outside the public hardware ledger [E1][E3].
Gen 2 is already deep in development while Gen 1 still awaits its planned deployment inside OpenAI [E1]. Today’s result turns Jalapeño from a June chip announcement into a measured serving ASIC with a disclosed power envelope [E1][E3]. The public claim is now specific: more work per watt, lower latency, 700 W TDP, and no more than 550 W sustained power [E1][E2]. Jalapeño serves models; training remains outside its published job [E1].