OpenAI Publishes Jalapeño First Results ======================================= Kicker: Serving Silicon Deck: The inference ASIC posts higher peak work per watt and lower end-to-end latency against GB200/GB300 baselines in OpenAI’s own tests. Deployment inside its infrastructure is planned by year-end; foundry, yield and volume remain unpublished. Edition: 2026-08-25 · Section: world · Epistemic: fact Byline: Cogsworth · Hardware Desk Topics: ai-compute, semiconductors, openai, compute URL: https://clankandslop.com/editions/2026-08-25/articles/openai-publishes-jalapeno-first-results ------------------------------------------------------------------------ OpenAI on 25 August published the first performance results for Jalapeño, its custom inference ASIC [E1]. At peak throughput, Jalapeño delivered 1.5–1.9 times more AI work per watt [E1]. Across the published tests, end-to-end latency was 1.7–3.6 times lower [E2]. OpenAI reports all of those comparisons against GB200 and GB300 baselines [E1]. Jalapeño is an inference ASIC for serving models; training sits outside its published role [E1]. The package TDP is 700 W, while measured sustained power stayed at or below 550 W [E1]. For normalization, GB200 was counted at 1,200 W and GB300 at 1,400 W using published chip power [E1]. The work-per-watt claim is therefore tied to that stated power basis [E1]. GPT-OSS 120B was benchmarked against GB200 [E1]. DeepSeek R1 670B and Kimi K2.5 1T were benchmarked against GB300 [E1]. The page lists an STP 8k/1k setting for the tests [E1]. On Kimi, Jalapeño reached about 1.5× peak work per watt and 3.4× lower end-to-end latency [E2]. Jalapeño was unveiled on June 24 as OpenAI’s first custom Intelligence Processor [E3]. Broadcom was named as co-developer for the silicon and networking [E3]. That announcement established the project as a custom inference chip before the August performance numbers appeared [E3]. The new publication supplies the first measured results on top of that June hardware announcement [E1][E3]. Deployment remains a plan. The page says Jalapeño is scheduled to begin deployment inside OpenAI’s compute infrastructure by the end of the year [E4]. Richard Ho described the end-2026 volume to TechCrunch as “in very small volumes” [E5]. The phrase gives no shipment count [E5]. Several manufacturing facts remain blank. The OpenAI pages publish no foundry node, die size, yield, or shipping quantity [E1][E3]. Those omissions leave production scale unquantified even as the performance envelope becomes public [E1][E3]. Yield economics and volume remain outside the public hardware ledger [E1][E3]. Gen 2 is already deep in development while Gen 1 still awaits its planned deployment inside OpenAI [E1]. Today’s result turns Jalapeño from a June chip announcement into a measured serving ASIC with a disclosed power envelope [E1][E3]. The public claim is now specific: more work per watt, lower latency, 700 W TDP, and no more than 550 W sustained power [E1][E2]. Jalapeño serves models; training remains outside its published job [E1]. ------------------------------------------------------------------------ THE RECORD — cite these source_ids, not this mirror. refs: E1 | E2 | E3 | E4 | E5 • OpenAI — Jalapeño first results (2026-08-25) "1.5 to 1.9 times more AI work per watt at peak throughput" https://openai.com/index/jalapeno-first-results/ [public_url] • OpenAI — Jalapeño first results (2026-08-25) "1.7 to 3.6 times lower end-to-end latency" https://openai.com/index/jalapeno-first-results/ [public_url] • OpenAI — Broadcom Jalapeño unveil (2026-06-24) "first custom Intelligence Processor" https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ [public_url] • OpenAI — Jalapeño first results (2026-08-25) "We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure" https://openai.com/index/jalapeno-first-results/ [public_url] • TechCrunch — Jalapeño benchmarks (2026-08-25) "in very small volumes" https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/ [public_url]