What it does
Jalapeño is OpenAI's first custom Intelligence Processor: an LLM-inference ASIC designed with Broadcom (and Celestica for board/rack industrialization per the unveil). The Stack artifact here is the first-silicon results stack: measured InferenceX operating points, power normalization, and the programming model OpenAI says AI can target.
Primaries:
- Jalapeño’s first results show industry-leading speed and efficiency in AI inference, OpenAI, 25 August 2026.
- OpenAI and Broadcom unveil LLM-optimized inference chip, OpenAI (unveil; engineering samples running ML workloads including GPT-5.3-Codex-Spark at production target frequency and power).
Measured claims (first-results primary, public models):
| Metric | Range vs comparison systems |
|---|---|
| AI work per watt at peak throughput | 1.5-1.9x |
| End-to-end latency | 1.7-3.6x lower |
| Highly interactive performance | 2.1-4.1x higher |
Models named: GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T (all in the first-results primary body and figure captions). Benchmark frame: SemiAnalysis InferenceX, matched user-experience / latency gates, results normalized to published chip power ratings. Figure captions name comparison package TDPs as GB200 1,200 W (GPT-OSS) and GB300 1,400 W (DeepSeek R1 and Kimi). Jalapeño package rating 700 W; OpenAI says measured sustained power stayed at or below 550 W on tested workloads. On Kimi (largest public model tested): ~1.5x peak performance per watt and 3.4x lower end-to-end latency vs the comparison system.
Architecture pitch (both posts): blank-slate LLM inference design; minimize data movement; keep KV cache and model state local; balance compute, memory, and networking across prefill vs decode; large network domain so a request can stay inside one connected system. Unveil: Broadcom silicon implementation and Tomahawk networking; multi-generation platform with Microsoft and other data-center partners beginning in 2026 (Hock Tan quote in unveil).
AI-in-the-loop design: OpenAI says AI helped move from initial design to tapeout in nine months, including arithmetic-circuit optimization. Programming target: local tensors, explicit communication, predictable synchronization so humans and AI can map work. Using Codex with GPT-Astra, three open-weight models outside the original production plan reached high performance in two months. Selected GPT-OSS attention and MoE blocks: AI-generated implementations 1.5-1.8x faster than prior human-expert versions (block-level; full-model speedup UNKNOWN from that claim).
Roadmap: begin deploying inside OpenAI compute by end of year; Gen 2 deep in development; Gen 3 taking shape. OpenAI says it will continue deploying NVIDIA and other partner accelerators for training and inference.
Why now
- First measured first-party OpenAI inference silicon on a public InferenceX frame (beyond unveil-day early-testing language).
- Puts a custom-ASIC path on the same operator radar as merchant GPU racks (GB300 / Vera Rubin NVL72), with explicit work/watt and latency claims.
- Documents a nine-month AI-assisted tapeout loop as a process claim operators may hear again from other labs.
- Clarifies that Jalapeño is an inference bet inside a multi-generation Broadcom collaboration, while training still cites external accelerators.
Maturity
| Layer | Signal |
|---|---|
| Engineering samples / lab ML workloads | Unveil: running at production target frequency and power |
| Public InferenceX first results | 25 Aug 2026 post |
| Internal OpenAI deployment | Planned by end of year |
| External sellable SKU / cloud listing | UNKNOWN |
| Process node / die details | UNKNOWN in primaries |
| Independent third-party board bring-up | UNKNOWN |
| Gen 2 / Gen 3 schedules | Named as in progress; dates UNKNOWN |
Pick vs
Watch Jalapeño closely when:
- your inference TCO is dominated by work/watt and interactive latency on large MoE / long-context agents;
- you need a full-stack lab (model + serving + chip) as a competitive reference for merchant GPU pricing;
- you evaluate whether AI-assisted RTL/kernel loops change ASIC schedule risk assumptions.
Stay on merchant GPUs (NVIDIA and others) when:
- you need purchasable capacity this quarter with public MLPerf Closed Division rows and broad ISV stacks;
- you require multi-cloud portability and CUDA/ROCm software depth Jalapeño does not yet offer externally;
- training remains your bottleneck (OpenAI itself continues partner accelerators for training).
Treat InferenceX Pareto claims as:
- OpenAI-reported, power-normalized against published TDPs. Reproduce on your traces before contract language.
Failure modes
- First-party only. Until external SKUs exist, you cannot buy the claimed watts. Planning on Jalapeño capacity outside OpenAI is premature.
- Comparison opacity. Body text says "leading commercially available AI systems"; figure captions name GB200 (GPT-OSS) and GB300 (DeepSeek R1 / Kimi) with published package TDPs. Always cite the post's operating points and caption SKUs. Avoid flattened "2x faster" slogans.
- TDP vs sustained. 700 W package vs ≤550 W measured matters for facility power. Facility design on nameplate alone will oversize; design on blog sustained alone may undersize for other models.
- Kernel tax. OpenAI still says each new model family needs new kernels. The two-month open-weight bring-up is impressive and still a recurring cost.
- Block-level AI speedups. 1.5-1.8x on selected blocks applies to those blocks only. Full-model guarantee: UNKNOWN.
- Supply and partner path. Unveil ties industrialization to Broadcom/Celestica and gigawatt-scale partners. Schedule risk sits in that chain. Details: UNKNOWN.
Operator read
Log Jalapeño as working first-party inference silicon with OpenAI-published InferenceX numbers, scheduled for internal ramp by year-end. Keep merchant GPU racks as the purchasable baseline. Re-read if OpenAI publishes an external product SKU, MLPerf submission, or third-party cloud listing.
