What it does

Jalapeño is OpenAI's first custom Intelligence Processor: an LLM-inference ASIC designed with Broadcom (and Celestica for board/rack industrialization per the unveil). The Stack artifact here is the first-silicon results stack: measured InferenceX operating points, power normalization, and the programming model OpenAI says AI can target.

Primaries:

Measured claims (first-results primary, public models):

MetricRange vs comparison systems
AI work per watt at peak throughput1.5-1.9x
End-to-end latency1.7-3.6x lower
Highly interactive performance2.1-4.1x higher

Models named: GPT-OSS 120B, DeepSeek R1 670B, Kimi K2.5 1T (all in the first-results primary body and figure captions). Benchmark frame: SemiAnalysis InferenceX, matched user-experience / latency gates, results normalized to published chip power ratings. Figure captions name comparison package TDPs as GB200 1,200 W (GPT-OSS) and GB300 1,400 W (DeepSeek R1 and Kimi). Jalapeño package rating 700 W; OpenAI says measured sustained power stayed at or below 550 W on tested workloads. On Kimi (largest public model tested): ~1.5x peak performance per watt and 3.4x lower end-to-end latency vs the comparison system.

Architecture pitch (both posts): blank-slate LLM inference design; minimize data movement; keep KV cache and model state local; balance compute, memory, and networking across prefill vs decode; large network domain so a request can stay inside one connected system. Unveil: Broadcom silicon implementation and Tomahawk networking; multi-generation platform with Microsoft and other data-center partners beginning in 2026 (Hock Tan quote in unveil).

AI-in-the-loop design: OpenAI says AI helped move from initial design to tapeout in nine months, including arithmetic-circuit optimization. Programming target: local tensors, explicit communication, predictable synchronization so humans and AI can map work. Using Codex with GPT-Astra, three open-weight models outside the original production plan reached high performance in two months. Selected GPT-OSS attention and MoE blocks: AI-generated implementations 1.5-1.8x faster than prior human-expert versions (block-level; full-model speedup UNKNOWN from that claim).

Roadmap: begin deploying inside OpenAI compute by end of year; Gen 2 deep in development; Gen 3 taking shape. OpenAI says it will continue deploying NVIDIA and other partner accelerators for training and inference.

Why now

  • First measured first-party OpenAI inference silicon on a public InferenceX frame (beyond unveil-day early-testing language).
  • Puts a custom-ASIC path on the same operator radar as merchant GPU racks (GB300 / Vera Rubin NVL72), with explicit work/watt and latency claims.
  • Documents a nine-month AI-assisted tapeout loop as a process claim operators may hear again from other labs.
  • Clarifies that Jalapeño is an inference bet inside a multi-generation Broadcom collaboration, while training still cites external accelerators.

Maturity

LayerSignal
Engineering samples / lab ML workloadsUnveil: running at production target frequency and power
Public InferenceX first results25 Aug 2026 post
Internal OpenAI deploymentPlanned by end of year
External sellable SKU / cloud listingUNKNOWN
Process node / die detailsUNKNOWN in primaries
Independent third-party board bring-upUNKNOWN
Gen 2 / Gen 3 schedulesNamed as in progress; dates UNKNOWN

Pick vs

Watch Jalapeño closely when:

  • your inference TCO is dominated by work/watt and interactive latency on large MoE / long-context agents;
  • you need a full-stack lab (model + serving + chip) as a competitive reference for merchant GPU pricing;
  • you evaluate whether AI-assisted RTL/kernel loops change ASIC schedule risk assumptions.

Stay on merchant GPUs (NVIDIA and others) when:

  • you need purchasable capacity this quarter with public MLPerf Closed Division rows and broad ISV stacks;
  • you require multi-cloud portability and CUDA/ROCm software depth Jalapeño does not yet offer externally;
  • training remains your bottleneck (OpenAI itself continues partner accelerators for training).

Treat InferenceX Pareto claims as:

  • OpenAI-reported, power-normalized against published TDPs. Reproduce on your traces before contract language.

Failure modes

  1. First-party only. Until external SKUs exist, you cannot buy the claimed watts. Planning on Jalapeño capacity outside OpenAI is premature.
  2. Comparison opacity. Body text says "leading commercially available AI systems"; figure captions name GB200 (GPT-OSS) and GB300 (DeepSeek R1 / Kimi) with published package TDPs. Always cite the post's operating points and caption SKUs. Avoid flattened "2x faster" slogans.
  3. TDP vs sustained. 700 W package vs ≤550 W measured matters for facility power. Facility design on nameplate alone will oversize; design on blog sustained alone may undersize for other models.
  4. Kernel tax. OpenAI still says each new model family needs new kernels. The two-month open-weight bring-up is impressive and still a recurring cost.
  5. Block-level AI speedups. 1.5-1.8x on selected blocks applies to those blocks only. Full-model guarantee: UNKNOWN.
  6. Supply and partner path. Unveil ties industrialization to Broadcom/Celestica and gigawatt-scale partners. Schedule risk sits in that chain. Details: UNKNOWN.

Operator read

Log Jalapeño as working first-party inference silicon with OpenAI-published InferenceX numbers, scheduled for internal ramp by year-end. Keep merchant GPU racks as the purchasable baseline. Re-read if OpenAI publishes an external product SKU, MLPerf submission, or third-party cloud listing.