OpenAI has partnered with Broadcom to develop "Jalapeño," an ASIC dedicated to LLM inference. This was achieved in a short timeframe of approximately 16 months from the start of design to completion of manufacturing.
This chip is not specialized for a specific model but is intended for general-purpose inference processing. SemiAnalysis conducted benchmarks in a lab under the supervision of OpenAI engineers.
In terms of token throughput per watt, Jalapeño outperformed chips from NVIDIA and AMD. These results were achieved using only single-token prediction, without utilizing multi-token prediction.
With the DeepSeek R1 model, it recorded over 700 tokens per second per user at a concurrency of 1. For GPT-OSS, it reached approximately 1,400 tokens per second.
While the figures were provided by OpenAI, SemiAnalysis verified them on actual hardware. Equivalent results to NVIDIA chips were also confirmed in GSM8k evaluations.
The chip employs HBM4, featuring specifications comparable to flagship GPUs. A demo was also presented showing Doom ported and running on the hardware.
Sources: OpenAI Jalapeño: Better than Nvidia Blackwell (HN 584pt, 381 comments) (HN Search (backfill), 2026-08-25)