OpenAI Publishes the First Jalapeño Benchmarks: The Comparison Conditions Matter
TL;DR
OpenAI says its Jalapeño system delivered 1.5 to 1.9 times the peak throughput with lower latency in InferenceX tests, but newer GPUs, speculative decoding, and system power were not included.
Two sets of evidence can test Jalapeño’s commercial value over the next three to six months. OpenAI would need to disclose rack-level power consumption and cost per million tokens, while an outside test would need to reproduce the performance gap using the same models, latency targets, and decoding settings. Without both, the first benchmark establishes speed under a particular configuration but does not establish that production inference is cheaper than on GPUs.
On August 25, 2026, OpenAI published the first results for Jalapeño, its custom inference chip. The company used SemiAnalysis’s InferenceX suite with GPT-OSS-120B, DeepSeek R1, and Kimi K2.5. It reported 1.5x–1.9x the peak throughput of the comparison systems and between 1.7 and 3.6 times lower end-to-end latency. In an ultra-low-latency setting, OpenAI reported a speed advantage of 2.1 to 4.1 times. These are vendor-run measurements, not operating statistics from a large customer deployment.
A Jalapeño rack contains 128 accelerators and provides 1.7 exaFLOPS of four-bit compute, 27.5 TB of HBM4, and nearly 2 PB per second of memory bandwidth. Each accelerator has 216 GB of HBM4 and 15.4 TB per second of memory bandwidth. OpenAI developed the ASIC with Broadcom and optimized it for inference on trained models. Training still requires the broader programmability of Nvidia or AMD GPUs. The company expects initial systems to begin appearing later in 2026, with volume production following in 2027.
Older GPUs and decoding choices narrow the claim
The Register notes that OpenAI compared Jalapeño against Nvidia’s GB200 NVL72 and GB300 NVL72 rack systems, introduced in 2024 and 2025. Nvidia Rubin and AMD MI455X, both expected to ramp in 2027, were not included. Those forthcoming GPUs are also designed for a mixture of training and inference, unlike the specialized Jalapeño chip. The comparison therefore says something useful about the selected systems, but it cannot settle how the custom ASIC will perform against its direct contemporaries.
OpenAI also excluded speculative decoding. That choice can make a hardware comparison cleaner because the benchmark is not helped by a smaller draft model predicting the output of a larger one. Production customers, however, pay for the complete platform. Their relevant question is which system delivers the required output quality at a target latency for the lowest total cost after software acceleration is enabled. A benchmark that turns off a widely used serving technique does not yet answer it.
Power is the other missing input. OpenAI has not disclosed consumption for the full system of 128 accelerators. Memory capacity, bandwidth, and FLOPS describe technical ceilings, while a data center pays for electricity, cooling, networking, software utilization, and failed or underused hardware as well. Jalapeño performs inference only. If workloads do not keep a specialized rack highly utilized, its peak performance advantage will not necessarily translate proportionally into a lower cost per token.
The most informative result before volume production in 2027 would combine energy per token, rack utilization, and failure rates under the same model and service-quality target. SemiAnalysis or another laboratory could then rerun the same InferenceX version on Rubin, MI455X, and Jalapeño with speculative decoding and full-system power included. If the reported 1.5x–1.9x throughput gap survives those conditions, OpenAI will have evidence that specialization produces a durable inference-cost advantage rather than an early benchmark lead.
Sources:
Related Articles
Huawei Unveils Ascend 960 SuperPoD: Optical Efficiency Still Needs Chips and Supply
Huawei details optical interconnects and an accelerated 2027 chip schedule. Claimed power savings exceed 550 kW, but testing and supply constraints still separate the architecture from delivered capacity.
a16z Raises $1.1B Machine Age Fund for AI Hardware, Data Centers, and Power
Andreessen Horowitz has raised the $1.1B Machine Age Fund for chips, memory, networking, storage, data centers, robotics, and home AI devices, while its deployment pace and portfolio remain undisclosed.