AI News

OpenAI Says Its Jalapeño Chip Beats Nvidia’s GB300 on Efficiency

OpenAI Benchmarks Its Own Chip Against the Nvidia It Still Buys
Jalapeno chip

OpenAI has given Nvidia fresh reason to watch its largest customers, after publishing early results from Jalapeño, its first custom inference chip. Developed with Broadcom, the processor is created for serving a large language models and aims to improve both performance and power efficiency. OpenAI says Jalapeño delivered more AI work per watt and lower end-to-end latency than Nvidia’s GB200 and GB300 systems, though the figures are OpenAI’s own and are not independently verified. OpenAI plans to begin deploying the chip in its own data centers by the end of 2026, with a larger ramp in 2027.

Although it says NVIDIA and other accelerators will continue to play a major role. Sam Altman summed up the development bluntly on X, writing: “We made a chip and it is fast.” The comment comes as OpenAI moves further towards acquiring more of the hardware and software stack behind its AI products. This gives the company greater adaptability over the costing and performance of inference at scale.

Why Is OpenAI Building Its Own AI Inference Chip?

Jalapeno was created specifically around the needs of modern future language model inference rather than a general-purpose accelerator. OpenAI says the chip was built considering compute, memory, networking, and software altogether. It was created with the goal of reducing data movement and communication delays during inference. The results are substantial because inference is becoming an increasingly crucial part of the AI infrastructure equation as use of AI models increases. These models them have to create responses quickly, by processing numerous requests and is where the much of the costs sit.

OpenAI says Jalapeno is designed to improve throughput and latency rather than forcing a mix-up between the two. OpenAI tested Jalapeno using InferenceX, a public benchmark from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across the three models, OpenAI reported 1.5 to 1.9 times more work per watt and 1.7 to 3.6 times lower end-to-end latency than Nvidia’s GB200 and GB300. On highly interactive workloads the advantage widened to 2.1 to 4.1 times. The comparison is noticeable, especially because Nvidia is both OpenAI’s main supplier and, as of last week, a backer of up to $105 billion in financing for its data centers.

OpenAI is not replacing Nvidia, but the development shows that one of the largest AI companies is building its own hardware for at least part of its inference workload. OpenAI says it will keep using Nvidia and other accelerators for both training and inference. That means Jalapeno is better viewed as an extra layer in OpenAI’s infrastructure strategy rather than a sudden replacement.

How Jalapeño Fits OpenAI’s Infrastructure Plan

OpenAI’s custom chip push could give it more autonomy over how its models are served as demand for AI continues to grow. The company says Jalapeno can produce more useful AI work from the same amount of power, potentially allowing it to serve more demand while decreasing the costs associated with delivering responses. The chip was developed quickly. OpenAI says AI models played a crucial role in the development, helping the team move from initial design to tape-out in nine months.

The models were used to explore implementation, shorten design and verification cycles, and optimize parts of the chip. OpenAI also used Codex with GPT-Astra to bring out extra OpenAI models to high performance in a relatively short period.

OpenAI and Broadcom first unveiled Jalapeño in June as the first of a multi-generation accelerator platform. Broadcom is helping with silicon implementation and networking technologies, while OpenAI has created the accelerator around its AI workloads. Broadcom said the partnership aims to support gigawatt-scale data center deployments from 2026. The first Jalapeno deployments are expected within OpenAI’s computing infrastructure by the end of 2026, with Gen 2 already in development and Gen 3 taking shape according to OpenAI. For Nvidia, the development does not immediately threaten its position in OpenAI infrastructure.

Why the Timing Matters for Nvidia

Currently, OpenAI is explicitly continuing to adopt Nvidia accelerators, however Jalapeno indicates that major AI customers increasingly want greater autonomy over specialized inference workloads. This could create more competition for Nvidia as custom silicon becomes a major part of the AI infrastructure market. The bigger story is not that OpenAI has stopped buying Nvidia hardware. It is that the organization is building the capability to choose where custom silicon can deliver better economics and performance.

If future Jalapeno generation continues improving on these outcomes, Nvidia may have to compete not only with other chip makers but also with increasingly capable hardware developers, the organizations buying its chips.

Khwaish Manwani
Khwaish Manwani, an inquisitive soul fond of words and driven by a profound interest in article writing that brings thoughts to life. Apart from her way with the words, she also pursues table tennis as a side passion.
You may also like
More in:AI News