OpenAI’s Chip Beats Nvidia’s GPUs on Benchmarks; Here’s Everything to Know About the New Jalapeño
Discover OpenAI’s new Jalapeño chip, benchmark results, and how it compares with Nvidia GPUs in AI performance and efficiency.

Just before launching its much-acclaimed Sol model in July, OpenAI announced it had developed its own inference chip. It partnered with American multinational semiconductor company Broadcom.
OpenAI had reported that its chip was developed from design to production in nine months— representing what it calls “the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors,” with the intention of making it widely available but accelerated by OpenAI’s models
Since then, the company has been testing the chip.

How semiconductors and foundries work is brilliantly
explained in this interview.
Incidentally, this week, reports emerged of Nvidia’s intention to purchase popular model-hosting service Hugging Face, sparking debate about whether the acquisition is a response to OpenAI and what looks like a breach of trust between the two companies after their much-publicized 2025 deal that tied OpenAI’s chip demand to Nvidia.
In this article, we want to take a closer look at this event following the recent update from OpenAI on the testing phase of its chip.
We’ll start by looking at OpenAI’s motivation, all the controversy —if you could say— around it, and what this could mean for the burgeoning AI ecosystem.
OpenAI’s Motivation was Speed—Sort of
If you came across the news that announced the development of the new chip online, you may know that several tech industry pundits have speculated about whether the motivation for this collaboration between OpenAI and Broadcom is beyond what was laid out in the initial announcement.
Some say the partnership represents a U-turn on the previously mentioned deal with Nvidia.
On its part, though, OpenAI had stated both speed and the need for tighter control over the complete AI stack as its motivation for developing its own chip.
]
The company has made it known that although its goal is developing frontier models and building products on top of them through its hardware arm, it is also seeing the need to design the infrastructure beneath this goal, such as chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience, to increase the odds for its models and company.
Because OpenAI intends to operate across the stack, each of these layers is expected to be optimized in its new chip around the same goal of making its models faster, more reliable, and more affordable for users.
Jensen’s Asia Interview
Beyond the corporate speak, it is important to point out that in a January 2026 interview outside a Taipei restaurant, Nvidia CEO Jensen Huang had denied reports that he was unhappy with OpenAI following an earlier report from The Wall Street Journal in what was the first sign of a crack in the company’s relationship.
In that interview, Jensen had said his company was ready to make its largest-ever investment in OpenAI.

Reacting to the news of OpenAI’s new chip has been
largely positive.
Although the difference between the timeline of the announcment of the new chip in June and that of Nvidia’s intention to help build 10 gigawatts of compute capacity and invest up to 100 billion in OpenAI in September is off by at least a month from the reported 9 months it took to develop Jalapeño—suggesting that designing the chip had begun before the partnership, we cant resits poking holes in this new announcement and its implications on the relationship of the two companies as well as the silicon valley ecosystem.
What we Know About Jalapeño
Regardless of our sentiments, OpenAI’s new chip —christened Jalapeño—was primarily designed for inference rather than training.
Perhaps this is a distinction that needs to be made, as some have pointed out.
Inference is the phase of AI where the weight —or code— of the model does actual work for end users.
In its newly released update, OpenAI reports that Jalapeño has shown industry-leading speed and efficiency in AI inference.

Transcript of OpenAI’s team announcement of Jalapeño can be
accessed from Hot Chips.
It had also previously reported that its own AI models helped save Jalapeño’s development when the design would not work.
“We introduced a new hardware programming environment, and more than half of the core is written in it: that’s the XLS hardware language and associated compiler infrastructure.
And this also helped us doubly, because it turned out that the AI we wanted to enable had a really good time writing programs in this language — because it kind of looks like Rust,” reported OpenAI’s hardware lead.
He continues:
“So between this advanced hardware language and AI, we were able to get the AI to write things that were correct by construction, quickly and easily explore the space.”
It was like having a sub-team of people optimizing PPA for us in the background…
One of the key stories from our program was that the content was not going to fit in the floor plan block — which is probably familiar to many people in the audience.
But through the use of AI to PPA-optimize, we were able to squeeze the content in, hit our target schedule, and deliver the data sheet that we thought we were going to achieve at the outset.
So AI really came to help us save the day at the end there.”
Exploring the Technicals
Beyond OpenAI’s older models being used to develop its chip, the company has said its latest models have been at the forefront of accelerating how it optimizes and programs it.
To understand how Jalapeño performs, the company tested it on InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request, and compared it to leading inference stack configurations.
According to the company’s report, they tested Jalapeño for different scenarios, from high-throughput serving to highly interactive, low-latency use cases, using all of its open-source GPT 120B parameter model, DeepSeek R1, and Kimi K2.5 1T as base testing models.
The result? Well, Jalapeño delivers both higher throughput and lower latency.
The breakthrough is similar to the use of a unified memory architecture in an integrated circuit, introduced with the M series chips in 2021. However, here, Jalapeño is being specifically built for a hyperspecialized role of inference load as opposed to general-purpose computing tasks.
In ‘integrated circuits’, such as the M series chips, both the CPU, GPU, and sometimes the Neural Processing Unit have been made into one.
The key numbers to note from this announcement are:
- Jalapeño does 1.5x to 1.9x more AI work per watt vs Nvidia chips, per reported tests
- It achieved 1.7x to 3.6x lower end-to-end latency across GPT-OSS, DeepSeek R1 and Kimi K2.5 1T
- Roughly 700W chip power, with reports of 700+ tokens/sec/user on DeepSeek R1
- In one report, there were claims of 104.3x more throughput per kW than Nvidia’s GB300 at a matched 169.41 tok/s/user decoding speed
Open Source is Why You Should Care
Even though it admittedly may be hard to see, as founders, the impact of these events on open source is why you should care.
Besides the fact that a secured value chain means consistent delivery of OpenAI’s services to general consumers and institutions alike, both of these reports have ultimately made the AI scene interesting again to watch.
The two most recent acquisitions of Nvidia using less than 1% of its market cap have positioned the company directly as competition for OpenAI’s specialized software Agent or Harness and also as a potential competitor for its main business.

Chamath shares Anthropic’s belief that there could be only a few
companies left in the world in the end?
As a matter of fact, pundits have speculated that Jensen is speedrunning the open-source AI model movement. Presumably with the intention of spreading out demand for its products away from the niche and tiny core of big companies to the wider populace.
This is going to be an incredible reversal of pace in the timeline of the evolution of AI if it plays out exactly this way.
And Nvidia doing an open-source self-driving project, and now owning Hugging Face, the leading open-source model distributor, makes this prediction plausible.
According to investor Chamath Palihapitiya in a recent podcast, “the generalized takeaway [here] should be that you’re going to see all of these businesses converge and they’re all going to be competing with each other in every which way possible.
So this idea that you have this very delineated separation of customer and supplier, I think, is melting away.”
We agree, and we think when that time comes it would only be all too good for founders as much as it would for general-purpose AI consumers.



