OpenAI Jalapeño: 1.9x Inference Efficiency vs Nvidia
Nvidia is not finished. But OpenAI just showed why selling the shovel is no longer enough when your biggest customers can build their own bloody shovel.
OpenAI has spent years handing Nvidia an enormous cheque. Now it has built a 700-watt reason to stop writing quite so many of them.
That does not mean Nvidia is cooked. Anyone declaring that after one set of company-released benchmarks needs to put the markets app down and go for a walk. But OpenAI’s Jalapeño chip is a serious warning: the biggest AI customers are no longer content to rent the most expensive part of their own business.
OpenAI’s claim: up to 1.9x more work per watt
On August 25, OpenAI published its first measured results for Jalapeño, its custom inference chip developed with Broadcom. The company says Jalapeño delivered 1.5x to 1.9x more AI work per watt at peak throughput than comparison systems across three public models: OpenAI’s GPT-OSS 120B, DeepSeek R1 and Moonshot AI’s Kimi K2.5 1T.
It also claimed 1.7x to 3.6x lower end-to-end latency, and 2.1x to 4.1x higher performance for highly interactive workloads.
Those are very spicy numbers, and OpenAI knows it. But the important point is not whether every decimal point survives a rival’s microscope. The important point is what OpenAI chose to optimise.
Jalapeño is not built to train frontier models. It is built for inference: the endless, expensive process of serving responses after a model has already been trained. Every ChatGPT answer, Codex task, API call and future agent taking 20 steps instead of one has to run through inference infrastructure.
That is where AI stops being a research project and starts being a cost of goods sold.
A model company can have the cleverest researchers on earth. It can still have a rotten business if every additional customer makes the infrastructure bill explode. That is the trap OpenAI is trying to escape.
This is not a chip announcement. It is a margin announcement.
Most people see a custom chip and think, “Right, OpenAI wants to beat Nvidia.” Too shallow.
OpenAI wants to turn AI from an expensive spectacle into an economic machine it can control. Faster responses matter. More capacity matters. But better performance per watt matters because energy, cooling, networking, data-centre space and scarce accelerator supply all cost real money.
If you can serve more useful work per unit of power, you can do one of three things: lower prices, hold prices and widen margins, or reinvest the savings into a better product. The best companies usually do all three at different times.
OpenAI says Jalapeño is rated at 700 watts, although measured sustained power stayed at or below 550 watts on the workloads tested. The company plans to deploy the chips in 128-chip racks; a full pod contains 2,048 ASICs. That 128-chip configuration is said to offer 1.7 exaflops of 4-bit compute and 27.5TB of HBM4 memory.
This is not somebody making a clever little accelerator board in a garage. It is industrial infrastructure, built around racks, memory, networking and data-centre deployment.
And that is precisely why it matters.
For years, AI has been sold as a model race: who has the smartest chatbot, the best image generator, the most impressive benchmark. The more durable race is now vertical integration. The winners will own enough of the stack — model, data, software, distribution, compute design and customer relationship — to decide where the economics land.
Apple understood this years ago with devices and silicon. Amazon understood it with cloud infrastructure. Google understood it with TPUs. OpenAI is now making the same move, only with far more demand pressure and far less room for error.
Nvidia still has the keys to the training kingdom
Before the Nvidia bears start ordering yachts, here is the inconvenient bit: Jalapeño does not train new models.
OpenAI remains deeply dependent on Nvidia and other suppliers for the massive training workloads required to create next-generation models. Its hardware chief, Richard Ho, has been explicit that Jalapeño is one component of a broader compute strategy that includes Nvidia, AMD, Cerebras and others.
That makes sense. Training and inference have different bottlenecks, different workloads and different economics. Building a specialised inference chip does not mean you wake up tomorrow and replace the most mature AI hardware ecosystem in the world.
There is another obvious caveat. OpenAI’s performance comparisons are against commercially available Nvidia Blackwell systems. Jalapeño will initially deploy only in very small volumes by the end of 2026, with more significant deployment expected in 2027. Nvidia will not sit around politely using last year’s kit while a customer tries to displace it.
So no, this is not a clean “Jalapeño beats Nvidia” story. That headline gets clicks because it is simple. It is also rubbish.
The real story is that Nvidia’s customers are becoming capable enough, rich enough and motivated enough to take the profitable parts of the workload in-house.
That matters even if Nvidia remains the dominant supplier for years.
The overlooked angle: OpenAI is building bargaining power
The immediate benefit of custom silicon is lower cost and better product performance. The overlooked benefit is negotiating leverage.
When you have only one realistic supplier for a critical input, you do not negotiate. You receive an allocation, smile through gritted teeth and pray your demand forecast is right.
When you can credibly run part of your workload on your own hardware — and can diversify across Nvidia, AMD, Cerebras, custom silicon and multiple cloud partners — every supplier conversation changes.
That does not mean OpenAI will stop buying Nvidia systems. It means OpenAI is less captive.
This is a lesson founders routinely miss. They hear “vertical integration” and immediately decide they need to own factories, warehouses and every line of code. Absolute nonsense. Owning everything is a good way to own a magnificent collection of fixed costs.
The point is to own or control the bottleneck that dictates your margin, your customer experience or your negotiating position.
For OpenAI, inference is clearly one of those bottlenecks. It has enormous usage, expensive agentic workloads ahead, and a consumer and enterprise product suite that lives or dies on speed, reliability and price. Building a chip tailored to its own workloads is rational. Frankly, at that scale, not doing it would be negligent.
The full-stack advantage is real — and dangerous
OpenAI says it used its own AI models to speed Jalapeño from initial design to manufacturing tape-out in nine months. Whether that is the fastest advanced ASIC cycle ever is less useful to you than the broader implication: AI is starting to improve the machinery that produces more AI.
That flywheel is powerful.
Better models can help engineers design and optimise chips. Better chips make model serving cheaper and faster. Better product performance creates more usage. More usage creates revenue and operational data. Revenue funds the next generation of infrastructure.
That is a serious compounding machine if it works.
It is also why small companies should stop pretending they will win by merely calling the same model APIs as everyone else. The raw model is becoming less scarce. The durable value is moving to proprietary workflow, customer trust, distribution, data feedback loops and operational speed.
If your AI strategy is “we added a chatbot,” you have not got a strategy. You have added a feature that your competitor can copy before lunch.
What this means for you
If you are a founder, do not rush out and build a chip. That would be a spectacularly expensive way to misunderstand this story.
Instead, do these four things tomorrow:
1. Measure your AI cost per valuable outcome. Not cost per token because nobody buys tokens. Measure cost per qualified lead, reconciled invoice, resolved support ticket, approved claim or completed task. If you do not know that number, you are playing with a very expensive toy.
2. Find your actual bottleneck. It may be model cost, but it may also be bad data, slow human approval, weak onboarding or customers who do not trust automation. Fix the constraint, not the fashionable thing.
3. Design for provider flexibility. Do not build your whole business so tightly around one model or cloud vendor that a price rise, outage or product shift can punch you in the throat. Abstract where it is sensible. Keep options.
4. Spend savings on a better customer result, not a press release. If cheaper or faster AI lets you deliver a product outcome in minutes instead of days, that is your advantage. Tell customers about the outcome. They do not care what flavour of silicon is under the bonnet.
OpenAI’s Jalapeño result is not the death of Nvidia. It is more interesting than that.
It is proof that the AI economy is moving from a land grab for chips into a harder fight over cost structure, supply control and who gets to keep the margin. The companies that win will not be the ones that talk most about AI.
They will be the ones that make it cheaper, faster and more useful — then quietly make a fortune while everyone else is still applauding the demo.