OpenAI has published the first performance results for Jalapeño, a custom chip it designed to run AI models rather than train them — the step known as "inference," where a trained model actually answers your questions.

According to OpenAI's own blog post, Jalapeño delivers "industry-leading speed and efficiency," with higher throughput and lower latency on modern models while using less power. In plainer terms: more answers per second, delivered faster, for less electricity.

The comparison drawing the most attention is against Nvidia. Per The Next Web, OpenAI says Jalapeño outperformed Nvidia's GB300 on both power and speed. Yellow.com framed the same claim as a win on two key AI metrics. Yahoo Finance reported that OpenAI claims its new chips can outperform Nvidia processors in tests.

A deeper technical account comes from SemiAnalysis, summarized by Techmeme, which describes Jalapeño as an ASIC — a chip built for one specific job — developed with Broadcom in 16 months. SemiAnalysis reports it beat chips from Nvidia, AMD and Google on multiple leading open-source models, and examines Jalapeño's total cost of ownership and throughput per megawatt alongside Nvidia's Rubin.

One caveat worth holding onto: the headline performance claims come from OpenAI itself, and the sources here describe results rather than independent verification.

Still, the direction matters. OpenAI has been one of Nvidia's most important customers, and inference is where the ongoing electricity and hardware bills pile up as more people use AI products. A 16-month build with Broadcom suggests custom silicon can be stood up quickly.

Why it matters: if a major AI company can design its own chips that run models faster and on less power, it weakens Nvidia's grip on the industry's most expensive input — and could eventually shape what AI services cost everyone else.