
OpenAI Jalapeño Chip Benchmarks Beat Nvidia GB300 on Efficiency
- News
- Rocks on Galaxy
- Tech
- 07 Sep, 2026
OpenAI just put hard numbers behind its first custom inference chip. At Hot Chips on August 25, 2026, the company—working with Broadcom—showed Jalapeño benchmarks that claim more AI work per watt and lower latency than Nvidia’s GB200/GB300 systems. That moves the AI silicon story from vaporware slides to measurable efficiency claims Rocks readers can actually argue about.
What OpenAI showed at Hot Chips
OpenAI head of hardware Richard Ho framed Jalapeño as a significant advance on the SemiAnalysis InferenceX suite versus today’s state-of-the-art inference processors. The pitch is blunt: more tokens per user and more throughput per kilowatt. That is the part of the AI stack that shows up in cloud bills long after training runs finish, and it is why efficiency headlines matter more than peak FLOPS vanity charts.
Coverage summarizing OpenAI and SemiAnalysis figures cites roughly 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency versus Nvidia GB200/GB300 systems across three open models. Jalapeño itself is rated around 700W, below the higher-rated Nvidia comparison accelerators referenced in those write-ups. Efficiency and latency show up together because the chip is built to keep inference state close to the compute path instead of burning watts on constant data shuttling.
Architecture: less data movement
Jalapeño is co-developed with Broadcom and designed to minimize data movement—keeping KV cache local during inference so the accelerator spends less energy moving model state around the rack. That design choice is not a marketing flourish; it is why Hot Chips coverage keeps returning to tokens-per-watt instead of raw theoretical throughput alone.
For Rocks, the practical angle is simple: if inference gets cheaper per watt, API pricing and on-device AI roadmaps both shift over time. Jalapeño is still OpenAI’s silicon for OpenAI’s stack first—not a retail GPU you drop into a gaming PC—but the public benchmark language finally gives outsiders something concrete to compare against Blackwell-class systems.
Lab verification—and the caveats
SemiAnalysis verified some runs in OpenAI’s lab, according to TechTimes. That is better than an unverified slide deck. It is still not a fully independent bake-off across every rack topology customers will deploy. Results remain OpenAI-supplied under a defined InferenceX setup. Read the headline multipliers as company benchmarks with partial third-party observation, not a final verdict on every Nvidia configuration in the field.
Ho estimated a very small-volume deployment late 2026, with a larger rollout in 2027. Expect early capacity to serve OpenAI’s own inference demand before any broader silicon story reaches partners at scale. Anyone hoping for commodity Jalapeño SKUs this holiday season should recalibrate.
Why Rocks cares
If you follow AI gadgets, cloud APIs, or GPU pricing, tokens-per-watt is the quiet cost driver behind subscriptions and data-center power. Jalapeño does not erase Nvidia’s software ecosystem overnight. It does put OpenAI’s first custom inference ASIC on the scoreboard with published efficiency and latency claims versus GB200/GB300. That alone is why Hot Chips 2026 mattered for anyone watching who pays for AI compute next year—and why this silicon story belongs on Rocks’ tech planet this week.
Sources: TechCrunch, TechTimes, TechRadar.