§ Claim under review · Benchmark
"OpenAI published measured performance results showing its custom AI inference chip Jalapeño outperforms Nvidia's GB300 on key inference metrics, delivering 1.5-1.9x more AI work per watt at peak throughput, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x higher performance for interactive workloads like AI agents, per the InferenceX SemiAnalysis benchmark"
Verdict
Mostly accurate
Confidence
MediumSummary
This one is mostly accurate. OpenAI really did publish these exact numbers for its Jalapeño inference chip on 25 August 2026, and the figures quoted in the post, 1.5 to 1.9 times more work per watt, 1.7 to 3.6 times lower latency, and 2.1 to 4.1 times better on highly interactive workloads, match OpenAI's own page word for word. InferenceX is a real public benchmark run by the research firm SemiAnalysis. The important missing context is that OpenAI ran the tests and supplied the numbers itself. SemiAnalysis watched some runs in OpenAI's lab and broadly backed the conclusion, but said it did not run the full benchmark suite, and it called the comparison against Nvidia's GB300 "somewhat incomplete and unfair" because the fair rival is Nvidia's newer Rubin platform, which was not tested. Two smaller points: the top of the work-per-watt range was measured against an older GB200, not a GB300, and the chip is still at engineering-sample stage with deployment only planned by the end of 2026. Also, the post's graphic credits the claim to Sam Altman, but it was published by OpenAI and presented by its hardware chief Richard Ho.
Why this verdict
Evidence
The claim's numbers are reproduced accurately from OpenAI's own published post. OpenAI states that across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T, "Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems," and that for highly interactive workloads it delivered 2.1 to 4.1 times higher performance . On Kimi K2.5, the largest public model tested, OpenAI reports roughly 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system.
The benchmark named is real. InferenceX, formerly InferenceMAX, is SemiAnalysis's open-source continuous inference benchmark research platform covering GB200 NVL72, GB300 NVL72, B200, MI355X and other accelerators.
The decisive caveat comes from the benchmark operator itself. SemiAnalysis wrote that all numbers were provided to them by OpenAI, that they verified the InferenceX runs in person in the lab but did not run the full suite of InferenceX benchmarks and had not seen AgentX results, AgentX being their preferred suite for chip comparison because its long-context multi-turn characteristics reflect realistic production cache behaviour . SemiAnalysis further stated that the comparison to Blackwell is "somewhat incomplete and unfair" because Jalapeño competes against chips like Rubin that also use HBM4, that Vera Rubin systems are shipping to customers now while OpenAI has nothing beyond engineering samples, and that the models tested are not on the open frontier . Separately, however, SemiAnalysis concluded that even compared with Rubin, Jalapeño's single-token-prediction output throughput per megawatt surpasses the Vera Rubin multi-token-prediction figures Nvidia and CoreWeave published in July , and described the part as "beating every Nvidia, AMD, and Google chip we have been able to test" .
Comparator detail matters. OpenAI's own chart captions show two different Nvidia comparators: GPT-OSS-120B at nominal 8k/1k, STP, against a GB200 at 1,200 W package TDP, and DeepSeek R1 and Kimi K2.5 against a GB300 at 1,400 W package TDP, with Jalapeño at 700 W . Tom's Hardware notes that Jalapeño was not tested against Vera Rubin, does not train models, and that the major comparisons ran Jalapeño's single-token prediction against GB300 configurations doing the same even though Nvidia deployments commonly use multi-token prediction in production .
Power normalisation used rated TDP, not measured draw. OpenAI revealed at Hot Chips that the processor is rated at 700 watts but measured sustained power stayed at or below 550 W on the workloads tested , and OpenAI normalised the benchmark results against each accelerator's published package TDP .
Deployment status: OpenAI's first custom AI chip is expected to begin deployment in the company's computing infrastructure by the end of the year . Nvidia's response was dismissive rather than a technical rebuttal: Jensen Huang brushed off the chip, saying he remains confident in Nvidia's technology and ability to supply the AI industry even as customers develop their own processors .
Findings
✓ What's accurate 7
- All three headline figures match OpenAI's published post verbatim: 1.5-1.9x work per watt at peak throughput, 1.7-3.6x lower end-to-end latency, and 2.1-4.1x on highly interactive workloads.
- OpenAI did publish measured results from working silicon, presented at Hot Chips 2026 by hardware VP Richard Ho, and Jalapeño is a real chip co-developed with Broadcom, unveiled in June 2026.
- InferenceX is a real, public, open-source SemiAnalysis inference benchmark, not an invented artifact.
- GB300 is a genuine comparator in the published tables, and OpenAI's chip chief said GB300 was the leading option on the benchmark used.
- The three models named in the post's caption, GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, are the models actually tested.
- The claimed deployment timing, beginning by end of 2026, matches OpenAI's stated plan.
- SemiAnalysis, the benchmark's operator, independently endorsed the broad conclusion, including that Jalapeño's figures exceed published Vera Rubin numbers on throughput per megawatt.
≈ What's misleading 7
- **Marketing as evidence:** the phrase "per the InferenceX SemiAnalysis benchmark" reads to a general audience as third-party measurement. The benchmark is third-party; the runs and numbers are not. SemiAnalysis states plainly that OpenAI supplied all numbers, that it verified runs in person but did not execute the full suite, and that it had not seen AgentX results. The claim's own opening, "OpenAI published," partly offsets this, which is why this is a framing gap rather than a fabrication.
- **Benchmark cherry picking:** the comparator was selected favourably. GB300 is an HBM3E Blackwell Ultra part; Jalapeño uses HBM4. The benchmark operator itself called the Blackwell comparison incomplete and unfair and named Vera Rubin, which is shipping to customers now, as the like-for-like peer. Vera Rubin was not tested.
- **Scale conflation:** the claim attributes the full 1.5-1.9x range to GB300. OpenAI's own chart captions show the GPT-OSS-120B comparison, which anchors the top of that range, was run against a 1,200 W GB200, a prior-generation part, not a GB300.
- **Harness mismatch:** Jalapeño's single-token-prediction results were set against Nvidia STP configurations, while production Nvidia deployments commonly use multi-token prediction. SemiAnalysis's Rubin comparison likewise sets Jalapeño STP against Rubin MTP figures. Also, the tested scenario is nominal 8k/1k single-turn, not the long-context multi-turn AgentX scenario the operator considers most representative of agentic production load, which is notable given the claim invokes "AI agents."
- **Cost compute omission:** all per-watt ratios are normalised to rated package TDP rather than measured draw. OpenAI disclosed that Jalapeño's sustained power stayed at or below 550 W. The normalisation choice is disclosed by OpenAI but disappears entirely from the social-post version.
- **Omitted qualifier:** Jalapeño is at engineering-sample stage and is inference-only. It cannot train models, the workload where Nvidia is unchallenged, and it is not yet deployed.
- **Misattribution (in the post's image, not the claim text):** the graphic headline reads that "CEO Sam Altman claims it outperforms Nvidia's GB300 in key tests." The results were published by OpenAI and presented by hardware VP Richard Ho. Altman's contribution on X was a short remark that the chip is fast, not the GB300 comparison.
? What's uncertain 5
- No fully independent, end-to-end reproduction of the full InferenceX suite on Jalapeño exists as of 2026-08-28. SemiAnalysis's witnessing is meaningful corroboration but is explicitly not a full independent run.
- AgentX results, the operator's preferred agentic scenario, have not been published for Jalapeño. This bears directly on the claim's "interactive workloads like AI agents" framing.
- Per-model harness details beyond the chart captions, such as batch and concurrency sweeps behind the 2.1-4.1x interactive figure, were not retrievable in full.
- I read the OpenAI results page, the SemiAnalysis analysis and the InferenceX site through search excerpts rather than full page loads. The excerpts contain the decisive verbatim sentences, but I did not review the complete documents.
- Whether shipped Jalapeño hardware sustains these ratios at scale, against Vera Rubin, is untested and unknowable now.
Sources
9 of 10 linked to recordsOpenAI, "Jalapeño's first results show industry-leading speed and efficiency in AI inference"
SemiAnalysis, "OpenAI Jalapeño: Better Than Nvidia Blackwell"
InferenceX by SemiAnalysis, live benchmark site
OpenAI and Broadcom, "OpenAI and Broadcom unveil LLM-optimized inference chip," June 24 2026
ServeTheHome, live coverage of the Hot Chips 2026 Jalapeño session
Tom's Hardware, benchmark write-up and Hot Chips deep dive
CNBC, analyst reaction and SemiAnalysis caveats
Bloomberg (via Yahoo Finance syndication), interview with OpenAI chip chief Richard Ho
The Register, DataCenterDynamics, The Next Web, Neowin, The Decoder
Benzinga / Yahoo Finance, Jensen Huang response