TrueSeeker AI · Verified claim report Case baaaf90a38 · 2026-08-26

§ Claim under review · Benchmark

"OpenAI's custom AI chip 'Jalapeño' benchmarks show up to 1.9x more work per unit of electricity than Nvidia's GB200 and GB300 systems, with responses up to 3.6x faster."

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

High
§

Summary

This one checks out on the numbers. On 25 August 2026, OpenAI published first benchmark results for Jalapeño, its custom inference chip built with Broadcom, reporting 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower response latency than Nvidia GB200 and GB300 systems, tested on SemiAnalysis's public InferenceX suite across three open models. The post quotes the top of those ranges correctly and credits OpenAI as the source. Three things the post leaves out matter. OpenAI generated the numbers itself, with the benchmark's author present at OpenAI's lab rather than running an arm's-length test, though that author published its own analysis agreeing with the result. The Nvidia systems compared are the older Blackwell generation, and SemiAnalysis itself calls that comparison incomplete and unfair because Nvidia's newer Rubin platform is the real peer and is already shipping, while Jalapeño is still at engineering-sample stage with deployment planned later this year and volume production in 2027. The efficiency figures also normalize by each chip's rated power rather than measured power, which OpenAI discloses and the post does not.

§

The readings

key figures from the evidence
3.6x

lower end-to-end latency vs Nvidia GB200/GB300

§

Why this verdict

The two numbers in the claim reproduce OpenAI's published figures exactly, the comparison targets are correctly named, the "up to" phrasing accurately signals these are range maxima, and the post attributes the source to OpenAI rather than passing the numbers off as independent. The operative proposition survives the contradiction test: the benchmark's own author corroborates rather than disputes the result, and even extends it past Rubin. I considered and rejected "Source exists but framing is misleading," because the distortions here are omissions of context that the claim's own attribution partly discloses rather than a reframing that reverses meaning, and I rejected "Accurate" because the missing comparator-generation caveat, the rated-TDP normalization, and the pre-deployment status are material to how a reader would weigh the result. Confidence is High rather than capped at Medium because the claim is explicitly scoped as vendor-reported, the primary vendor artifact was retrieved, and the benchmark operator published a concurring analysis; as of 2026-08-26 no contrary result has appeared.
§

Evidence

The claim's numbers are real and match the source. OpenAI presented Jalapeño at Hot Chips on 2026-08-25 and published first benchmark results the same day. Jalapeño delivered 1.5 to 1.9 times more throughput per kilowatt and 1.7 to 3.6 times lower end-to-end latency than Nvidia's GB200 and GB300 rack systems on SemiAnalysis's public InferenceX suite, with a 700W part going up against accelerators rated at 1,200W and 1,400W , and the tests covered three open models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's 1-trillion-parameter Kimi K2.5 .

On methodology, OpenAI states directly: "To compare the systems consistently, we normalized the results using each accelerator's published chip power rating. Jalapeño is rated at 700 watts, although its measured sustained power remained at or below 550 watts on the workloads tested." OpenAI also frames the comparison set as leading commercially available AI systems across the tested operating range, from high-throughput serving to highly interactive, low-latency use .

Provenance of the numbers is mixed rather than independent: the results come from tests using SemiAnalysis's public InferenceX benchmark, OpenAI provided the numbers, and SemiAnalysis verified some runs on-site in the lab . SemiAnalysis describes its access: OpenAI invited it to look at the chip, visit the labs, and benchmark it with the InferenceX suite . Its conclusion is favorable: "In general first generation chips are not competitive, but OpenAI bucks the trend by being industry leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open source models."

SemiAnalysis also volunteers the central caveat the post omits: it acknowledged the comparison with Blackwell is "somewhat incomplete and unfair," as Jalapeño is really competing against Nvidia's newer Rubin, which uses HBM4 while Blackwell uses HBM3e, and added that "Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño."

Tom's Hardware notes Jalapeño was not tested against Vera Rubin, the Nvidia platform slated to power the first gigawatt of Nvidia systems OpenAI agreed to deploy . Partially offsetting this, SemiAnalysis reports Jalapeño still squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization Jalapeño has not adopted , and Jalapeño posted its numbers without multi-token prediction or speculative decoding while some comparison systems did rely on those optimizations .

§

Findings

✓ What's accurate 6

  • The figures 1.9x and 3.6x are exactly what OpenAI published. They are the top ends of ranges OpenAI itself states as 1.5x to 1.9x and 1.7x to 3.6x, and the post's "up to" phrasing signals that correctly.
  • The comparison targets named in the post, GB200 and GB300, are the systems OpenAI actually benchmarked against.
  • "More work per unit of electricity" is a fair plain-English rendering of throughput per kilowatt, and "responses faster" is a fair rendering of lower end-to-end latency.
  • The attribution "Source: OpenAI" is correct. The post does not launder vendor numbers as independent findings.
  • The caption's secondary claim about AI-assisted design matches OpenAI's wording: AI played a direct role in Jalapeño's development, enabling the team to move from initial design to tapeout in nine months.
  • The caption's deployment timing matches reporting: deployment in OpenAI data centers is planned for later in 2026.

≈ What's misleading 6

  • Benchmark cherry picking (favorable comparator): the post presents the win against GB200 and GB300 without noting that these are Blackwell-generation, HBM3e parts, while Jalapeño uses HBM4. The benchmark's own author calls the Blackwell comparison "somewhat incomplete and unfair" and identifies Rubin as the proper peer. This is partly mitigated because SemiAnalysis separately reports Jalapeño ahead of Vera Rubin on perf per megawatt, but the post carries neither the caveat nor the mitigation.
  • Marketing as evidence: "benchmarks show" reads as a neutral test result. In fact OpenAI generated the numbers on someone else's suite, with the suite's author present by invitation. That is stronger than a pure vendor claim and weaker than an independent run, and the post's phrasing collapses the distinction.
  • Omitted qualifier: the perf-per-watt figures are normalized by rated package TDP, not measured power, and Jalapeño's measured draw was materially below its rating. The comparison is also chip-level rather than full rack-level power. OpenAI discloses this; the post does not.
  • Omitted qualifier: Jalapeño ran without multi-token prediction or speculative decoding while some comparison systems used them, which cuts in Jalapeño's favor and is worth stating alongside the ratio.
  • Unreleased as released: the headline framing invites reading Jalapeño as a fielded competitor. It is at engineering-sample stage, with data center deployment planned later in 2026 and volume production in 2027, against Nvidia parts shipping in production today.
  • Omitted qualifier (secondary claim): the caption's "nine months from first design work to manufacturing-ready" is OpenAI's own framing, but SemiAnalysis dates the program differently, reporting that design work began in the middle of 2024, going from initial team hiring to manufacturing tape-out in about 16 months. The nine months is a design-to-tapeout window inside a longer program, not the whole effort.

? What's uncertain 5

  • Whether the same ratios hold at rack and datacenter scale under production load. All published figures are pre-deployment lab results on engineering samples.
  • The full per-model breakdown behind the ranges. Published fragments indicate the top-line ratios vary by model, for example on Kimi K2.5 1T, approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency, so 1.9x and 3.6x are not the same model's numbers as each other in every case.
  • Exactly which runs SemiAnalysis independently verified versus which OpenAI supplied. Reporting says "some runs," without a boundary.
  • Whether Nvidia or any third party will contest the figures. No Nvidia response was found as of 2026-08-26.
  • Measured wall power for the Nvidia comparison systems, which would allow a measured-to-measured efficiency comparison rather than a rated-TDP one.
Distortion flags benchmark cherry picking marketing as evidence omitted qualifier unreleased as released
§

Sources

8 of 8 linked to records
[1]

OpenAI, "Jalapeño's first results show industry-leading speed and efficiency in AI inference," openai.com/index/jalapeno-first-results/

primary vendor, interested party for comparative claims
https://openai.com/index/jalapeno-first-results/ ↗
[2]

SemiAnalysis, "OpenAI Jalapeño: Better Than Nvidia Blackwell"

primary independent evaluator with published benchmark suite, but hosted/invited access
https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia ↗
[3]

Tom's Hardware, "OpenAI's 700W Jalapeño ASIC outpaces 1,400W Nvidia flagship GPU"

secondary named-outlet trade journalism
https://www.tomshardware.com/tech-industry/semiconductors/openai-says-its-jalapeno-chip-beats-nvidias-gb300-in-first-published-benchmarks ↗
[4]

TechCrunch, "OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show"

secondary named-outlet journalism, includes press-call quotes
https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/ ↗
[5]

The Register, "OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast"

secondary named-outlet trade journalism
https://www.theregister.com/systems/2026/08/25/openais-upcoming-jalapeno-chip-looks-like-itll-be-an-inference-beast/5292052 ↗
[6]

DataCenterDynamics, "OpenAI details Jalapeño AI chip, with 700W TDP"

secondary trade press
https://www.datacenterdynamics.com/en/news/openai-details-jalape%C3%B1o-ai-chip-with-700w-tdp/ ↗
[7]

OpenAI/Broadcom June 2026 unveiling post

primary vendor
https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ ↗
[8]

The New Stack, on the AI-assisted design timeline

secondary trade press
https://thenewstack.io/openai-jalapeno-inference-chip/ ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →