§ Claim under review · Benchmark
"NVIDIA's AVO agent system achieved a perfect 100.00 RHAE score on the public ARC-AGI-3 benchmark using Claude Opus 5, solving all 183 levels across 25 environments without any instruction, according to NVIDIA." Secondary claim on the post's image slide, investigated below: "NVIDIA's AI system scored a perfect 100 on ARC-AGI-3, solving all 183 problems without any instruction."
Verdict
Mostly accurate
Confidence
MediumSummary
NVIDIA really did publish this. On August 21, 2026 its developer blog reported that its AVO agent system, running Anthropic's Claude Opus 5, scored 100.00 on the RHAE metric across all 183 levels of the 25 public ARC-AGI-3 environments. The Instagram caption's numbers are correct and it does say "public" and "according to NVIDIA," which is better than most coverage. The important missing context is that ARC Prize, which runs the benchmark, did not administer or verify this run, and that ARC Prize has said it will never report public-set scores on its official leaderboard because a harness built with knowledge of those environments can reach 100%. To prove that point ARC Prize itself released an open-source harness that scores 100% on the same public set by replaying human moves. NVIDIA states plainly that these are not results on the semi-private or private competition sets and that the jump from Claude Opus 5's separately measured 30.2% is not a controlled comparison. What remains unknown is whether AVO would perform anywhere near this on the held-out sets, and what the run cost in compute and time.
The readings
key figures from the evidenceNVIDIA's self-reported AVO score on ARC-AGI-3 public set
Claude Opus 5 alone, ARC Prize's own administered public-set run
Why this verdict
Evidence
The underlying result is real and the numbers in the claim match NVIDIA's own writeup. NVIDIA states that using Claude Opus 5, AVO completed the full 25-environment public set with a 100.00 RHAE score, solving all 183 levels in 6,624 environment actions, and that VISTA reports 7,542 environment actions with Claude Opus 5 while completing the same 183 public-set levels, approximately 12% more.
NVIDIA also states that this cross-system comparison should not be interpreted as a controlled ablation, because the two systems differ in agent backend, observation representation, memory, context management, and other implementation details.
NVIDIA scopes the result itself. Its blog says the results cover the 25-environment ARC-AGI-3 public set using the official scorecard and RHAE metric, and are not results on the semi-private or fully private competition sets.
The benchmark operator treats public-set scores as non-probative. The ARC-AGI-3 technical report defines task-specific overfitting to include any agent created with knowledge of public ARC-AGI-3 environments and then evaluated on those same environments, whether trained on them or using a harness handcrafted or specifically configured by someone with knowledge of the public environments; it states such agents can in principle achieve a 100% score on the public set, and that to demonstrate this ARC Prize released an open-source harness scoring 100% on all public environments using human replay; because it is impossible to ensure designers do not use the public environments in their work, and because the public set is materially easier than the private set, ARC Prize says it will never report public set scores of any system on the official leaderboard.
The 30% comparator is real and is also a public-set number, but produced under ARC Prize's own administration. ARC Prize reports that as of July 24, 2026, Claude Opus 5 (High) is the highest-performing model on ARC-AGI-3 at 30.2%, having completed five additional Public Demo environments that no model had previously beaten. Trade coverage draws the distinction explicitly: the AVO figure is not a score on the ARC Prize leaderboard, since verified numbers there including Claude Opus 5's 30.2% come from runs the foundation independently administers, whereas AVO's result was generated and reported by NVIDIA's own team using its own reimplementation of the task interface.
No independent verification of the AVO number was found. ARC Prize states that self-reported community scores are untrustworthy by design, that it does not run community leaderboard code or verify scores against the semi-private sets, and that for ARC-AGI-3 it collects a required scorecard_url and derives the score from it. Reaction was split: on X, many congratulated NVIDIA while others dismissed the result as overfitting on public data.
Findings
✓ What's accurate 5
- NVIDIA published this result on its own developer blog and its official AI account on August 21, 2026. The announcement is real, not fabricated.
- The specific figures check out against NVIDIA's writeup: 100.00 RHAE, 183 levels, 25 environments, Claude Opus 5 as the underlying model.
- The claim correctly says "public" and correctly attributes the result to NVIDIA. Both qualifiers matter and both are present.
- The agent genuinely received no stated rules or goals within the episode, per NVIDIA's description of the setup.
- The caption's "Claude Opus 5 alone was previously reported at roughly 30% under a different setup" is accurate, and "under a different setup" preserves the caveat NVIDIA itself makes.
≈ What's misleading 5
- Omitted qualifier: the claim reports a self-reported, unverified number without saying so. ARC Prize did not run or verify this result, and its own policy is that self-reported scores are untrustworthy by design. A reader hears "achieved" where the evidence supports "NVIDIA reports it achieved."
- Benchmark cherry picking: the claim, and the coverage generally, presents a public-set result as a result on ARC-AGI-3. The public set is the one split the benchmark operator says it will never report on the official leaderboard, precisely because a harness built with knowledge of those environments can reach 100%. ARC Prize proved the point by publishing a human-replay harness that scores 100% on the same set. Without that context, a ceiling score on the practice set reads as a solved benchmark.
- Harness mismatch: the implied 30% to 100% jump compares ARC Prize's administered run of Opus 5 through the official model interface against NVIDIA's agent running through NVIDIA's own reimplemented interface. NVIDIA states this is not a controlled ablation. The Instagram caption partially discloses this with "under a different setup," but the framing still invites the reader to attribute the entire gap to the AVO architecture.
- Omitted qualifier, on the image slide specifically: the headline reads "scored a perfect 100 on ARC-AGI-3" with no mention of the public set, and says "183 problems" rather than 183 levels. On its own, that slide would earn a "source exists but framing is misleading" verdict, since it removes the single qualifier that determines what the number means.
- Cost compute omission: neither the claim nor most coverage carries the resource picture. Environment action counts are published, but the token spend, dollar cost per environment, and wall-clock time for a long-horizon agent loop with persistent memory and supervision were not published in a form I could locate. ARC Prize's own leaderboards are built around cost per task for exactly this reason.
? What's uncertain 5
- Whether NVIDIA submitted a scorecard URL to the ARC-AGI Community Leaderboard, and whether ARC Prize has responded to or commented on the AVO result. No ARC Prize statement on it was found as of 2026-08-22.
- Whether AVO's task interface or configuration was informed by knowledge of the public environments. NVIDIA says it reimplemented the task interface independently and adopted direct-interaction design principles from VISTA, but the writeup does not settle whether any tuning used public-environment knowledge, which is the exact condition ARC Prize defines as task-specific overfitting.
- Whether the result would hold on the semi-private or private sets. Unknown, and NVIDIA does not claim it would.
- Whether AVO's code for this run is public and reproducible. NVIDIA describes the architecture, but I could not confirm a release that would let a third party rerun it.
- Full harness settings: reasoning effort, retry policy, number of attempts per environment, and cost.
Sources
9 of 9 linked to recordsNVIDIA Technical Blog, "NVIDIA AVO Reaches 100% on ARC-AGI-3..." (vendor artifact of record for this result; retrieved via search index with verbatim passages)
ARC Prize, ARC-AGI-3 Technical Report (benchmark operator's rules on public-set scoring and task-specific overfitting)
ARC Prize results page, Claude Opus 5, 30.2% on ARC-AGI-3 as of July 24, 2026
ARC Prize Community Leaderboard repository and policy pages (self-reported scores, not verified, scorecard_url required for ARC-AGI-3)
NVIDIA AI official account post, August 21, 2026
OfficeChai, "NVIDIA's Coding Agent AVO Scores 100% On ARC-AGI Benchmark"
The New Stack, "Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100%."
Crypto Briefing report
Hacker News discussion thread