TrueSeeker AI · Verified claim report Case 4701873284 · 2026-08-19

§ Claim under review · Benchmark

"BDH-CQ, a 150M-parameter model that reasons silently in latent space instead of writing chain-of-thought text, scored 29.5% on ARC-AGI-1 at $0.0007 per task, and runs 11x cheaper than GPT-5.6 Luna (Low) even after OpenAI's 80% price cut, for about a 5-point accuracy tradeoff (29.5% vs 34.2%)"

Circulating claim, as submitted.

Verdict

Source exists but framing is misleading

Confidence

Medium
§

Summary

The paper behind this post is real and the numbers are quoted correctly. A 150-million-parameter model called BDH-CQ, from the AI lab Pathway, did report scoring 29.5% on the public ARC-AGI-1 test at a computed cost of $0.0007 per task, and Pathway did publish the 11x cost comparison against GPT-5.6 Luna. The problem is what the comparison leaves out. Luna (Low) is the weakest setting of OpenAI's cheapest model tier, and the ARC Prize Foundation's own verified testing shows that same Luna model scoring 90.7% on ARC-AGI-1 at max effort, so the real accuracy gap is closer to 61 points than 5. The two cost figures are also measured differently: Pathway calculated its own from its hardware time at an assumed GPU rate with no profit margin, while Luna's comes from retail API pricing. Pathway itself says BDH-CQ is specialized for this kind of visual puzzle while the models it is compared to are general-purpose, so it is not a substitute for them at any price. The score is self-reported on a public test set rather than verified by the benchmark operator, and the paper reportedly does not publish a contamination check, so how much of the result reflects genuine reasoning versus familiarity with similar training data remains open.

§

The readings

key figures from the evidence
29.5 %

BDH-CQ self-reported pass@2 on ARC-AGI-1 public set

90.7 %

GPT-5.6 Luna verified score at max effort, ARC-AGI-1 semi-private

34.2 %

GPT-5.6 Luna (Low) score used in vendor's comparison

§

Why this verdict

Every figure in the claim traces correctly to a real preprint and a real vendor release, and the claim is unusually faithful in that it reports the accuracy gap rather than burying it. But the comparison it repeats is built on a chosen comparator, the weakest configuration of OpenAI's budget tier, and on two cost figures computed by different methods, one of them the vendor's own unmargined hardware time. Against the benchmark operator's verified number for the same GPT-5.6 Luna model after the same price cut, 90.7% on ARC-AGI-1, the "5-point tradeoff" framing collapses. I considered and rejected "Mostly accurate," because the comparator and cost-basis omissions materially change what a reasonable reader takes away rather than merely simplifying it; "Partially accurate but misleading," because no individual number in the claim is wrong; and "False," because the underlying result is genuine and correctly transcribed. Confidence is Medium, not High: the benchmark comparison rests on vendor-run, self-reported public-set numbers with no benchmark-operator verification, and I did not read the full PDF. As of 2026-08-19. ---
§

Evidence

The paper is real and the numbers are real. The arXiv abstract states directly: "A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of \$0.0007 per task. This operating point breaks through the previously reported ARC-AGI-1 cost-accuracy Pareto frontier, establishing a new state of the art in benchmark cost efficiency." The paper describes the mechanism the post describes: inputs presented at inference time continuously update the model's recurrent memory, and the model then solves a query through iterative computation in a high-dimensional latent space, without verbalizing its intermediate reasoning .

The vendor's own release carries the comparison the claim repeats: BDH-CQ scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.0007 per task, and runs approximately 11 times as cheaply per task as GPT 5.6 Luna (Low), even after accounting for OpenAI's 80% price cut of 5.6 Luna on July 30th, with Luna scoring 34.2% against BDH-CQ's 29.5% .

The 80% price cut is independently confirmed by the benchmark operator: on July 30, 2026, OpenAI announced an 80% price reduction for GPT-5.6 Luna .

Two facts in the primary sources sharply reframe the comparison. First, the cost figures are not measured the same way. Pathway itself discloses: Pathway's cost is computed from measured hardware time, while comparison costs are those reported to the leaderboard and may reflect API pricing for generalist models . The underlying rate is an assumption, not a price paid: inference takes roughly 0.85 H200 GPU-seconds per task at an assumed $3 per H200-hour, which is what pulls the price so low .

Second, "Luna (Low)" is the weakest configuration of OpenAI's cheapest tier, not GPT-5.6's reasoning performance. The benchmark operator's own verified figure for the same model after the same price cut is far higher: at max reasoning effort under the new pricing, Luna scores 90.7% on ARC-AGI-1 Semi-Private at $0.07/task and 59.6% on ARC-AGI-2 Semi-Private at $0.18/task .

Pathway also states the scope limit the post drops. On its own research page: "BDH-CQ is specialized for this kind of visual reasoning. While the commercial models on the leaderboard are general-purpose systems, the result matters because meaningful reasoning performance was achieved with a small fraction of the inference compute used by leading reasoning systems."

On verification status, ARC Prize draws a bright line: scores are self-reported unless noted otherwise, results on the ARC-AGI-1 and ARC-AGI-2 semi-private sets are run and verified by ARC Prize, and everything else is scored on a public set and self-reported . BDH-CQ's 29.5% is a public-set number. Pathway's own page invites, rather than reports, API-level external validation.


§

Findings

✓ What's accurate 7

  • The paper exists and is correctly attributed. arXiv:2608.09888, submitted 10 August 2026, authors Engdahl, Kosowski, Chorowski and others, from Pathway.
  • 150M parameters, latent-space iterative reasoning without verbalized chain-of-thought: accurately described.
  • 29.5% and $0.0007 per task are the paper's actual headline figures, quoted correctly.
  • The 34.2% figure for GPT-5.6 Luna (Low) and the roughly 11x post-price-cut cost ratio are exactly what Pathway published. The arithmetic checks: $0.040 × 0.20 ÷ $0.0007 = 11.4x.
  • The 80% OpenAI price cut on GPT-5.6 Luna, dated 30 July 2026, is confirmed independently by ARC Prize.
  • The claim states the accuracy gap rather than hiding it. Most aggregator coverage kept the multiplier and dropped the gap; this claim does not.
  • The paper does contain the controlled-intervention analysis the post mentions.

≈ What's misleading 6

  • **Benchmark cherry picking:** the claim compares against GPT-5.6 Luna (Low), the lowest-reasoning-effort setting of OpenAI's cheapest tier, and presents the resulting 4.7-point gap as the accuracy cost of switching. The benchmark operator's verified figure for the same model after the same price cut is 90.7% on ARC-AGI-1 at $0.07 per task. Against that configuration the gap is about 61 points, not 5. The comparator choice, not the result, produces the "near-frontier" impression.
  • **Harness mismatch:** BDH-CQ's 29.5% is self-reported on the 400-task public evaluation set. ARC Prize's verified Luna figures are on the semi-private set. The claim presents the two numbers as though they came off one board under one protocol.
  • **Cost compute omission:** the two cost figures are computed on different bases and are not directly comparable. BDH-CQ's $0.0007 is Pathway's own hardware time at an assumed $3 per H200-hour, carrying no serving margin, no availability overhead, and no failed-call retries. Luna's figure is retail API pricing, which includes OpenAI's margin. Pathway discloses this; the claim does not. The post's slide text goes further, describing the computed cost as "enabling direct cross-model comparison," which is the opposite of what the disclosure supports.
  • **Omitted qualifier:** "scored 29.5%" drops pass@2, drops that this is the highest of three effort settings (Low gives 21%), and drops that the score is self-reported rather than ARC Prize verified.
  • **Capability extrapolation:** the surrounding post frames this as "near-frontier reasoning" and a "1000x smaller" model buying most of a frontier model's ability. Pathway's own page says BDH-CQ "is specialized for this kind of visual reasoning" while the leaderboard comparators are general-purpose systems. A model that does ARC grids and nothing else is not substitutable for GPT-5.6 Luna at any price, which is the practical meaning "11x cheaper" conveys to a reader.
  • **Eval contamination, unresolved rather than demonstrated:** the training mixture reportedly blends privately curated examples with public ARC-derived datasets, with no published deduplication or contamination analysis, and the evaluation is on the public split. The post's slide asserts flatly that "ARC-AGI-1 is designed to resist memorization and test genuine few-shot reasoning," which describes the benchmark's design intent and does not establish that this particular result is uncontaminated.

? What's uncertain 5

  • I retrieved the arXiv abstract page and the PDF's contributions section, but did not read the full paper end to end. The training-mixture composition, the three-effort-setting breakdown, and the reproducibility limitations come from secondary and tertiary analyses of the paper, not from my own reading of the full text. Those specific details are reported at Medium confidence.
  • The independence of the two named reproductions cannot be settled from public sources. Pathway's blog presents Richard Zhong as a reproducer, while syndicated versions of the same release describe him as a co-author. If the auditors are co-authors, this is not independent verification.
  • Whether the 34.2% Luna (Low) figure is a public-set or semi-private-set number is not resolvable from the pages I reached, which makes the exact comparability of the two accuracy figures indeterminate.
  • No ARC Prize verified entry for BDH-CQ was found, and Pathway's page invites external API validation rather than reporting it. I state this as not found rather than as confirmed absence.
  • Secondary claim, not investigated: reporting that Pathway raised additional funding at a $500M valuation, bringing total seed funding to $30M. That is a business claim requiring separate sourcing analysis.
Distortion flags benchmark cherry picking harness mismatch cost compute omission omitted qualifier capability extrapolation eval contamination demo to product conflation
§

Sources

10 of 10 linked to records
[1]

arXiv:2608.09888, "BDH-CQ: In-Context Learning with Recurrent Latent Reasoning," Engdahl, Kosowski, Chorowski et al., submitted 10 Aug 2026 (abstract page and PDF contributions section retrieved; full PDF not read end to end)

primary preprint, unrefereed, authored by the vendor
https://arxiv.org/abs/2608.09888 ↗
[2]

Pathway research page, "Reasoning at a Fraction of the Compute"

primary vendor
https://pathway.com/research/introducing-bdh-cq ↗
[3]

ARC Prize official results page, GPT-5.6 Luna re-test after the 30 July 2026 price cut

primary benchmark operator of record
https://arcprize.org/results/openai-gpt-5-6-luna-2026-07-30 ↗
[4]

ARC Prize Verified Testing Policy and Community Leaderboard

primary benchmark operator of record
https://arcprize.org/policy ↗
[5]

Pathway press release via Business Wire, 11 Aug 2026

primary vendor-issued wire release
https://www.businesswire.com/news/home/20260811268264/en/ ↗
[6]

pathwaycom/arc-task-gen repository

primary vendor code release
https://github.com/pathwaycom/arc-task-gen ↗
[7]

emergentmind paper analysis of 2608.09888

tertiary automated paper-analysis aggregator, uncorroborated on detail
https://www.emergentmind.com/papers/2608.09888 ↗
[8]

explainx.ai analysis, "Pathway BDH-CQ: 150M Model, 11x Cheaper Than GPT-5.6"

secondary named tech commentary
https://explainx.ai/blog/pathway-bdh-cq-150m-post-transformer-arc-agi-august-2026 ↗
[9]

AI Weekly alert on BDH-CQ

secondary trade newsletter
https://aiweekly.co/alerts/150m-bdh-cq-hits-295-on-arc-agi-1-for-00007-a-task ↗
[10]

Analytics India Magazine coverage

secondary named-outlet trade press
https://analyticsindiamag.com/ai-news/pathway-claims-11x-lower-ai-reasoning-costs-with-150m-parameter-model ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →