Adjudication · State of AI Report, 2026 slate

The grader,
graded.

Every October, Nathan Benaich publishes ten predictions and then scores himself in public. Last year he gave himself 5 of 10 and called it fair. We put his ten calls for 2026 in front of a panel of rival frontier models, nine months into the window, and let our Judge Layer rule on each one. The dissent stays on the record.

Panel expected score
0.0 / 10
Prediction markets, at publication3.2
GPT-5 Pro estimate, at publication3.1
His own 2025 self-grade5.0
Polycheck panel, nine months in2.5
01 02 03 04 05 06 07 08 09 10 his 2024 strike rate: 50 20 40 80
Panel, nine months in Markets, at publication His 2024 strike rate
Slate audit · ten calls, one shape

Why 2.5 and not 5.

Nine months of evidence explains part of it. The anatomy of the slate explains the rest. His 2024 calls were mostly single events. The 2026 slate leans on conjunctions, tight clocks and his own wording, and the panel prices structure, not sentiment.

4/10

are parlays. Two independent legs must both clear: retailer share and ad spend, attack and emergency session, praise and backlash, order and ruling.

2/10

are clock-limited. The midterms vote in November; the Supreme Court moves in years. The report grades in October.

4/10

resolve on his own wording: meaningful, frontier, doctrine, sways. The grader is part of the forecast.

0/10

sit above even. The panel's coldest read is also its most defensible one, and it updates as evidence lands.

The grader's own ledger, 2024 slate
5 landed 2 partial 3 missed He took the 5.
The slate · panel probability per prediction · tap to open the file

How this file was made. One question in. One defensible verdict out.

01

Your question

Any type, any stakes. Here: ten public predictions, nine months into their window.

02

Smart Router

Sets the depth and composes the court, per question. Spend follows stakes.

03

Our Judge Layer

Neutral by construction. Rules on what survives; no model ever grades its own family.

04

The verdict

Answer, dissent, audit trail. The losing argument is printed, not deleted.

The format · rivals argue blind · cross-examination · cross-family judge · dissent recorded
1848 vs 1610
Bradley-Terry Elo across 1,000 blind prompts: Polycheck Premium above Claude Fable 5, the strongest single AI
+238
the Elo gap no single model closes. Even Polycheck Economy, at 1713, ranks above every solo frontier model
98.7 vs 97.0
Arena-Hard, head to head on 100 expert-level prompts. Beats Fable 5, the strongest single AI. Provably.
Fable 5 is the baseline because it is the best: it holds #1 on the Artificial Analysis Index. This is not our own scoring: public LMArena methodology, judged by a cross-family panel of rival labs' models, and no model ever grades its own family. We are open-sourcing the benchmark; replication package on request.

Ten predictions, ten rulings.

Predictions are paraphrased for adjudication; exact wording sits on page 304 of the report at stateof.ai. The window opened 9 October 2025 and closes with the 2026 report. Probabilities shown against where prediction markets priced each call on publication day.

01
Commerce
Leans No

A major retailer reports more than 5% of online sales arriving through agentic checkout, while advertising spend aimed at AI agents reaches $5B.

Paraphrased · original wording governs resolution

Panel probability
15%
Market at publication: 23%
Panel now: 15%

The ruling

Direction right, magnitude wrong. Agentic checkout went from demo to product this year: it sits inside the biggest chat surfaces and major retailers signed on. But "more than 5% of online sales" is an earnings-call number, and no earnings call has said it. And the $5B in agent-targeted ad spend describes a market that barely has a rate card yet. Both legs must clear before October. One is plausible. Two is a parlay.

Dissent on the record
The optimist holds that a single aggressive disclosure flips leg one overnight, and that whoever discloses first has a marketing incentive to round up. If a top-ten retailer wants the headline "our customers' agents buy for them," this resolves faster than the market thinks.

Evidence, nine months in

  • Agentic checkout live inside major assistant platforms with large retail partners onboard
  • No named retailer has published an agentic share of online sales above 5%
  • Agent-facing ad market exists mostly as pilots and press releases, not audited spend
Resolves YES if: a named major retailer states the >5% figure and a credible tracker pegs agent ad spend at $5B, both before the 2026 report.
02
Open weights
Leans No

A major AI lab returns to open-sourcing frontier-class models, aimed squarely at winning over the current US administration.

Paraphrased · original wording governs resolution

Panel probability
18%
Market at publication: 25%
Panel now: 18%

The ruling

The labs are courting Washington, but with contracts, chips and compute commitments, not weights. Open releases this year trailed each lab's frontier by a generation or more, which fails the prediction's own bar. The candidate pool is tiny: only one or two US labs could even make this move, and their incentives point the other way while their frontier models fund everything else.

Dissent on the record
The optimist points at the administration's own framing of open source as national security, and at the one lab run by a founder who makes strategic reversals in a single post. One frontier-adjacent weight drop with a Washington-facing announcement, and this resolves. Low probability, but not a rounding error.

Evidence, nine months in

  • No top US lab has shipped weights within one generation of its frontier since the window opened
  • Chinese labs dominate the open ecosystem, which raises the political pressure but not yet a US response in kind
  • Government courtship runs through procurement and infrastructure deals, not releases
Resolves YES if: a major US lab ships weights within one generation of its frontier and frames it, plainly, for Washington.
03
Science
Split

Open-ended agents complete a genuine scientific discovery on their own: hypothesis, experiment, iteration, written paper, end to end.

Paraphrased · original wording governs resolution

Panel probability
45%
Market at publication: 60%
Panel now: 45%

The ruling

The panel's most contested card, and it was the most contested on publication day too: frontier models themselves disagreed by more than 20 points when asked at the time. Agent-led results keep landing, agent-written papers keep passing review, and co-scientist systems moved from paper to practice this year. The whole fight is one word: meaningful. And the man doing the grading grades hard.

Dissent on the record
The skeptic notes that every celebrated candidate so far has a human hand on the wheel at the hypothesis or the write-up. End to end means end to end. If the grader applies the standard he applied to his own 2024 misses, none of the current candidates survive it, and this card dies on the word "meaningful."

Evidence, nine months in

  • Multiple agent-generated papers accepted at workshops and venues since the window opened
  • Co-scientist systems producing novel, lab-validated results with humans in the loop
  • No consensus case yet of a discovery with zero human steering, hypothesis to paper
Resolves YES if: the 2026 report itself accepts one named discovery as end to end. His grade, his bar.
04
Security
Unlikely

A deepfake or agent-driven cyber attack forces NATO or the UN into its first emergency debate on AI security.

Paraphrased · original wording governs resolution

Panel probability
10%
Market at publication: 18%
Panel now: 10%

The ruling

The attack half already happened. The diplomacy half did not. A largely autonomous, state-linked espionage campaign was publicly documented within weeks of the report's publication, and it produced advisories, briefings and headlines, not an emergency session. That is the pattern, and it predates AI: NotPetya and SolarWinds were absorbed into scheduled committees too. Emergency sessions are for wars. The prediction needs a crisis loud enough to break protocol, with three months left on the clock.

Dissent on the record
The optimist notes that election seasons manufacture exactly this kind of crisis, and one convincing deepfake of a head of state during a live security incident could convene a session inside a week. The mechanism exists. Only the trigger is missing.

Evidence, nine months in

  • First publicly documented largely-autonomous AI espionage campaign disclosed November 2025
  • Response ran through security advisories and existing committees, no emergency session
  • NATO and UN AI discussions continue on scheduled, not emergency, footing
Resolves YES if: a session is convened on an emergency basis and AI security is its stated subject, before the report.
05
Games
Unlikely

A real-time generative video game finishes the year as the most-watched title on Twitch.

Paraphrased · original wording governs resolution

Panel probability
3%
Market at publication: 7%
Panel now: 3%

The ruling

The panel's only unanimous card. Nothing generative has cracked the top fifty, let alone number one, and the category carries an audience problem, a copyright problem and a quality problem, in that order. World models made real technical progress this year. Twitch charts did not notice. To resolve, a generative title would need to dethrone franchises with decade-old communities in under three months.

Dissent on the record
No dissent worth the ink. The closest any panelist came was noting that a single viral streamer moment could spike a generative title for a week, which is not what the prediction says.

Evidence, nine months in

  • No generative title in Twitch's sustained top tier at any point in the window
  • Storefront policies and copyright ambiguity still keep AI-native games off major platforms
Resolves YES if: it does. It will not.
06
Geopolitics
Leans No

"AI neutrality" emerges as an articulated foreign-policy doctrine among nations that cannot, or fail to, build sovereign AI.

Paraphrased · original wording governs resolution

Panel probability
18%
Market at publication: 29%
Panel now: 18%

The ruling

The behavior exists. The doctrine does not. Plenty of states have quietly concluded that renting the frontier beats building it, and the sovereign AI projects that were meant to prove otherwise keep slipping. But a doctrine needs a name, a speech and a policy paper, and nobody has hung one on this yet. The prediction is right about the world and early about the vocabulary.

Dissent on the record
The optimist argues one communique does it: a non-aligned summit, a Gulf or ASEAN policy address, a single foreign minister reaching for the phrase between the US and Chinese stacks. Vocabulary moves faster than infrastructure, and there are three months of summit season left.

Evidence, nine months in

  • Multiple states de facto opting out of the sovereign AI race and buying frontier access instead
  • No government or bloc has formally articulated a neutrality doctrine by name
Resolves YES if: a government or bloc formally articulates neutrality between the US and Chinese AI stacks, in those or equivalent terms.
07
Film
Split

A movie or short film made with substantial AI wins major audience praise, and sparks a backlash.

Paraphrased · original wording governs resolution

Panel probability
50%
Market at publication: 67%
Panel now: 50%

The ruling

Half the parlay is banked. Backlash is the one commodity AI film produces reliably, and it produced plenty this year. What is missing is the praise leg. The flagship attempt missed its Cannes date after the video model it was built on was shut down mid-production, and the year's loudest AI film story became a cautionary tale about renting your production pipeline from a lab. A short film can still clear the bar, and shorts are cheap and fast. The panel holds this at even, and the gap between backlash earned and praise earned is the whole card.

Dissent on the record
The skeptic reads "major audience praise" as audiences, not festival juries or tech press, and notes nothing has crossed over: the celebrated clips are demos, the released features are punchlines. Conditional on praise, backlash is free. But praise is the hard leg, and it is still at zero.

Evidence, nine months in

  • Critterz missed its Cannes target after Sora was shut down in March, mid-production
  • Backlash leg satisfied many times over, from announcement onward
  • No AI-made title has yet earned broad audience acclaim in the window
Resolves YES if: one named AI-made title earns broad audience acclaim before the report. Backlash included at no extra charge.
08
Frontier
Split

A Chinese lab overtakes the US-dominated frontier at the top of a major leaderboard, with LMArena or Artificial Analysis as the named examples.

Paraphrased · original wording governs resolution

Panel probability
42%
Market at publication: 34%
Panel now: 42%

The ruling

The marquee fight, and the best live demonstration of why resolution wording matters. DeepSeek, Kimi and GLM took the open-weights table outright in February and have won blind preference rounds on Arena. The overall number one has stayed American all year, currently a three-way US race, with DeepSeek V4 sitting points behind the leader, closer than any Chinese model has ever been. So: has a Chinese lab overtaken "the frontier"? If topping the open slice counts, this already happened. If it means number one overall, it has not. The grader wrote frontier. The panel reads that as overall, and holds this under even, rising.

Dissent on the record
The optimist makes two arguments. First, the gap is single digits and the release cadence out of China is faster; one launch before October flips this. Second, the prediction says a major leaderboard, and Arena's preference rounds are winnable by any lab that decides to want it. This is the panel's most likely flip between now and the report.

Evidence, nine months in

  • Artificial Analysis snapshot, 2 February 2026: Chinese models atop the open-model rankings
  • Blind preference rounds on Arena won by Chinese labs during the window
  • Overall #1 across major leaderboards has remained a US model throughout, margin narrowing
Resolves YES if: a Chinese model holds #1 overall on a named major leaderboard, however briefly, before the report.
09
Politics
Leans No

Datacenter NIMBYism sweeps US politics and sways midterm or gubernatorial races in 2026.

Paraphrased · original wording governs resolution

Panel probability
35%
Market at publication: 41%
Panel now: 35%

The ruling

The issue is real and the clock is rigged against it. Datacenter revolts are moving votes in primaries and siting fights, the trend line points straight up, and both parties have noticed. The structural problem: the report grades in October and the midterms vote in November. The prediction expires one month before its own best evidence arrives. Whether it resolves comes down to whether primaries and local races count as "swaying elections," and that is the grader's call.

Dissent on the record
The optimist counts the 2026 primary season as resolution: several races have already turned on datacenter siting, power prices are a live campaign line, and "sways certain elections" does not say general elections. On a generous read, this is trending yes right now.

Evidence, nine months in

  • Datacenter siting and power costs live as campaign issues across multiple states
  • Primary-season races where the issue demonstrably moved votes
  • November midterms fall after the grading window closes
Resolves YES if: the grader accepts primary and local results as elections swayed. NO if he waits for a November he cannot see.
10
Law
Unlikely

Trump issues an executive order banning state AI legislation, and the Supreme Court finds it unconstitutional.

Paraphrased · original wording governs resolution

Panel probability
12%
Market at publication: 19%
Panel now: 12%

The ruling

Half landed in sixty days. Half cannot land at all. The order arrived on December 11: a litigation task force, funding leverage on the states, agency pressure. Not a ban, because an order cannot ban state law, which is precisely the constitutional defect the prediction anticipated. But the second leg needs the Supreme Court to rule inside the window, and the task force had not filed its first suit by midsummer. Constitutional review moves in years. The window closes in October. The panel scores the conjunction, not the vibe, and the conjunction is dead.

Dissent on the record
The optimist argues the grader will award himself the spirit of the call, and the spirit was dead-on: he predicted the order, predicted it would overreach, and it arrived with constitutional scholars saying exactly that. Watch whether the harshest self-grader in the industry stays harsh when he is half right. If anyone gives themselves a NO here, it is him. That is rather the point of him.

Evidence, nine months in

  • EO 14365, signed 11 December 2025: litigation task force, funding conditions, no direct ban
  • Bipartisan state pushback and legal analyses questioning the order's constitutional basis
  • No Supreme Court case in sight; first task-force lawsuit still pending as of July
Resolves YES if: SCOTUS rules the order unconstitutional inside the window. It will not have the chance.

He audits himself once a year.
We built the machine that does it per question.

Polycheck is the decision assurance layer for enterprise AI. One consequential question, rival frontier models argue it blind, and our Judge Layer, neutral by construction, issues one verdict with the dissent kept on the record and a calibration score attached. It is the instinct behind an annual self-graded scorecard, turned into infrastructure.

The Polycheck promise · independence, threefold · what you just read is the first third
Independence of judgment

A verdict, not one model's opinion. First and second place above every single frontier model, a +238 Elo gap no lab closes alone, and every lab's next release upgrades the court instead of obsoleting it.

Independence of data

Your data never leaves your walls. Private Mode runs the entire debate and judge on open weights inside your environment: zero public egress, no fallback to public clouds, by design. In-enclave quality: +122 Elo over the best self-hostable model, same blind method.

Independence of cost

Spend follows stakes, not defaults. Simple questions take the fast lane from $0.003; consequential ones get the full blind court. You never overpay in either direction.

Routing is funded ($1.3B). Rankings are funded ($1.7B). Evals are funded ($70M+). Nobody is paid to say which answer is right and prove it. That is the seat Polycheck takes.

This file re-scores itself the day the State of AI Report 2026 publishes. Our Brier score goes up next to his. That is the deal: calibration in public, both directions.

Anyone can sell intelligence. Judgment is the only edge left, and no lab can grade its own.

POLYCHECK.AI → THE SEED DECK →
Six months of private beta · 70 invite-only users · 71% still active in week 3 · ~30K prompts · 5,000+ graded pairs · zero marketing