FINDINGS
DOES TALKING IT OUT IMPROVE DECISIONS?
We asked the same AI to predict how real Supreme Court cases would come out, three different ways: one model alone, nine AI judges voting without speaking to each other, and the full court that argues before voting. All three graded on the same cases against the real Court's rulings — plus the row any honest scoreboard needs: the Court reverses most cases it agrees to hear, so "always guess reverse" is the score to beat.
data table
WHY WE ONLY COUNT BRAND-NEW CASES
The AI read about old cases during training — asking it about them is like giving someone a test they've already seen the answers to. It scores near-perfect on history not because it's smart, but because it remembers. So the scoreboard only counts cases decided after the AI's knowledge ends (June 2025).
data table
CASE STUDY — THE JUNE 29 REMOVAL-POWER DOUBLE-HEADER
On June 29, 2026 the real Court decided two presidential-removal cases the same morning, in opposite directions: it let the President fire an FTC commissioner (Trump v. Slaughter, 6–3, overruling Humphrey's Executor) but kept a Federal Reserve governor in her seat (Trump v. Cook, 5–4). We ran both past the panel on July 4 — after the rulings, but blind to them: the models' knowledge ends June 2025, and the panel saw only the briefing memos. It split the pair the same way the Court did.
WHO CHANGED THEIR MIND (85 FLIPS ACROSS 24 CASES)
Each judge's instructions say exactly when it may switch sides — and each switch had to be argued out loud, never just announced. The pattern proves the debate was real: judges built to stay cautious switched constantly, judges built to stand firm almost never did.
data table
WHAT THIS MEANS FOR AI TEAMS (AND HUMAN ONES)
All numbers render live from scoreboard/results.json — the same file the
benchmark writes. Tape integrity: podcast, scoreboard, and courtroom read one immutable
events.jsonl per case; juror words are never rewritten.