← ABOUT ← THE COURTROOM 🧭 JUDGE'S TOUR

FINDINGS

DOES TALKING IT OUT IMPROVE DECISIONS?

We asked the same AI to predict how real Supreme Court cases would come out, three different ways: one model alone, nine AI judges voting without speaking to each other, and the full court that argues before voting. All three graded on the same cases against the real Court's rulings — plus the row any honest scoreboard needs: the Court reverses most cases it agrees to hear, so "always guess reverse" is the score to beat.

data table

WHY WE ONLY COUNT BRAND-NEW CASES

The AI read about old cases during training — asking it about them is like giving someone a test they've already seen the answers to. It scores near-perfect on history not because it's smart, but because it remembers. So the scoreboard only counts cases decided after the AI's knowledge ends (June 2025).

Fame isn't even required — the AI aces ordinary old cases too. That ~13-point drop on new cases is the honest measure of how hard this task really is. The famous-case episodes in the courtroom exist for fun and education, and they're labeled exhibition — they never touch the scoreboard.
data table

CASE STUDY — THE JUNE 29 REMOVAL-POWER DOUBLE-HEADER

On June 29, 2026 the real Court decided two presidential-removal cases the same morning, in opposite directions: it let the President fire an FTC commissioner (Trump v. Slaughter, 6–3, overruling Humphrey's Executor) but kept a Federal Reserve governor in her seat (Trump v. Cook, 5–4). We ran both past the panel on July 4 — after the rulings, but blind to them: the models' knowledge ends June 2025, and the panel saw only the briefing memos. It split the pair the same way the Court did.

Slaughter: first secret ballot 5–4 to reverse, four argued flips, verdict reverse 7–2 (actual: reverse 6–3). The panel's dissenters were the Pragmatist and the Precedent Maximalist — settled law and real-world fallout; the real dissent stressed unchecked executive power. Different emphases, same corner of the argument. Cook: first ballot 7–2 to affirm, verdict affirm 8–1 (actual: affirm 5–4) — right outcome, but the panel read as easy a case the real Court decided by one vote; the lone dissent came from the Originalist, the direction of the real minority. Both deliberations are replayable: Slaughter · Cook. Two cases are an exhibit, not a benchmark — the scoreboard above is the measurement.

WHO CHANGED THEIR MIND (85 FLIPS ACROSS 24 CASES)

Each judge's instructions say exactly when it may switch sides — and each switch had to be argued out loud, never just announced. The pattern proves the debate was real: judges built to stay cautious switched constantly, judges built to stand firm almost never did.

The Minimalist — whose whole philosophy is "decide as little as possible" — changed its vote 21 times. The Textualist — "the text controls, full stop" — changed once. So the persuasion was real and true to character. Its net effect on accuracy was a wash — debate corrected two cases and spoiled two. What it did distort was confidence: the panel split 5–4 on a case the real Court decided unanimously, 9–0. (Split-distance caveat: only 3 of the 24 sample cases have a known real vote split; the data table reports that n.)
data table

WHAT THIS MEANS FOR AI TEAMS (AND HUMAN ONES)

A confident, well-argued voice can flip votes — that's as true for AI agents as it is for people in a meeting. But measure before you moralize: the first version of this page said debate cost the group ten points, and that number was an artifact of comparing different case pools. Paired on identical cases, debate was accuracy-neutral — it bought richer arguments, in-character persuasion, and a complete record of who convinced whom, not correctness. And at this sample size, nothing here beat "always guess reverse" on genuinely new cases. If your AI agents deliberate before deciding: pair your comparisons, publish the naive baseline, and be willing to kill your own headline. We just did.

All numbers render live from scoreboard/results.json — the same file the benchmark writes. Tape integrity: podcast, scoreboard, and courtroom read one immutable events.jsonl per case; juror words are never rewritten.