Everyone building multi-agent AI eventually pays for the same assumption: that making agents debate produces better decisions than asking one model directly. Debate costs tokens, time, and complexity — but does it actually help?
Split Decision tests that on one of the few public sets of genuinely hard, formally graded judgment calls: real Supreme Court cases. Nine AI jurists — each locked to a distinct judicial philosophy — read the same case, vote in private, argue for up to five rounds, and file a verdict. Their predictions are scored against what the real Court actually did. For cases the Court hasn't decided yet, the panel's ruling stands as a falsifiable public prediction.
The answer surprised us: deliberation made the panel more persuasive and less accurate — and we published that instead of hiding it. The full numbers are on the findings page.
HOW THE SOCIETY WORKS
- The Clerk reads a real case — filings, facts, the question presented — and writes a neutral bench memo.
- Nine jurists vote in private first. No one has spoken yet, so these votes are each philosophy's honest, independent read.
- They argue. The Foreperson (who never votes) picks the sharpest disagreements and makes those jurists face each other directly.
- They vote in private again. Any changed vote is a measured act of persuasion — argued in the transcript, never just announced.
- After up to five rounds, the panel files its verdict. Decided cases get graded against the real Court; pending cases go on the record as predictions.
- A second society — two AI journalists — cuts the deliberation into a podcast. They may frame and react, but every clip of jurist speech is played verbatim from the log. They are never allowed to re-voice the court.
- Everything you can watch, hear, or score comes from one immutable event log per case. The pixel courtroom just replays it.
MEET THE AGENTS — AND READ THEIR ACTUAL PROMPTS
Every agent below is steered by one system prompt — its whole constitution. Click any card to read it verbatim, including the anti-sycophancy trigger that tells each jurist exactly when it is allowed to change its vote. Nothing else makes them who they are.
THE HONEST FINDING
(n=94 cases)
(n=93)
(n=24)
More debate, better arguments, worse predictions. Eloquent majorities pulled correct independent votes into 5–4 splits on cases the real Court decided 9–0. If your agents can be talked out of correct answers, deliberation is a liability — independence is a feature you have to engineer for. See the full analysis →
TWO RULES WE COULDN'T BREAK
The podcast, the scoreboard, and the courtroom all read the same events.jsonl. The journalists select clips by event index — they can characterize what a jurist said, but they can never rewrite, re-voice, or paraphrase it. After every production run we diff the log: zero content changes, every time.
The model already knows history — it scores ~96% on pre-2025 cases because it's remembering, not predicting. So the benchmark only uses cases decided after June 2025, plus pending cases nothing could have memorized. Famous landmark episodes exist for civic education and are labeled exhibition, never benchmark.
LISTEN & WATCH
▶ The pixel courtroom — every deliberation replays from the log, with "The Record" transcript beside it.
🎙 Podcast mode — the newsroom covers the chamber: anchors at the desk, hard cuts to verbatim tape.
📡 RSS feed — subscribe in any podcast app.
Three predictions are currently on the record for cases the real Court hasn't decided. We'll be graded whether we like it or not.