JUDGE'S TOUR

split decision · two agent societies, one immutable record · 90 seconds to orient

Nine AI jurists with locked judicial philosophies deliberate real Supreme Court cases; two AI journalists cover the fight as a podcast. Everything you can watch, hear, or score comes from one immutable event log per case. The five stops below hit every mode — each says what to try and what it proves.

1 Watch a deliberation replay ~2 min

The pixel courtroom replays the hero case — Younge v. Fulton County D.A., a prediction filed before the real Court rules. Press play; audio drives the visuals.

Look for: The Record panel on the right — it IS the event log, rendered. Click any line to seek. Watch the vote board when the Pragmatist flips (affirm → reverse).

▶ open the hero replay problem value · falsifiable bet
2 Podcast mode — the second society ~2 min

Two AI journalists cut each deliberation into a broadcast: news desk, then a hard cut into courtroom tape. They may frame the tape; they may never rewrite a word of it — we diff the record after every episode (0 content diffs / 193 events).

Look for: navy STUDIO lines in The Record between verbatim tape lines — two societies, one shared factual record.

🎙 open podcast mode innovation · agents auditing agents
3 Convene the panel yourself — live ~1 min to start

One button starts a real, uncached deliberation on this Alibaba Cloud instance. Events stream in the moment each agent produces them. A full session runs ~10–15 minutes — start it, continue the tour, come back for the verdict.

Look for: private ballots landing before anyone speaks (that's the anti-sycophancy design — flips are measured persuasion, not mimicry). One session at a time; if the panel is in recess, another judge beat you to it.

⚖ open live bench technical depth · agents run on cloud
4 The findings — an honest negative result ~2 min

We benchmarked the society against the real Court on post-training-cutoff cases, paired on the same 24 cases: solo model 75%, silent jury 66.7%, deliberating society 66.7% — debate moved persuasion, not accuracy (sign test p=1.0), and nothing beat the always-guess-reverse baseline (75%). Our first readout claimed debate cost 10 points; that was a cross-pool artifact, and the correction is published in the exhibit itself.

Look for: the contamination guard (post-cutoff benchmark vs. famous-case memorization check), the baseline row, and the on-the-record correction.

📊 open the findings technical depth · measurement over vibes
5 Read the agents' actual prompts ~1 min

Every jurist is steered by one system prompt — its whole constitution. On the landing page, click any agent card to read it verbatim, including the exact trigger that permits a vote change ("you change your vote ONLY when…").

Look for: the architecture diagram further up the landing page — the dots are the actual data flow. The same diagram is under ARCHITECTURE inside the player.

🧠 meet the agents presentation · nothing hidden
Verify it yourself
· Podcast RSS (4 episodes, 3 pending-case predictions on the record): feed.xml
· Every replay reads episodes/<case>/events.jsonl — the same log that scored the benchmark; the podcast never re-voices it
· Source, prompts, pipeline, and reproduction steps: github.com/tmoody1973/split-decision (Apache-2.0)
· Backend: SAS instance (Singapore) runs the deliberation pipeline + this site; OSS hosts the podcast
Qwen Cloud Global AI Hackathon · Track 3 · Agent Society