Nine AI jurists with locked judicial philosophies deliberate real Supreme Court cases; two AI journalists cover the fight as a podcast. Everything you can watch, hear, or score comes from one immutable event log per case. The five stops below hit every mode — each says what to try and what it proves.
The pixel courtroom replays the hero case — Younge v. Fulton County D.A., a prediction filed before the real Court rules. Press play; audio drives the visuals.
Look for: The Record panel on the right — it IS the event log, rendered. Click any line to seek. Watch the vote board when the Pragmatist flips (affirm → reverse).
▶ open the hero replay problem value · falsifiable betTwo AI journalists cut each deliberation into a broadcast: news desk, then a hard cut into courtroom tape. They may frame the tape; they may never rewrite a word of it — we diff the record after every episode (0 content diffs / 193 events).
Look for: navy STUDIO lines in The Record between verbatim tape lines — two societies, one shared factual record.
🎙 open podcast mode innovation · agents auditing agentsOne button starts a real, uncached deliberation on this Alibaba Cloud instance. Events stream in the moment each agent produces them. A full session runs ~10–15 minutes — start it, continue the tour, come back for the verdict.
Look for: private ballots landing before anyone speaks (that's the anti-sycophancy design — flips are measured persuasion, not mimicry). One session at a time; if the panel is in recess, another judge beat you to it.
⚖ open live bench technical depth · agents run on cloudWe benchmarked the society against the real Court on post-training-cutoff cases, paired on the same 24 cases: solo model 75%, silent jury 66.7%, deliberating society 66.7% — debate moved persuasion, not accuracy (sign test p=1.0), and nothing beat the always-guess-reverse baseline (75%). Our first readout claimed debate cost 10 points; that was a cross-pool artifact, and the correction is published in the exhibit itself.
Look for: the contamination guard (post-cutoff benchmark vs. famous-case memorization check), the baseline row, and the on-the-record correction.
📊 open the findings technical depth · measurement over vibesEvery jurist is steered by one system prompt — its whole constitution. On the landing page, click any agent card to read it verbatim, including the exact trigger that permits a vote change ("you change your vote ONLY when…").
Look for: the architecture diagram further up the landing page — the dots are the actual data flow. The same diagram is under ARCHITECTURE inside the player.
🧠 meet the agents presentation · nothing hiddenepisodes/<case>/events.jsonl — the same log that scored
the benchmark; the podcast never re-voices it