The night I listened to nine machines argue
A few nights ago I sat in my kitchen and listened to nine voices argue about whether a president can fire the officials who regulate him. One voice was gravelly and Southern and impatient with everyone. One kept asking what the words meant when they were written. One wanted to know who would live with the consequences. They interrupted each other. They conceded points. About twelve minutes in, a voice I had built to be cautious changed its vote out loud and named the exact argument that moved it.
None of them are people. I had made all of them a couple of days earlier, for a hackathon.
I make radio for a living, so I think about voices and arguments more than is probably healthy. When Qwen Cloud announced a hackathon with an "agent society" track, my first thought wasn't a productivity tool. It was a question I couldn't find a straight answer to: everyone is wiring AI agents together and letting them talk, and almost nobody can tell you whether the talking helps. Ask and you get vibes. I wanted a number. And because of the radio thing, I wanted you to be able to hear the answer, not just read it.
Supreme Court cases turned out to be the honest test I was looking for. The Court grades in public. It tells you who won and by what vote, and there is no arguing with the answer key.
Three days on Qwen Cloud
My first API call hit Qwen Cloud on the morning of July 4th at 9:28 a.m. By submission the log showed 11,271 calls and about 35 million tokens, all inside the free tier. I know that because every single call appends a line to a cost log. That file became the project's conscience more than once.
Casting the models felt like casting a show. The nine jurists run on qwen3.7-plus, which turned out to have plenty of judgment for arguing doctrine at a fraction of the cost. The foreperson and the clerk run on qwen3.7-max, because the moderator should not be the least capable voice in the room. The boring formatting calls go to qwen3.6-flash. Nobody notices the intern, which is the point.
The voices were the part that felt like magic, and they almost didn't happen. I had planned to use CosyVoice. A live test on day one returned an error that, after some digging, meant the model only exists in the Beijing region. The detour led me to Qwen's Voice Design, where you describe a voice in prose the way a casting director would: "older male voice, gravelly, deliberate, slight Southern cadence, dry wit." It hands you back that person. I made twelve โ nine judges, a foreperson, two news anchors โ and sat there listening to preview clips like demo tapes, sending some back. Twenty years of radio and I have never cast a show that fast.
Not everything charmed me. The image model blocked my first episode-art prompts because I described the crimes in the cases instead of the scene I wanted drawn. Fair enough; lesson learned. Less fair: a content-moderation block once killed a deliberation forty minutes in, mid-round, because the engine retried the same blocked content like it was a network hiccup. Now moderation fails fast and loud, and only genuinely transient errors retry. The pipeline got more resilient because it got hurt.
Deployment had its own surprise. Alibaba's object storage force-downloads HTML pages on the default endpoint, which I discovered only when my deployed site downloaded itself instead of loading. The web app moved to nginx on a small Simple Application Server instance, which turned out to be a gift: with real compute behind the site, I could add a live bench where a visitor presses a button and convenes an actual deliberation, nine models arguing in real time on that server, nothing cached. Judges can make the machine perform on demand. On radio we call that going live, and it is always the part that makes you sweat.
The part where I was wrong
My first published scoreboard said that debate made the panel ten points less accurate. I was proud of that number. An honest negative result, I told myself. Look how rigorous.
Then I ran a mock judging pass against my own project, and one of the reviewer personas took my benchmark apart. The ten-point gap compared two different sets of cases. Paired on the identical 24, the story collapsed: silent jury, 66.7 percent. Debating court, 66.7 percent. The debate fixed two cases and broke two. And nothing โ not the solo model, not the jury, not the arguing court โ beat the dumbest possible strategy on that sample, which is to always guess "reverse," because the Supreme Court reverses most cases it takes.
I fixed the math and left the correction on the findings page, in public, next to the numbers it corrects. It stung to write. It is also the part of the project I trust the most.
The control experiment made the lesson sharper. I swapped the nine ideologues for nine neutral analysts, no philosophies at all, and ran the same 24 cases again overnight. Same accuracy, to the decimal. What changed was the character of the room: the neutral panel herded, reaching unanimous verdicts on half the cases in a round and a half, while the ideologues argued for three rounds and flipped along philosophical lines โ the cautious Minimalist changed its vote 21 times, the immovable Textualist once. Ideology shapes the debate. The model sets the score.
What the machines showed me
Ten years ago I wrote a piece for Radio Milwaukee arguing that what my city was missing was empathy. I did not expect that thought to be waiting for me inside a hackathon project.
As an exhibition, outside the scoreboard, I had the panel re-argue Plessy v. Ferguson, the 1896 case that blessed "separate but equal." Nine judicial philosophies, faithfully applied, upheld it 7 to 2. The real Court went 7 to 1. The two dissents in my simulation came from the judges built to ask who bears the burden โ the Civil Libertarian and the Living Constitutionalist, standing roughly where Justice Harlan stood alone. Method, honestly followed, walked straight into one of the worst decisions in American history. The objection came from the corner where empathy lives. You can listen to it happen. I am still not sure how I feel about how easy it was.
The panel earns back some respect on new ground. On June 29 the real Court decided two presidential-removal cases the same morning, in opposite directions: it let the President fire an FTC commissioner and kept a Federal Reserve governor in her seat. My panel, working blind from briefing memos โ these cases sit past the models' training cutoff โ split the pair the same way, dissents surfacing from the philosophically matching corners. Two cases prove very little. They do make you lean in.
The deal
Three predictions now sit on the record for cases the Court has not decided. When the rulings come down, the scoreboard updates in public, whether it flatters me or not. That is the deal I think anyone building with AI agents should be willing to take: put the number where people can check it, publish the baseline, and keep the receipts when you are wrong. We demand that much from our institutions. The machines will argue either way โ beautifully, it turns out. The honesty is still on us.
Watch the courtroom, convene the live bench, or subscribe to the podcast. The code is open source.
Built solo, July 4โ8 2026, on Qwen Cloud (qwen3.7-plus, qwen3.7-max, qwen3.6-flash, Voice Design, qwen3-tts, wan2.6-t2i) and Alibaba Cloud (SAS + OSS), for the Qwen Cloud Global AI Hackathon โ Agent Society track. Every number here renders from the same logs the scoreboard reads: the findings.