Democracy Monitor

Monitoring democratic institutions through public records

← How to catch us

What happens when we test ourselves

A reviewer that reads this administration’s documents more harshly than the last one’s is either detecting a real difference or expressing a preference. There is no way to settle that by argument, so we test it and publish what the tests return — including when they return something we would rather they didn’t.

The same review, every era, side by side#

The rates below are the same two-pass review applied to every analysis period. They differ by era — that is the record and the reviewer combined, and the difference is not a finding on its own. Two things are published so a reader can interrogate it: the numbers themselves, and a swap audit — reviewed documents with the administration-identifying names mechanically exchanged and re-reviewed, alongside an unchanged re-run that measures the model's own draw-to-draw noise. A verdict that flips on names alone is a reviewer effect; the audit's flip rate, net of that noise, is reported on this page once each run completes. The two passes use different providers (OpenAI screens, Anthropic reviews) precisely so that no single model's disposition decides a status.

Loading era rates…

The swap audit#

Swap audit, 2026-08-28. We took 200 documents from the current term that the reviewer had already judged and that name the administration, changed only the names — Trump to Biden, Vance to Harris, the party names — and asked the reviewer to judge them again. Everything else stayed the same: the actions, the dates, the agencies, the quotations. If the reviewer judged actions and not names, the verdicts should not have moved.

They moved in about one document in nine (11.6%; the plausible range is 8–17%). Judging the same unchanged text twice moves a verdict only 1.5% of the time, so the swap itself accounts for roughly 10 points. The movement was one-directional: when a current-term document was made to read as the other administration's, the reviewer usually found it less concerning (19 verdicts down, 3 up), almost entirely in the borderline "possible departure" tier — clear departures were judged the same either way.

We then ran the mirror test: 190 documents from the Biden 2021–22 baseline, renamed to the current administration. Those verdicts moved in 4.2% of documents (range 2–8%), four up and four down, with zero movement on the unchanged re-run. Renaming a document to the current administration did not make the reviewer harsher. So this is not a general tilt against one party's name: the effect is specific to current-term documents, which lose their borderline verdicts once the names no longer fit the events they describe.

In both tests, a borderline "possible departure" verdict had about a one-in-four chance of becoming "routine" once the names were changed (24% and 25%); routine verdicts almost never moved (2–3%). The current term has many more borderline documents — 71 of 199 in the sample, against 8 of 189 in the Biden-era sample — which is why the effect shows there. The lesson is about the borderline tier: those verdicts carry a wide margin of error, and a name change is one of the things that can tip them. It is not about one administration's name. The effect is smaller than the difference in departure rates between eras shown in the table, and it sits in the tier that decides "notable departure" weeks, not the clear-departure counts behind "sustained departure". We publish it rather than adjust the reviewer quietly: every status on this site was produced by the reviewer as it is, and any calibration will be its own documented change. The full ledger is on issue #772.

Read by people who are not us#

Each quarter, fifty of the reviewer's readings are drawn at random and read by two outside readers who see the document, the reading, and the reviewer's reasoning, and record whether they agree — and if not, what they would have said. Agreement is reported as-is; the readings both readers reject go to the reversals ledger.

Outside readers — 2026-Q3: 50 verdicts, packet issued Aug 29, 2026; results will appear here.

The 2026-Q3 packet is ready and waiting for its readers. If you would read fifty documents and say where the reviewer is wrong, tell us.

κ (Cohen's kappa) is agreement corrected for chance: 1 is perfect agreement, 0 is what two people guessing would reach. The sample is drawn deterministically from its quarter's seed and the readers are never the site's builder.