Monitoring democratic institutions through public records
Democracy Monitor is an open-source system that tracks signs of executive-power centralization across U.S. government institutions. It reads publicly available government documents — federal regulations, court filings, press releases, legislative reports — and uses AI content assessment as its primary detection method, supported by three descriptive context methods, to identify when institutional norms may be shifting.
The system is designed to surface patterns worth human examination, not render definitive judgments. All assessments trace to specific documents, reproducible metrics, and published thresholds.
The stance behind every measurement here is witness, not verdict: document the shift in how America governs itself without judging it — see what this site is, and is not. The same instruments point at every administration; this page is where that claim is checkable.
Democracy Monitor ingests documents from multiple source types, covering different facets of government activity:
| Source | What It Provides | Update Cadence |
|---|---|---|
| Federal Register | Executive orders, proposed and final rules, notices, presidential documents | Daily |
| GovInfo | Congressional reports, public laws, presidential documents (CPD) | Every few days |
| CourtListener | Federal court opinions, plus case docket metadata (RECAP) feeding the litigation tracker | Every few days |
| DOJ Press Releases | Department of Justice press releases across divisions | Every few days |
| DHS/ICE/CBP Press Releases | Operational press releases from DHS headquarters, ICE (full newsroom, including local enforcement operations), and CBP (national media releases) | Weekly |
| Congressional Record (CREC) | Senate and House floor speeches with speaker attribution | Daily when in session |
| Congressional Hearings (CHRG) | Committee hearing transcripts from seven committees (Judiciary, Oversight, Homeland Security, Appropriations, Armed Services, Administration, Intelligence) | Weekly; transcripts publish months after hearings, so past weeks gain documents as transcripts become available |
| Inspector General (OIG) | Audit reports and investigations from 11 Inspectors General: HHS, DOJ, SSA, and DHS directly; OPM, TIGTA, Treasury, State, EAC, FEC, and the Intelligence Community via oversight.gov | Every few days |
| LegiScan | Federal legislative bill tracking via bulk datasets | Periodic |
| FEC | Federal Election Commission advisory opinions and Matters Under Review | Weekly |
DHS/ICE/CBP capture scope: the homeland-security press corpus is captured in full within a deliberate scope, with no keyword or relevance filtering at ingest. From DHS headquarters, only the Press Releases news type is collected (speeches, testimony, fact sheets, and blog posts are not). From ICE, every newsroom release is collected, including local enforcement-operation announcements — often the most detection-relevant content. From CBP, national media releases are collected while port-level local media releases (routine seizure and trade notices) are excluded by their URL class. Other DHS component newsrooms (USCIS, TSA, FEMA, Secret Service) are not monitored. Releases cross-posted by DHS headquarters and a component newsroom are deduplicated to the component original. Every captured release is stored with its full text and screened by the AI document-review layer downstream — relevance triage happens in assessment, never at ingest. Because the agencies' live newsroom listings only reach back to January 20, 2025, earlier releases are recovered from the agencies' own sitemaps where their robots.txt policies permit, and from the Internet Archive's Wayback Machine where they do not; every source's robots.txt is re-verified programmatically on each weekly ingest run.
What counts as a document: nearly every stored document carries a complete body and is what document counts, search, and AI assessment operate on. Federal litigation is tracked separately: rather than storing a body-less row per docket filing (most filing texts sit behind the PACER paywall), each monitored case lives in a dedicated case-tracker table with its court, filing and termination dates, subject matter, and current posture — sourced from CourtListener's bulk docket data and refreshed weekly for active cases. Court opinions, which do have retrievable text, remain full documents in the corpus. The remaining metadata-only document rows are news-rhetoric records from the GDELT event database (headline-level signals whose full articles we do not republish) and a small set of documents whose bodies are unobtainable (for example, reports an agency stopped publishing openly); these are excluded from document counts, search results, and all detection layers so a body-less row can never masquerade as a substantive document. Anyone loading the downloadable dump will see the case tracker as its own table (as of August 2026, replacing roughly 283,000 docket-entry stub rows that previously inflated the raw row count) plus these metadata records.
Source ingestion is health-checked on every weekly run. A source that fails to fetch is marked unavailable; one that succeeds but returns zero documents for two consecutive checks is marked silent. Unhealthy sources surface as alerts on the System pages and roll up to the site-wide data-integrity level, which is shown on the overview page and gates the weekly email digest — no digest is sent for a week whose ingest looks degraded.
Congressional Record granularity (disclosed August 2026): the Congressional Record arrives from GovInfo at finer granularity for the current term (individual speeches) than for 2019–2024 (multi-topic whole-day sections), so AI review has examined proportionally less of the older floor-speech record. A measured audit (August 2026) bounds the effect as small: sampled older floor content, when individually reviewed, confirmed as erosion evidence at roughly one-sixth the current-term rate — floor speeches earn their evidentiary weight by discussing a sitting administration's contemporaneous actions, which re-reading historical debate does not reproduce. Earlier terms' concern levels are therefore best read as floors sitting modestly below their true values (scattered single-point weeks, not a broad shift). The older record is being made individually searchable; full historical re-review remains a documented, deliberately deferred option.
Coverage parity (July 2026): historical coverage and counting gaps were repaired in July 2026 — the court-scoped opinion layer and federal-legislation tracking now extend uniformly across all monitored periods (a correction that raised concern statuses in 147 historical weeks once previously missing court decisions and bills were assessed), and weekly document counts now count substantive documents only, under the same rule in every period. One inherent difference remains and cannot be repaired: public court-record archives digitized fewer documents for 2017–2018 than for later years, so court-document volume in those years reflects the source archives themselves. Weekly concern statuses use fixed, absolute criteria within each week and are unaffected. A full accounting is maintained in the project's coverage-parity audit.
Retrieval-relevance correction (July 2026): Federal Register full-text term queries for the Press Freedom category had matched administrative boilerplate (Privacy Act statements, paperwork notices) in routine documents from unrelated agencies. A verified title-and-abstract relevance filter now screens these at fetch time, and 17,241 historical off-topic documents (2017–2026) were annotated and excluded from assessment, statistics, search, and exports — annotated, not deleted, and every exclusion is recorded in a public drop ledger. Recomputing nine years of Press Freedom history with the corrected corpus changed 2 week-statuses (one week rose to Elevated, one ConfirmedConcern week was revised to Elevated), confirming that detection had been driven by real signal rather than the noise.
The system monitors 14 institutional categories, aligned to frameworks used by V-Dem and Freedom House for measuring democratic governance:
| Category | What It Monitors |
|---|---|
| Civil Service | Protection of career government workers from political dismissal |
| Fiscal Independence | Congressional control over spending; whether appropriated funds are being withheld |
| Executive Oversight | Independence and functioning of Inspectors General |
| Hatch Act | Separation of government work from partisan political activity |
| Judicial Independence | Executive compliance with court orders |
| Military Constraints | Restrictions on domestic military deployment |
| Rulemaking Autonomy | Independence of regulatory agencies from political interference |
| Executive Actions | Volume and pace of presidential orders and directives |
| Information Availability | Public access to government data, reports, and websites |
| Elections | Fair administration of elections, voter access, election infrastructure integrity |
| Media Freedom | Press access, FOIA compliance, threats to independent journalism |
| Law Enforcement | Selective or political use of federal prosecution authority |
| Civil Liberties | Protection of constitutional rights, due process, and equal protection |
| Immigration Enforcement | Detention, removal, asylum restrictions, and enforcement apparatus patterns |
Structural anomaly detection is fully deterministic and uses only document metadata — no text analysis. It compares the current week's document patterns against historical baselines across six dimensions:
Each dimension produces a z-score. The composite structural score is a weighted average with exponential dampening for mild z-scores (to avoid noise from routine variation) and a cap on JSD outliers. A long-horizon component tracks cumulative deviation over 12 weeks to detect slow-building trends that wouldn't appear in any single week.
The AI document review uses artificial intelligence to read and evaluate individual documents. To reduce single-provider bias, it uses a two-pass design with different AI providers:
The weekly status is determined by absolute Pass 2 classification counts — no cross-administration baseline comparison is needed:
Pass 2 also records two descriptive classifications for each concerning document: the mechanism of change (formal override, operational hollowing, or noncompliance/refusal — stored as the "erosion type") and the actor — which institutional actor performs the erosion-relevant action: the federal executive, Congress, the judiciary, or a state/local government. The actor is whoever performs the action, not the document's author or venue: a court opinion documenting a federal agency's defiance of court orders attributes to the federal executive, while a ruling that itself removes a protection attributes to the judiciary. Actor attribution is context only — it does not change how any document is assessed or how weekly concern status is computed. To guarantee that, attribution runs as a separate lightweight classification pass, fully decoupled from the assessment prompt: a controlled experiment showed that embedding attribution in the assessment prompt measurably shifted outcomes, so the assessment prompt is kept unchanged. How attribution should shape the dashboard's headline framing is an open product question that will be decided from the attributed data itself.
An audit sample (3% of unflagged documents) is independently reviewed by Pass 2 to estimate false negative rates — how many concerning documents Pass 1 might be missing. Across historical baselines, the audit false negative rate ranges from 0% (Biden 2021) to under 1% (Trump 2017–2018), indicating that Pass 1 screening correctly filters the vast majority of routine documents while catching most documents that warrant closer review.
Thematic drift uses embedding-based analysis to detect when the topics discussed in a category shift away from recent norms. It operates on an intra-administration rolling window (8 weeks):
During the bootstrap period (first weeks of a new administration), confidence is reduced because the rolling window lacks sufficient history for meaningful comparison.
Data reprocessing. When scoring, filtering, or counting rules change, all historical periods are reprocessed under the new rules, so cross-era comparisons remain valid — rule changes do not create breaks in the data. When court-record collection was reworked in February 2026, document counts were made consistent in July 2026 by defining the counting population with a documented classifier applied uniformly to all periods (the counting_scope flag in the published data). If a future collection change cannot be reconciled this way, the volume-based research views mark it with ▲ and suppress findings that overlap it. Concern statuses are derived from document content against absolute thresholds and are verified to remain comparable across every change (each pipeline change is gated on producing zero unexplained status flips), so the concern chart and status timeline carry no breaks.
AI document review drives the weekly status for each category. Structural anomaly, silence detection, and thematic drift scores are preserved as descriptive metadata but do not influence the status.
| Status | Meaning | How it's set (Pass 2 counts) |
|---|---|---|
| Consistent with norms | Document review within the baseline range. No departures detected. | 0 clear-departure documents and at most 1 possible-departure |
| Notable departure from norms | Two-pass document review flags departures from baseline practice, with Pass 2 corroboration. | ≥1 clear-departure, or ≥2 possible-departure documents |
| Sustained departure from norms | High Pass 2 rate of clear-departure documents (>20%). Warrants close examination. | ≥2 clear-departure, or ≥3 departure documents with a >20% departure rate |
AI document review is the sole active detection method driving concern status. Structural anomaly, silence detection, and thematic drift provide descriptive context but do not influence the concern status.
All anomaly detection requires a reference period for comparison. The system maintains eight historical baselines — every year of the two preceding administrations:
| Baseline | Period | Role |
|---|---|---|
| Biden 2022 | Year 2 of term | Primary baseline — chosen for stability and comprehensive source coverage |
| Biden 2021 | Year 1 of term | First-year-in-term comparison |
| Biden 2023 | Year 3 of term | Late-term comparison |
| Biden 2024 | Year 4 of term | Election-year comparison |
| Trump 2017 | Year 1 of term | Cross-administration, first year |
| Trump 2018 | Year 2 of term | Cross-administration, same cycle year as primary |
| Trump 2019 | Year 3 of term | Cross-administration, late term |
| Trump 2020 | Year 4 of term | Cross-administration, election year |
All eight baselines cover the same core data sources (Federal Register, CourtListener, DOJ, GovInfo, FEC, LegiScan, OIG) under uniform routing and filtering rules — see the coverage-parity note above for the July 2026 repairs that made this true across every period.
Cycle-year adjustment: First-year administrations systematically differ from second-year administrations (higher executive order volume, more personnel changes). Cycle adjustment factors account for these predictable differences so that expected seasonal patterns don't trigger false positives.
Keywords were Democracy Monitor's original detection mechanism, but as the detection architecture evolved, their role changed. Keywords now serve as contextual annotations — they help explain what the system is detecting, but they do not drive the concern status.
Each category has curated keyword dictionaries organized by severity tier (capture, drift, warning). An administration-specific keyword overlay adds time-bounded terms relevant to the current administration. Baselines use only the core keyword set to avoid anachronistic false positives.
The system continuously monitors the availability of its data sources. Six "canary" sources — critical feeds whose absence would significantly degrade analysis — are tracked with special attention.
| Level | Meaning |
|---|---|
| High | All or nearly all sources responding normally |
| Moderate | Some degradation or canary source concerns |
| Low | Significant source unavailability |
| Critical | Majority of sources unavailable |
When source availability drops below critical thresholds, data coverage scores are capped to prevent high-confidence assessments based on incomplete data. A critical source health level caps the maximum confidence at 30%.
For categories at Elevated status or above, the system generates plain-language narrative summaries explaining what the detection system found and why. Narratives are produced in two versions:
Categories at Stable status use a template-based summary rather than AI generation, since there is nothing unusual to explain.
Research mode on the Search page answers questions from the documentary record. Retrieval is hybrid: semantic similarity finds documents about the question's topic, while corpus-validated keyword expansion finds documents that use different vocabulary for the same subject (a question about “Schedule F” also searches the era's actual terms — the expanded terms are disclosed as “Also searched” chips above the results). Comparative questions retrieve evenly from each administration named, and every era's results balance primary sources (orders, rules, opinions, bills) with congressional discussion.
The written answer is generated by an AI model grounded exclusively in the retrieved documents, with every claim cited back to a numbered document. Three safeguards apply: statements about missing coverage are scoped to the retrieved set, never generalized to the whole corpus; our own automated-review classifications are attributed explicitly when referenced, never presented as document content; and after generation, every quoted passage is machine-checked verbatim against the stored document text — the result appears under the answer (“✓ verified” or a caution when a quote could not be matched). Answers are generated fresh for each query, so wording varies between runs; the cited documents, which you can open directly, are the ground truth.
The following are the production prompts used in the detection and narrative pipelines. Where template variables are used, they have been replaced with example values from the Government Worker Protections (Civil Service) category to show what the AI actually receives. You can evaluate the prompts for bias, test them yourself against the same documents, and provide specific feedback if you think an instruction is unfair.
Built-in fairness controls: Every concern raised by the system requires counter-arguments ranked by plausibility — the most likely benign explanation comes first. A separate AI provider (GPT-4o) independently reviews each narrative for overstatement, missing balance, and unsupported claims before publication. The editorial review criteria are shown in the Narrative Editorial Review prompt below.
Using different AI providers for each pass (OpenAI for screening and editorial review, Anthropic for detailed assessment and drafting) ensures epistemic independence — neither provider can reinforce the other's biases.
These are the production prompts. Template variables shown above are replaced with actual data at runtime. The prompt source code is available in the open-source repository.
All scoring thresholds, dimension weights, and configuration constants are defined in a single file (lib/methodology/scoring-config.ts). Key values include:
The methodology constants are also available programmatically via the /api/methodology JSON endpoint. The full database can be restored locally for reproduction via pnpm db:init (see the Data page).