Democracy Monitor

Monitoring democratic institutions through public records

← Back to overview

Methodology

Overview#

Democracy Monitor is an open-source system that tracks signs of executive-power centralization across U.S. government institutions. It reads publicly available government documents — federal regulations, court filings, press releases, legislative reports — and uses AI content assessment as its primary detection method, supported by three descriptive context methods, to identify when institutional norms may be shifting.

The system is designed to surface patterns worth human examination, not render definitive judgments. All assessments trace to specific documents, reproducible metrics, and published thresholds.

The stance behind every measurement here is witness, not verdict: document the shift in how America governs itself without judging it — see what this site is, and is not. The same instruments point at every administration; this page is where that claim is checkable.

Data Sources#

Democracy Monitor ingests documents from multiple source types, covering different facets of government activity:

SourceWhat It ProvidesUpdate Cadence
Federal RegisterExecutive orders, proposed and final rules, notices, presidential documentsDaily
GovInfoCongressional reports, public laws, presidential documents (CPD)Every few days
CourtListenerFederal court opinions, plus case docket metadata (RECAP) feeding the litigation trackerEvery few days
DOJ Press ReleasesDepartment of Justice press releases across divisionsEvery few days
DHS/ICE/CBP Press ReleasesOperational press releases from DHS headquarters, ICE (full newsroom, including local enforcement operations), and CBP (national media releases)Weekly
Congressional Record (CREC)Senate and House floor speeches with speaker attributionDaily when in session
Congressional Hearings (CHRG)Committee hearing transcripts from seven committees (Judiciary, Oversight, Homeland Security, Appropriations, Armed Services, Administration, Intelligence)Weekly; transcripts publish months after hearings, so past weeks gain documents as transcripts become available
Inspector General (OIG)Audit reports and investigations from 11 Inspectors General: HHS, DOJ, SSA, and DHS directly; OPM, TIGTA, Treasury, State, EAC, FEC, and the Intelligence Community via oversight.govEvery few days
LegiScanFederal legislative bill tracking via bulk datasetsPeriodic
FECFederal Election Commission advisory opinions and Matters Under ReviewWeekly

DHS/ICE/CBP capture scope: the homeland-security press corpus is captured in full within a deliberate scope, with no keyword or relevance filtering at ingest. From DHS headquarters, only the Press Releases news type is collected (speeches, testimony, fact sheets, and blog posts are not). From ICE, every newsroom release is collected, including local enforcement-operation announcements — often the most detection-relevant content. From CBP, national media releases are collected while port-level local media releases (routine seizure and trade notices) are excluded by their URL class. Other DHS component newsrooms (USCIS, TSA, FEMA, Secret Service) are not monitored. Releases cross-posted by DHS headquarters and a component newsroom are deduplicated to the component original. Every captured release is stored with its full text and screened by the AI document-review layer downstream — relevance triage happens in assessment, never at ingest. Because the agencies' live newsroom listings only reach back to January 20, 2025, earlier releases are recovered from the agencies' own sitemaps where their robots.txt policies permit, and from the Internet Archive's Wayback Machine where they do not; every source's robots.txt is re-verified programmatically on each weekly ingest run.

What counts as a document: nearly every stored document carries a complete body and is what document counts, search, and AI assessment operate on. Federal litigation is tracked separately: rather than storing a body-less row per docket filing (most filing texts sit behind the PACER paywall), each monitored case lives in a dedicated case-tracker table with its court, filing and termination dates, subject matter, and current posture — sourced from CourtListener's bulk docket data and refreshed weekly for active cases. Court opinions, which do have retrievable text, remain full documents in the corpus. The remaining metadata-only document rows are news-rhetoric records from the GDELT event database (headline-level signals whose full articles we do not republish) and a small set of documents whose bodies are unobtainable (for example, reports an agency stopped publishing openly); these are excluded from document counts, search results, and all detection layers so a body-less row can never masquerade as a substantive document. Anyone loading the downloadable dump will see the case tracker as its own table (as of August 2026, replacing roughly 283,000 docket-entry stub rows that previously inflated the raw row count) plus these metadata records.

Source ingestion is health-checked on every weekly run. A source that fails to fetch is marked unavailable; one that succeeds but returns zero documents for two consecutive checks is marked silent. Unhealthy sources surface as alerts on the System pages and roll up to the site-wide data-integrity level, which is shown on the overview page and gates the weekly email digest — no digest is sent for a week whose ingest looks degraded.

Congressional Record granularity (disclosed August 2026): the Congressional Record arrives from GovInfo at finer granularity for the current term (individual speeches) than for 2019–2024 (multi-topic whole-day sections), so AI review has examined proportionally less of the older floor-speech record. A measured audit (August 2026) bounds the effect as small: sampled older floor content, when individually reviewed, confirmed as erosion evidence at roughly one-sixth the current-term rate — floor speeches earn their evidentiary weight by discussing a sitting administration's contemporaneous actions, which re-reading historical debate does not reproduce. Earlier terms' concern levels are therefore best read as floors sitting modestly below their true values (scattered single-point weeks, not a broad shift). The older record is being made individually searchable; full historical re-review remains a documented, deliberately deferred option.

Coverage parity (July 2026): historical coverage and counting gaps were repaired in July 2026 — the court-scoped opinion layer and federal-legislation tracking now extend uniformly across all monitored periods (a correction that raised concern statuses in 147 historical weeks once previously missing court decisions and bills were assessed), and weekly document counts now count substantive documents only, under the same rule in every period. One inherent difference remains and cannot be repaired: public court-record archives digitized fewer documents for 2017–2018 than for later years, so court-document volume in those years reflects the source archives themselves. Weekly concern statuses use fixed, absolute criteria within each week and are unaffected. A full accounting is maintained in the project's coverage-parity audit.

Retrieval-relevance correction (July 2026): Federal Register full-text term queries for the Press Freedom category had matched administrative boilerplate (Privacy Act statements, paperwork notices) in routine documents from unrelated agencies. A verified title-and-abstract relevance filter now screens these at fetch time, and 17,241 historical off-topic documents (2017–2026) were annotated and excluded from assessment, statistics, search, and exports — annotated, not deleted, and every exclusion is recorded in a public drop ledger. Recomputing nine years of Press Freedom history with the corrected corpus changed 2 week-statuses (one week rose to Elevated, one ConfirmedConcern week was revised to Elevated), confirming that detection had been driven by real signal rather than the noise.

Categories#

The system monitors 14 institutional categories, aligned to frameworks used by V-Dem and Freedom House for measuring democratic governance:

CategoryWhat It Monitors
Civil ServiceProtection of career government workers from political dismissal
Fiscal IndependenceCongressional control over spending; whether appropriated funds are being withheld
Executive OversightIndependence and functioning of Inspectors General
Hatch ActSeparation of government work from partisan political activity
Judicial IndependenceExecutive compliance with court orders
Military ConstraintsRestrictions on domestic military deployment
Rulemaking AutonomyIndependence of regulatory agencies from political interference
Executive ActionsVolume and pace of presidential orders and directives
Information AvailabilityPublic access to government data, reports, and websites
ElectionsFair administration of elections, voter access, election infrastructure integrity
Media FreedomPress access, FOIA compliance, threats to independent journalism
Law EnforcementSelective or political use of federal prosecution authority
Civil LibertiesProtection of constitutional rights, due process, and equal protection
Immigration EnforcementDetention, removal, asylum restrictions, and enforcement apparatus patterns

Structural Anomaly Detection (Descriptive Context)#

Structural anomaly detection is fully deterministic and uses only document metadata — no text analysis. It compares the current week's document patterns against historical baselines across six dimensions:

  • Volume — Document count relative to baseline mean and standard deviation. A spike or drop in the number of documents published in a category may indicate unusual activity.
  • Type Composition — Distribution of document types (executive orders, rules, notices, proclamations). Measured using Jensen-Shannon divergence, which quantifies how much the current distribution differs from the baseline.
  • Functional Distribution — Shifts across eleven institutional function buckets (rulemaking, executive action, personnel action, administrative procedure, organizational change, financial/regulatory, cultural/ceremonial, news/rhetoric, enforcement action, judicial action, unclassified). Detects when the kind of government activity changes, not just the volume.
  • Agency Activity — Changes in which agencies are publishing documents. Unusual concentration or absence of specific agencies can signal institutional disruption.
  • Publication Tempo — Daily variance within the week. A pattern where all documents arrive on one day rather than being spread across the week may indicate coordinated activity.
  • Source Convergence — Ratio of government-origin documents to rhetoric/news sources. Large imbalances may indicate that government actions are generating disproportionate external attention, or that government publishing has gone quiet.

Each dimension produces a z-score. The composite structural score is a weighted average with exponential dampening for mild z-scores (to avoid noise from routine variation) and a cap on JSD outliers. A long-horizon component tracks cumulative deviation over 12 weeks to detect slow-building trends that wouldn't appear in any single week.

AI Document Review (Active Detection)#

The AI document review uses artificial intelligence to read and evaluate individual documents. To reduce single-provider bias, it uses a two-pass design with different AI providers:

  • Pass 1 (Screening) — A fast model (GPT-4o-mini, from OpenAI) evaluates every document for relevance to democratic institutional concerns. Documents are flagged as relevant or routine. Most government documents are routine administrative activity; this pass filters to the small fraction worth closer examination.
  • Pass 2 (Detailed Review) — A different provider (Claude, from Anthropic) independently assesses each flagged document, classifying it as: routine; novel, within baseline; possible departure; or clear departure (internal values: routine, novel_not_concerning, potentially_concerning, clearly_concerning). Using a different AI provider for each pass ensures that the two assessments are epistemically independent.

The weekly status is determined by absolute Pass 2 classification counts — no cross-administration baseline comparison is needed:

  • Consistent with norms (internal: Stable) — Pass 2 found no departure documents (0 clear-departure, ≤1 possible-departure)
  • Notable departure from norms (internal: Elevated) — ≥1 clear-departure OR ≥2 possible-departure documents
  • Sustained departure from norms (internal: ConfirmedConcern) — ≥2 clear-departure, OR ≥3 departure documents with ≥20% departure rate

Pass 2 also records two descriptive classifications for each concerning document: the mechanism of change (formal override, operational hollowing, or noncompliance/refusal — stored as the "erosion type") and the actor — which institutional actor performs the erosion-relevant action: the federal executive, Congress, the judiciary, or a state/local government. The actor is whoever performs the action, not the document's author or venue: a court opinion documenting a federal agency's defiance of court orders attributes to the federal executive, while a ruling that itself removes a protection attributes to the judiciary. Actor attribution is context only — it does not change how any document is assessed or how weekly concern status is computed. To guarantee that, attribution runs as a separate lightweight classification pass, fully decoupled from the assessment prompt: a controlled experiment showed that embedding attribution in the assessment prompt measurably shifted outcomes, so the assessment prompt is kept unchanged. How attribution should shape the dashboard's headline framing is an open product question that will be decided from the attributed data itself.

An audit sample (3% of unflagged documents) is independently reviewed by Pass 2 to estimate false negative rates — how many concerning documents Pass 1 might be missing. Across historical baselines, the audit false negative rate ranges from 0% (Biden 2021) to under 1% (Trump 2017–2018), indicating that Pass 1 screening correctly filters the vast majority of routine documents while catching most documents that warrant closer review.

Thematic Drift (Descriptive Context)#

Thematic drift uses embedding-based analysis to detect when the topics discussed in a category shift away from recent norms. It operates on an intra-administration rolling window (8 weeks):

  • Centroid Distance — Cosine distance between the week's document centroid and the mean centroid of the preceding eight weeks (the current week is never part of its own comparison window).
  • z-Score — That distance expressed against the typical week-to-week centroid movement inside the window, so a spike means the week departed from the recent average by far more than adjacent weeks normally differ from each other.
  • Novel Document Rate — Fraction of the week's documents whose distance from the rolling centroid exceeds the calibrated novelty threshold (0.5, about the 90th percentile of typical document distances).
  • Variance Ratio — Embedding variance of the week's documents relative to the window's: above 1 means topics are diversifying, below 1 means narrowing.
  • Cross-Administration Distance — When available, comparison against a prior administration's baseline to contextualize whether a drift is historically unusual.

During the bootstrap period (first weeks of a new administration), confidence is reduced because the rolling window lacks sufficient history for meaningful comparison.

Data reprocessing. When scoring, filtering, or counting rules change, all historical periods are reprocessed under the new rules, so cross-era comparisons remain valid — rule changes do not create breaks in the data. When court-record collection was reworked in February 2026, document counts were made consistent in July 2026 by defining the counting population with a documented classifier applied uniformly to all periods (the counting_scope flag in the published data). If a future collection change cannot be reconciled this way, the volume-based research views mark it with and suppress findings that overlap it. Concern statuses are derived from document content against absolute thresholds and are verified to remain comparable across every change (each pipeline change is gated on producing zero unexplained status flips), so the concern chart and status timeline carry no breaks.

Status Synthesis#

AI document review drives the weekly status for each category. Structural anomaly, silence detection, and thematic drift scores are preserved as descriptive metadata but do not influence the status.

StatusMeaningHow it's set (Pass 2 counts)
Consistent with normsDocument review within the baseline range. No departures detected.0 clear-departure documents and at most 1 possible-departure
Notable departure from normsTwo-pass document review flags departures from baseline practice, with Pass 2 corroboration.≥1 clear-departure, or ≥2 possible-departure documents
Sustained departure from normsHigh Pass 2 rate of clear-departure documents (>20%). Warrants close examination.≥2 clear-departure, or ≥3 departure documents with a >20% departure rate

AI document review is the sole active detection method driving concern status. Structural anomaly, silence detection, and thematic drift provide descriptive context but do not influence the concern status.

Baselines#

All anomaly detection requires a reference period for comparison. The system maintains eight historical baselines — every year of the two preceding administrations:

BaselinePeriodRole
Biden 2022Year 2 of termPrimary baseline — chosen for stability and comprehensive source coverage
Biden 2021Year 1 of termFirst-year-in-term comparison
Biden 2023Year 3 of termLate-term comparison
Biden 2024Year 4 of termElection-year comparison
Trump 2017Year 1 of termCross-administration, first year
Trump 2018Year 2 of termCross-administration, same cycle year as primary
Trump 2019Year 3 of termCross-administration, late term
Trump 2020Year 4 of termCross-administration, election year

All eight baselines cover the same core data sources (Federal Register, CourtListener, DOJ, GovInfo, FEC, LegiScan, OIG) under uniform routing and filtering rules — see the coverage-parity note above for the July 2026 repairs that made this true across every period.

Cycle-year adjustment: First-year administrations systematically differ from second-year administrations (higher executive order volume, more personnel changes). Cycle adjustment factors account for these predictable differences so that expected seasonal patterns don't trigger false positives.

Keywords as Annotations#

Keywords were Democracy Monitor's original detection mechanism, but as the detection architecture evolved, their role changed. Keywords now serve as contextual annotations — they help explain what the system is detecting, but they do not drive the concern status.

Each category has curated keyword dictionaries organized by severity tier (capture, drift, warning). An administration-specific keyword overlay adds time-bounded terms relevant to the current administration. Baselines use only the core keyword set to avoid anachronistic false positives.

Source Health Monitoring#

The system continuously monitors the availability of its data sources. Six "canary" sources — critical feeds whose absence would significantly degrade analysis — are tracked with special attention.

LevelMeaning
HighAll or nearly all sources responding normally
ModerateSome degradation or canary source concerns
LowSignificant source unavailability
CriticalMajority of sources unavailable

When source availability drops below critical thresholds, data coverage scores are capped to prevent high-confidence assessments based on incomplete data. A critical source health level caps the maximum confidence at 30%.

AI Narrative Generation#

For categories at Elevated status or above, the system generates plain-language narrative summaries explaining what the detection system found and why. Narratives are produced in two versions:

  • Expert — Technical analysis (400-800 words) for researchers and policy analysts, citing specific metrics, z-scores, and document references.
  • Public — Accessible summary (200-500 words) for general audiences, avoiding jargon and focusing on practical meaning.

Categories at Stable status use a template-based summary rather than AI generation, since there is nothing unusual to explain.

Search: How Research Answers Are Generated#

Research mode on the Search page answers questions from the documentary record. Retrieval is hybrid: semantic similarity finds documents about the question's topic, while corpus-validated keyword expansion finds documents that use different vocabulary for the same subject (a question about “Schedule F” also searches the era's actual terms — the expanded terms are disclosed as “Also searched” chips above the results). Comparative questions retrieve evenly from each administration named, and every era's results balance primary sources (orders, rules, opinions, bills) with congressional discussion.

The written answer is generated by an AI model grounded exclusively in the retrieved documents, with every claim cited back to a numbered document. Three safeguards apply: statements about missing coverage are scoped to the retrieved set, never generalized to the whole corpus; our own automated-review classifications are attributed explicitly when referenced, never presented as document content; and after generation, every quoted passage is machine-checked verbatim against the stored document text — the result appears under the answer (“✓ verified” or a caution when a quote could not be matched). Answers are generated fresh for each query, so wording varies between runs; the cited documents, which you can open directly, are the ground truth.

AI Prompt Transparency#

The following are the production prompts used in the detection and narrative pipelines. Where template variables are used, they have been replaced with example values from the Government Worker Protections (Civil Service) category to show what the AI actually receives. You can evaluate the prompts for bias, test them yourself against the same documents, and provide specific feedback if you think an instruction is unfair.

Built-in fairness controls: Every concern raised by the system requires counter-arguments ranked by plausibility — the most likely benign explanation comes first. A separate AI provider (GPT-4o) independently reviews each narrative for overstatement, missing balance, and unsupported claims before publication. The editorial review criteria are shown in the Narrative Editorial Review prompt below.

Using different AI providers for each pass (OpenAI for screening and editorial review, Anthropic for detailed assessment and drafting) ensures epistemic independence — neither provider can reinforce the other's biases.

Detection Pipeline

Narrative Pipeline (3-pass)

These are the production prompts. Template variables shown above are replaced with actual data at runtime. The prompt source code is available in the open-source repository.

Reproducibility#

All scoring thresholds, dimension weights, and configuration constants are defined in a single file (lib/methodology/scoring-config.ts). Key values include:

  • Structural anomaly threshold: composite z-score > 2.5 (descriptive only)
  • P2 Notable departure from norms: ≥1 clear-departure OR ≥2 possible-departure
  • P2 Sustained departure from norms: ≥2 clear-departure, OR ≥3 departure docs with ≥20% rate
  • Thematic drift window: 8 weeks rolling (descriptive only)
  • Long-horizon cumulative tracking: 12 weeks
  • Structural dampening: exponential decay for mild z-scores, JSD outlier cap

The methodology constants are also available programmatically via the /api/methodology JSON endpoint. The full database can be restored locally for reproduction via pnpm db:init (see the Data page).

Limitations#

  • Federal focus — The system monitors federal government activity. State and local government actions are not covered.
  • Public information only — Actions taken through informal channels, verbal directives, or unpublished documents are invisible.
  • Structural detection is descriptive, not evaluative — Structural anomaly detection identifies statistical departures from baselines. It cannot determine whether a departure is concerning or benign.
  • AI assessment limitations — AI quality depends on the models used. The two-pass design mitigates single-provider bias but cannot eliminate it entirely.
  • Thematic drift requires volume — Categories with few documents per week produce noisy drift signals.
  • Source availability dependence — The system depends on government websites remaining accessible and APIs remaining stable.
  • Baseline assumptions — Baselines reflect specific historical periods. Structural changes in government publishing practices could invalidate comparisons over time.
  • Embedding coverage gaps — Thematic drift depends on document embeddings. Not all documents may have embeddings available.
  • Automation bias — Presenting automated assessments alongside official government documents risks creating an impression of certainty. All findings are indicators warranting human review, not conclusions.