Skip to content
CIVIC HERALD

Methodology & trust · working spec v1

How we work, and how we show it.

Trust is the whole game. If we can't point to where a claim came from, we don't make the claim. Here is exactly how the data gets collected, summarized, scored, and checked — and the rules that keep it honest.

Where the data comes from

We start with the government's own records. Everything on Civic Herald is built from primary, public-domain federal sources — not from a pundit, a party, or a press release. When a number is on the page, you can follow it back to the office that published it.

  • Congress.gov — bill metadata, status timeline, official CRS summaries, sponsors, and subjects.
  • GovInfo — full bill text in structured XML, the input to the plain-language pipeline.
  • House Clerk · Senate LIS — per-legislator roll-call votes, normalized across both chambers.
  • FEC OpenFEC — campaign finance: the "who funds them" data.

State and local sources follow the same discipline — official first, with a citation attached — as coverage expands.

How the pipeline works

Automated labor, prioritized by what people need. A bill becomes law about one time in twenty, so we don't summarize all ten thousand of them with equal effort. The pipeline ranks work by activity and attention, drafts a plain-language briefing of what a bill does and who it touches, and tags each provision against a fixed issue taxonomy.

The model does the toil — reading dense legislative text and proposing structured outputs. It does not get the last word. Until a person has reviewed a claim, it is shown as provisional, clearly labeled, never presented as settled fact.

Derived vs. fetched

A model's guess never masquerades as a fact. We keep a hard line between two kinds of information: data we fetched from an official source, and analysis a model derived from it. Both can appear side by side, but they never wear the same clothes.

We never let a model's estimate sit on the page looking like an official number.

Fetched data carries its source and the time it was retrieved. Derived data carries the run that produced it and its review status. A cost figure from the Congressional Budget Office is labeled as such; a modeled estimate is labeled as a modeled estimate.

How alignment is scored

We measure against your values — not a house line. Onboarding asks where you stand on the issues you care about. A bill's alignment with you is computed at read time from your stated values and the bill's per-issue effects — and it is never stored on the bill, because the same bill aligns differently with different people.

The method is symmetric: the exact same scoring runs for every user and every politician, with no special handling for either side. We keep relevance ("is this on your radar?") separate from alignment ("does it match your values?") — two different questions that deserve two different answers.

Human review & balance

The pipeline does the toil; people guard the integrity. Trained reviewers audit the data itself — confirming or correcting plain-language summaries, issue tags, effect directions, and divergence flags. Accusatory features stay gated behind that review: a machine's provisional judgment about a named person is not enough to publish.

Reviewers are recruited in equal number from the left, the center, and the right. Each rates their own political lean before they serve, and we publish it. They review blind to a claim's source or sponsor, and scores are averaged across the groups so that no single side sets a number.

Provenance on everything

Every claim carries its source, version, and freshness. Provenance isn't a footnote here; it's a load-bearing part of the page.

  • Example — "Bill text: GovInfo XML (38pp) · summarized by model v.2026.05 · last refresh 6h ago."
  • Citations — summaries point back into the bill's own sections, not vague paraphrases.
  • Review stamp — "Reviewed by a balanced panel · 3 days ago," shown only when it's true.

If a part of a claim is missing — no source, no review — the page says so rather than papering over it.

Neutrality governance

Neutrality is structural, not a slogan. A scorecard is only useful if people who disagree with each other can both trust the math. So the neutrality is built into the process, not asserted in a mission statement.

A scorecard is only trustworthy if a conservative and a progressive both believe the numbers weren't cooked for the other side.

We blind extraction to party, hold to a strict non-advocacy charter, and borrow our safeguards — balanced cohorts, public donor disclosure, an editorial firewall — from the institutions that set the standard for nonpartisan civic data.

Privacy & your data

Your values, your activity, your profile — yours. Your stated positions power your scores and nothing else. You can download your data or delete it, and personalized scores are computed fresh each time you read, never sold and never used to advocate.

We'd rather show less and be unimpeachable than show more and be doubted. This spec is a living document — it evolves as our methods are reviewed and externally audited. Found an error, or want to scrutinize the math? That's the point. See who funds us or become a reviewer.

Living document

Scrutinize the math.

This spec evolves as our methods are reviewed and externally audited. Found an error, or want to check our work? That's the point.