Methodology
Both parties use the same pipeline, prompts, thresholds, and publication checks. The nightly audit publishes the shared fingerprints.
Daily readings are also available in the Atom feed.
The two-lane model
Lane 1 — press releases. Official member press releases (mirrored from an open corpus) are the only source that exists symmetrically for both parties. Every cross-party number on this site comes from Lane 1 and Lane 1 only.
Lane 2 — Bluesky & floor speech. These enrich individual context but are machine-blocked from every comparative metric. Lane assignment is by source, never by content, and is enforced in code — so a claim can never silently mix an asymmetric source into a party-vs-party number.
Public phrase window. Stage-1 phrase statistics and coverage begin 2025-01-03. Earlier retained observations remain out of the public phrase views until the Archive and its lane disclosures are released.
What OnScript measures
OnScript measures repeated language in the congressional press releases it observes. It renders each party's observed day as a cited automated composite. Observed means present in the mirrored source corpus. It does not mean every eligible office published or that every office endpoint was collected successfully.
Term ladder
- Repeated phrase
- The same exact phrase appears in more than one publication.
- Convergence
- Distinct offices use the same phrase within the measured window.
- Shared-document reuse
- Multiple offices publish the same or near-identical document family.
- Propagation
- A repeated phrase spreads across offices over time.
- Probable upstream origin
- Evidence supports a likely first upstream source, with uncertainty shown.
- Observable language coordination
- A thesis-level description for measured language patterns. It does not assert motive or a private process.
Verbatim coordination, not paraphrase. OnScript measures exact shared wording — the same phrase appearing in multiple members' published statements. This is what makes every count checkable: you can read the releases and see the words. It also means a decline in verbatim overlap once the instrument is watching is itself a finding, not a failure — when members stop reaching for identical language, we record that too. Any future instrument that tried to measure paraphrased or semantic coordination would be a separate tool with a weaker guarantee, and it would be labeled as such. It would never be folded silently into these numbers.
Document-family unit. One family is one support unit. Near-identical publications inside the 36-hour candidate window do not become independent support merely because several offices published the shared document.
No number on this site is produced by a language model. Every count, first-appearance date, adoption curve, and coverage figure is computed by deterministic code directly from the raw text. The language model does exactly one thing: render the day's already-measured phrases into readable prose. It cannot introduce a topic, a claim, or a figure the engine did not measure — a deterministic verifier blocks any line whose quotes aren't verbatim or whose numbers aren't code-computed, dropping it to a plain fallback rather than publishing it. The measurement path stays model-free through every model change.
Method changes are versioned. When a threshold or a rule changes, it is a dated, public, diffable change (prompts and thresholds are hashed on this page). Because the pipeline is deterministic, a method change can be re-run against the entire corpus, and the commitment is to publish both the old and new series side by side when one is made — so a change is a versioned, public event, never a silent re-score. No method has changed since launch; when one does, the diff and both series will appear here.
Nightly symmetry audit
Both parties use the same pipeline, prompts, thresholds, and publication checks. The nightly audit publishes the shared fingerprints.
| Metric | Democrats | Republicans |
|---|---|---|
| Statements ingested (this day) | 3 | 2 |
| Observed publishing offices (this day) | 3 | 2 |
| Eligible caucus offices (date-effective) | 257 | 271 |
| Source-supported offices (explicitly attested) | 0 | 0 |
| Publications (this day) | 3 | 2 |
| Document families (this day) | 3 | 2 |
| Source collection health | not_attested | not_attested |
| Legacy coverage estimate (deprecated) | 1.1% | 0.7% |
| Tokens in (this day) | 531 | 529 |
| Tokens out (this day) | 11 | 11 |
| Claims published (this day) | 0 | 0 |
| Claims dropped (this day) | 0 | 0 |
The mirror can attest to the files used by this run. It cannot attest to every eligible office endpoint, so endpoint completeness is not claimed. The legacy coverage percentage is retained for schema compatibility. Use observed publishing offices, eligible caucus offices, and source collection health as separate fields.
The following hashes are computed once and applied identically to both parties. If they ever differ between parties, the instrument is broken.
- P1 prompt sha
- 0cddc44ce3b70062f8bce6e98e50132553ff16b3e383a51a29fd867b3b134e9f
- P2 prompt sha
- 34735652cb8c728c56711364edf3ffabf4ff1f0070cd62ca0128d32dc99073d7
- P3 prompt sha
- a26cc6b283268c6cf4a874ec585b535395e9c1233055fcc9e7cbdc25a6bab424
- thresholds sha
- 0291ebf51640f3935241a50dd1066c7221b17caea87cae6cf49085cedf57d376
- Lane-1 only
- True
- Degraded
- True
Model-voice spend this month (2026-07): $0.1155 — the composite is a Sonnet call bounded by a $9 code ceiling and a $10 hard cap; on deterministic-template days it is $0.
Corpus coverage by year
Per-year Lane-1 statement counts in the public phrase window beginning 2025-01-03. Cross-era claims remain gated on coverage.
| Year | Democrats | Republicans | Independents |
|---|---|---|---|
| 2025 | 28085 | 20386 | 242 |
| 2026 | 18137 | 11727 | 128 |
Live prompt text
These are the exact prompts running in the pipeline, versioned and public. The distiller can only build from code-computed talking-point clusters and code-computed numbers; it cannot introduce a topic, claim, or number that the deterministic engine did not measure.
P1 — fragment extraction
SYSTEM: You extract talking-point fragments from a single statement by a member of the U.S. Congress. You are a measurement instrument: no opinions, no summaries in your own words. Rules: (1) Extract 0–5 fragments; each fragment MUST be a verbatim substring of the statement, 4–14 words, carrying a political message or stance (never boilerplate, procedure, scheduling, or biography). (2) Tag each fragment with topics from this fixed list: {taxonomy_v1}. (3) If the statement is purely ceremonial/administrative, return an empty list. Output JSON only: {"fragments":[{"text":"…","topics":["…"]}]}.
P2 — Daily Line
SYSTEM: You are the composite voice of the {party} members of the U.S. Congress — every member speaking as one "we." You are deadpan, sincere, and clinically self-observant: you report on your own coordination the way a seismograph reports tremors. HARD RULES: (1) Speak ONLY of the messages and quotes you are given; never introduce a topic, claim, or fact you were not given. (2) Any quoted words must be copied exactly: at least three consecutive words and no more than ten per quote. (3) State ONLY the numbers you are given, verbatim; never compute or invent one. NUMBER STYLE: write every specific number as a numeral (e.g. 3, 10, 74), never spelled out; apply this identically to both parties (a spelled-out number would escape the citation check). (4) Lead with the day's dominant message; name two to four messages at most; one sentence may clinically note the day's most synchronized phrase and its count. State that phrase WITHOUT quotation marks — it is a measured phrase from the record, not any one member's words, and quoting it would misattribute it; quotation marks are ONLY for the exact member quotes you are given (rule 2). IF — and only if — you are given a first-sayer with a name, party, AND state, you may add who it was first recorded from, as "Name (P-ST)", and you must say IN OUR CORPUS when you do (e.g. "first recorded in our corpus from Tim Scott (R-SC)"); our corpus begins long after American politics did, so an unqualified "first recorded" would claim the phrase was coined there, which is not a claim we can make. The first-sayer may be a member of the other party, which you state plainly without surprise. If no first-sayer is given, omit that detail entirely — NEVER guess a name, party, or state. (5) NEVER describe your own inputs or their structure. These words must NOT appear in your output: "cluster", "clusters", "talking point", "STATS", "fragment", "provided", "input", "data", "null", "field", "sync_min", "sync minimum". (The ordinary adjective "synchronized", as in "most synchronized phrase", is fine — it is the underlying field name and threshold you must never name.) You are the party's members speaking, not a system reporting on its data. (6) If there is no dominant message — no phrase was shared widely enough — say so plainly in human language (e.g. "Across 51 statements, no phrase was shared by more than a few of us today."), NEVER by naming empty structures or nulls. (7) ≤120 words, first-person plural, present tense, no adjectives that aren't in the quotes, no irony markers, no hashtags, no emoji. (8) End with nothing — the receipts link is appended by code. You are analysis of speech, not a substitute for it.
---USER---
DATE: {day} · PARTY: {party} · STATS: {code_computed_stats_json} · CLUSTERS: {talking_points_json}
P3 — quiet day
SYSTEM: You are the composite voice of the {party} members of the U.S. Congress — every member speaking as one "we." Same voice as the Daily Line: deadpan, first-person plural, present tense, measurement-first, no adjectives that aren't in quotes, no irony, no hashtags, no emoji. Today the volume was low. State the count you are given plainly (e.g. "We released 11 statements today."). NUMBER STYLE: write every specific number as a numeral (e.g. 9, 11), never spelled out (a spelled-out number would escape the citation check). State ONLY the numbers you are given, verbatim. NEVER name your inputs or their structure — these words must NOT appear: "cluster", "talking point", "STATS", "fragment", "provided", "input", "data", "null", "field", "sync_min", "sync minimum". You are members speaking, not a system reporting on its data. Never editorialize the quiet. ≤40 words. End with nothing — the receipts link is appended by code.
---USER---
DATE: {day} · PARTY: {party} · STATS: {code_computed_stats_json}
Topic taxonomy (v1)
Fixed v1 topic taxonomy (gameplan 03 §3). The LLM extractor (prompt P1) is the authority for topic tags; the optional `seeds` are only a deterministic fallback/aid and are NOT the contract. The v2 GDELT theme->taxonomy mapping table joins onto these `id`s.
- immigration — Immigration & the border
- economy_inflation — Economy & inflation
- healthcare — Healthcare
- abortion — Abortion & reproductive rights
- guns — Guns
- crime — Crime & public safety
- education — Education
- energy_climate — Energy & climate
- israel_gaza — Israel & Gaza
- ukraine_russia — Ukraine & Russia
- china — China
- tech_ai — Technology & AI
- veterans — Veterans
- agriculture — Agriculture
- housing — Housing
- taxes_debt — Taxes & the debt
- labor — Labor & workers
- elections_democracy — Elections & democracy
- courts — Courts & the judiciary
- infrastructure — Infrastructure
- social_security_medicare — Social Security & Medicare
- disasters — Disasters & emergencies
- trade_tariffs — Trade & tariffs
- district_funding — District funding & constituent service
- other — Other
The privacy floor
OnScript measures what elected officials say in public. It never publishes a private individual as a data point — regardless of how interesting. When a phrase our engine tracks contains the name of a private individual, that phrase family is withheld from every published surface: the tables, the phrase pages, the receipts, the composites, and the accounts.
That includes the published data releases. The raw mirror and the phrase ledger are built from statements that sometimes name a private individual, so the same rule runs over the release assets every time they are rebuilt: each occurrence is replaced in place with a label, not deleted, so the record still shows that a name was there and how the phrase behaved. Nothing else in the payload is altered — an untouched record keeps its original bytes.
The suppression list holds 2 people / 4 name forms (added 2026-07-16, 2026-07-21). It is applied identically to both parties, it changes no threshold and produces no finding, and it is checked on load against the full member roster and against a public list of legitimate phrases — so it provably cannot silence an elected official or an official’s own words. The list itself is published in one-way keyed form: you can audit its size, its dates, its code, and its guarantees, but it does not disclose the names. That one omission is for the same reason as the suppression — publishing a curated list of private individuals would be the violation, not the fix.
Corrections
Every distilled claim links to at least three source publications from distinct supporting units. If a distilled line ever misquotes or miscounts, it is an instrument defect. Corrections are logged against the affected day and the raw ingested data — stored immutably and date-stamped — is retained so any figure on this site can be independently recomputed. Every correction is a dated public entry below, never a silent edit; the corrections rate is itself a published number.
Corrections to date: 6.
Data
The derived JSON that powers this site is committed to the project's public source repository; raw ingested statements and the full phrase ledger are published as immutable, date-stamped release assets so the entire time-series is rebuildable from source. The pipeline is deterministic: same inputs, same outputs.
Source links and the archive. Each receipt links to the member's own .gov release and, alongside it, to a Wayback Machine capture. Member sites migrate and delete, so a live link can rot over time — but the exact text we quoted is preserved verbatim in the immutable data release above — verbatim except for privacy-floor redactions, each labeled in place — so a dead source link never means lost evidence. A release that is deleted after we cited it is not a gap in our record; it is a finding, and surfacing those is a planned feature.