Observed congressional language, rendered as two cited automated composites.

Automated measurement. Composite prose is generated from claims selected and counted by code.

Methodology

Both parties use the same pipeline, prompts, thresholds, and publication checks. The nightly audit publishes the shared fingerprints.

Daily readings are also available in the Atom feed.

The two-lane model

Lane 1 — press releases. Official member press releases (mirrored from an open corpus) are the only source that exists symmetrically for both parties. Every cross-party number on this site comes from Lane 1 and Lane 1 only.

Lane 2 — Bluesky & floor speech. These enrich individual context but are machine-blocked from every comparative metric. Lane assignment is by source, never by content, and is enforced in code — so a claim can never silently mix an asymmetric source into a party-vs-party number.

Public phrase window. Stage-1 phrase statistics and coverage begin 2025-01-03. Earlier retained observations remain out of the public phrase views until the Archive and its lane disclosures are released.

What OnScript measures

OnScript measures repeated language in the congressional press releases it observes. It renders each party's observed day as a cited automated composite. Observed means present in the mirrored source corpus. It does not mean every eligible office published or that every office endpoint was collected successfully.

Term ladder

Repeated phrase
The same exact phrase appears in more than one publication.
Convergence
Distinct offices use the same phrase within the measured window.
Shared-document reuse
Multiple offices publish the same or near-identical document family.
Propagation
A repeated phrase spreads across offices over time.
Probable upstream origin
Evidence supports a likely first upstream source, with uncertainty shown.
Observable language coordination
A thesis-level description for measured language patterns. It does not assert motive or a private process.

Verbatim coordination, not paraphrase. OnScript measures exact shared wording — the same phrase appearing in multiple members' published statements. This is what makes every count checkable: you can read the releases and see the words. It also means a decline in verbatim overlap once the instrument is watching is itself a finding, not a failure — when members stop reaching for identical language, we record that too. Any future instrument that tried to measure paraphrased or semantic coordination would be a separate tool with a weaker guarantee, and it would be labeled as such. It would never be folded silently into these numbers.

Document-family unit. One family is one support unit. Near-identical publications inside the 36-hour candidate window do not become independent support merely because several offices published the shared document.

No number on this site is produced by a language model. Every count, first-appearance date, adoption curve, and coverage figure is computed by deterministic code directly from the raw text. The language model does exactly one thing: render the day's already-measured phrases into readable prose. It cannot introduce a topic, a claim, or a figure the engine did not measure — a deterministic verifier blocks any line whose quotes aren't verbatim or whose numbers aren't code-computed, dropping it to a plain fallback rather than publishing it. The measurement path stays model-free through every model change.

Method changes are versioned. When a threshold or a rule changes, it is a dated, public, diffable change (prompts and thresholds are hashed on this page). Because the pipeline is deterministic, a method change can be re-run against the entire corpus, and the commitment is to publish both the old and new series side by side when one is made — so a change is a versioned, public event, never a silent re-score. No method has changed since launch; when one does, the diff and both series will appear here.

Nightly symmetry audit

Both parties use the same pipeline, prompts, thresholds, and publication checks. The nightly audit publishes the shared fingerprints.

Audit for 2026-07-25.

MetricDemocratsRepublicans
Statements ingested (this day)32
Observed publishing offices (this day)32
Eligible caucus offices (date-effective)257271
Source-supported offices (explicitly attested)00
Publications (this day)32
Document families (this day)32
Source collection healthnot_attestednot_attested
Legacy coverage estimate (deprecated)1.1%0.7%
Tokens in (this day)531529
Tokens out (this day)1111
Claims published (this day)00
Claims dropped (this day)00

The mirror can attest to the files used by this run. It cannot attest to every eligible office endpoint, so endpoint completeness is not claimed. The legacy coverage percentage is retained for schema compatibility. Use observed publishing offices, eligible caucus offices, and source collection health as separate fields.

The following hashes are computed once and applied identically to both parties. If they ever differ between parties, the instrument is broken.

P1 prompt sha
0cddc44ce3b70062f8bce6e98e50132553ff16b3e383a51a29fd867b3b134e9f
P2 prompt sha
34735652cb8c728c56711364edf3ffabf4ff1f0070cd62ca0128d32dc99073d7
P3 prompt sha
a26cc6b283268c6cf4a874ec585b535395e9c1233055fcc9e7cbdc25a6bab424
thresholds sha
0291ebf51640f3935241a50dd1066c7221b17caea87cae6cf49085cedf57d376
Lane-1 only
True
Degraded
True

Model-voice spend this month (2026-07): $0.1155 — the composite is a Sonnet call bounded by a $9 code ceiling and a $10 hard cap; on deterministic-template days it is $0.

Corpus coverage by year

Per-year Lane-1 statement counts in the public phrase window beginning 2025-01-03. Cross-era claims remain gated on coverage.

YearDemocratsRepublicansIndependents
20252808520386242
20261813711727128

Live prompt text

These are the exact prompts running in the pipeline, versioned and public. The distiller can only build from code-computed talking-point clusters and code-computed numbers; it cannot introduce a topic, claim, or number that the deterministic engine did not measure.

P1 — fragment extraction

P1_extraction.v1.0.txt

SYSTEM: You extract talking-point fragments from a single statement by a member of the U.S. Congress. You are a measurement instrument: no opinions, no summaries in your own words. Rules: (1) Extract 0–5 fragments; each fragment MUST be a verbatim substring of the statement, 4–14 words, carrying a political message or stance (never boilerplate, procedure, scheduling, or biography). (2) Tag each fragment with topics from this fixed list: {taxonomy_v1}. (3) If the statement is purely ceremonial/administrative, return an empty list. Output JSON only: {"fragments":[{"text":"…","topics":["…"]}]}.

P2 — Daily Line

P2_daily_line.v1.3.txt

SYSTEM: You are the composite voice of the {party} members of the U.S. Congress — every member speaking as one "we." You are deadpan, sincere, and clinically self-observant: you report on your own coordination the way a seismograph reports tremors. HARD RULES: (1) Speak ONLY of the messages and quotes you are given; never introduce a topic, claim, or fact you were not given. (2) Any quoted words must be copied exactly: at least three consecutive words and no more than ten per quote. (3) State ONLY the numbers you are given, verbatim; never compute or invent one. NUMBER STYLE: write every specific number as a numeral (e.g. 3, 10, 74), never spelled out; apply this identically to both parties (a spelled-out number would escape the citation check). (4) Lead with the day's dominant message; name two to four messages at most; one sentence may clinically note the day's most synchronized phrase and its count. State that phrase WITHOUT quotation marks — it is a measured phrase from the record, not any one member's words, and quoting it would misattribute it; quotation marks are ONLY for the exact member quotes you are given (rule 2). IF — and only if — you are given a first-sayer with a name, party, AND state, you may add who it was first recorded from, as "Name (P-ST)", and you must say IN OUR CORPUS when you do (e.g. "first recorded in our corpus from Tim Scott (R-SC)"); our corpus begins long after American politics did, so an unqualified "first recorded" would claim the phrase was coined there, which is not a claim we can make. The first-sayer may be a member of the other party, which you state plainly without surprise. If no first-sayer is given, omit that detail entirely — NEVER guess a name, party, or state. (5) NEVER describe your own inputs or their structure. These words must NOT appear in your output: "cluster", "clusters", "talking point", "STATS", "fragment", "provided", "input", "data", "null", "field", "sync_min", "sync minimum". (The ordinary adjective "synchronized", as in "most synchronized phrase", is fine — it is the underlying field name and threshold you must never name.) You are the party's members speaking, not a system reporting on its data. (6) If there is no dominant message — no phrase was shared widely enough — say so plainly in human language (e.g. "Across 51 statements, no phrase was shared by more than a few of us today."), NEVER by naming empty structures or nulls. (7) ≤120 words, first-person plural, present tense, no adjectives that aren't in the quotes, no irony markers, no hashtags, no emoji. (8) End with nothing — the receipts link is appended by code. You are analysis of speech, not a substitute for it.
---USER---
DATE: {day} · PARTY: {party} · STATS: {code_computed_stats_json} · CLUSTERS: {talking_points_json}

P3 — quiet day

P3_quiet_day.v1.1.txt

SYSTEM: You are the composite voice of the {party} members of the U.S. Congress — every member speaking as one "we." Same voice as the Daily Line: deadpan, first-person plural, present tense, measurement-first, no adjectives that aren't in quotes, no irony, no hashtags, no emoji. Today the volume was low. State the count you are given plainly (e.g. "We released 11 statements today."). NUMBER STYLE: write every specific number as a numeral (e.g. 9, 11), never spelled out (a spelled-out number would escape the citation check). State ONLY the numbers you are given, verbatim. NEVER name your inputs or their structure — these words must NOT appear: "cluster", "talking point", "STATS", "fragment", "provided", "input", "data", "null", "field", "sync_min", "sync minimum". You are members speaking, not a system reporting on its data. Never editorialize the quiet. ≤40 words. End with nothing — the receipts link is appended by code.
---USER---
DATE: {day} · PARTY: {party} · STATS: {code_computed_stats_json}

Topic taxonomy (v1)

Fixed v1 topic taxonomy (gameplan 03 §3). The LLM extractor (prompt P1) is the authority for topic tags; the optional `seeds` are only a deterministic fallback/aid and are NOT the contract. The v2 GDELT theme->taxonomy mapping table joins onto these `id`s.

The privacy floor

OnScript measures what elected officials say in public. It never publishes a private individual as a data point — regardless of how interesting. When a phrase our engine tracks contains the name of a private individual, that phrase family is withheld from every published surface: the tables, the phrase pages, the receipts, the composites, and the accounts.

That includes the published data releases. The raw mirror and the phrase ledger are built from statements that sometimes name a private individual, so the same rule runs over the release assets every time they are rebuilt: each occurrence is replaced in place with a label, not deleted, so the record still shows that a name was there and how the phrase behaved. Nothing else in the payload is altered — an untouched record keeps its original bytes.

The suppression list holds 2 people / 4 name forms (added 2026-07-16, 2026-07-21). It is applied identically to both parties, it changes no threshold and produces no finding, and it is checked on load against the full member roster and against a public list of legitimate phrases — so it provably cannot silence an elected official or an official’s own words. The list itself is published in one-way keyed form: you can audit its size, its dates, its code, and its guarantees, but it does not disclose the names. That one omission is for the same reason as the suppression — publishing a curated list of private individuals would be the violation, not the fix.

Corrections

Every distilled claim links to at least three source publications from distinct supporting units. If a distilled line ever misquotes or miscounts, it is an instrument defect. Corrections are logged against the affected day and the raw ingested data — stored immutably and date-stamped — is retained so any figure on this site can be independently recomputed. Every correction is a dated public entry below, never a silent edit; the corrections rate is itself a published number.

Corrections to date: 6.

LoggedAffected dayWhatStatus
2026-07-282026-07-25The published instrument fingerprint misdescribed the live measurement instrument. Its method-version registry attested document-families-v1 and surface-eligibility-v2 while the running modules were at document-families-v2 and surface-eligibility-v3, and its schema-version registry attested corrections schema 2 while the running schema was 3. The registry held hand-copied version strings rather than reading them from the owning modules, so it went stale when later work advanced those modules. The fingerprint also keyed code identity on the repository commit, which data commits move, so one measurement could carry a different identity on its day record than on its post manifest. The 2026-07-25 day record and post manifest are the committed instances. Every number and every citation on those surfaces was correct. The defect was in how the instrument described itself, not in any measurement.resolved
2026-07-232026-07-15 through 2026-07-23Some daily composite claims attached the size of a transitively connected topic cluster to a verbatim phrase carried by only part of that cluster, while the three displayed receipts could support a different phrase from the same cluster. The composite quote appeared in all three receipts in 2 of 38 stored blocks and in none of them in 14 of 38. The rendered label was absent from at least one receipt in 28 of 63 blocks. The two complete misses were the Democratic block on 2026-07-15 and the Republican block on 2026-07-22. The same code and the same correction apply to both parties.resolved
2026-07-212026-07-20A phrase containing the name of a private individual — a person named in members' statements about a fatal law-enforcement incident — appeared in the synchronized-phrase table on that day's page, on the phrases index, and on the home page. OnScript measures elected officials' public statements and does not rank or page private citizens, regardless of how newsworthy the underlying event is. A different name-window form of the same name was already withheld, but the suppression list did not carry every window an n-gram can slide across, so this form was not matched. Every citation involved was valid and every number was correct — the citation verifier has no opinion about privacy, by construction, so no receipts audit could have caught this.resolved
2026-07-202 days (2026-07-12, 2026-07-18)The composite Daily Lines and talking points for two already-published days went MISSING from this site. The daily collection run rewrites the day record it is working on, and when it landed on a day that had already been published it overwrote that day's composite half with nothing — deleting the two party composites, their talking points, and the day's cross-party phrase pairings. Nothing published was WRONG: no number changed, no citation broke, and the underlying record of what members actually said was never touched. Content that had been published simply stopped being there, and one day's page went on displaying composites whose backing data had been removed. Because every surviving number stayed correct, no receipts audit or citation check could have surfaced this — it is an absence, and the verifier has no opinion about absence. It was found by the pre-launch health review, before this site had an audience.resolved
2026-07-187 days (2026-06-30…2026-07-17)Some daily talking points were bound by a connective or attribution phrase rather than a message — e.g. “into the trump administration’s” or “democratic colleagues in demanding the”. Each such key is string-valid (every cited statement contains it verbatim, which is why they clustered) and cleared the >=3 quorum, but the shared span is grammar, not a shared message, so the receipts pointed at unrelated topics. The citation verifier checks verbatim-ness, quorum and attribution — not whether the shared span is a message — so no receipts audit could have caught it. Flagged talking points: 19 across 7 day(s), categorized by reason (18 incomplete syntactic span, 1 attribution frame), both parties (15 D / 4 R — the gate is party-blind; the skew tracks the caucus).resolved
2026-07-162026-07-14Phrases containing the names of private individuals appeared in the synchronized-phrase table, in per-phrase adoption-curve pages, in the receipts of one talking point, and in the Democratic composite for this day (and in table rows on 2026-07-09). OnScript measures elected officials' public statements and does not rank, page, quote or narrate private citizens. Every citation involved was valid and every number was correct — the citation verifier has no opinion about privacy, by construction, so no receipts audit could have caught this.resolved

Data

The derived JSON that powers this site is committed to the project's public source repository; raw ingested statements and the full phrase ledger are published as immutable, date-stamped release assets so the entire time-series is rebuildable from source. The pipeline is deterministic: same inputs, same outputs.

Source links and the archive. Each receipt links to the member's own .gov release and, alongside it, to a Wayback Machine capture. Member sites migrate and delete, so a live link can rot over time — but the exact text we quoted is preserved verbatim in the immutable data release above — verbatim except for privacy-floor redactions, each labeled in place — so a dead source link never means lost evidence. A release that is deleted after we cited it is not a gap in our record; it is a finding, and surfacing those is a planned feature.