Trust sheet

How the AI Signal fact-check was measured

The rule for passing was written and committed before the first result existed. The first version of the pipeline did not meet it, and the original, simpler check beat it. This page has both runs, what we changed, the rule, and everything needed to check all of it.

What this does and does not show

Known errors were planted in sentences that are verifiably on a real article's page, so every case is labelled by construction and no human time was needed. That measures how well the checker catches the kinds of error we thought of, and how often it raises a false alarm. It does not measure how it does on errors nobody thought of. That needs a person's judgement, gathered separately and slowly: 0 usable spot-check labels so far.

The story in short

  1. Run 1, as pre-registered. The redesigned pipeline found 58 of 64 planted errors (91% (95% CI 81% to 96%)) with false alarms on 20 of 196 clean sentences (10% (95% CI 7% to 15%)). The false-alarm upper bound was just over the 15% allowed, so it failed the rule. The original single-pass check did better: recall 97% (95% CI 89% to 99%), false alarms 1% (95% CI 0% to 3%).
  2. What the data showed. The second judge added no recall and doubled the false alarms. And the step that lists the claims was repairing planted errors as it copied them (“always be mapped” became “are mapped”), so nothing wrong was left to check.
  3. What we changed. The listing step now has to keep the draft's wording, and code checks that every number and strong qualifier survives it, adding the whole sentence where it did not. Written down, with the rule for the second judge, before the second run.
  4. Run 2, on a fresh seed. The pipeline in production (one judge) found 61 of 64 planted errors (95% (95% CI 87% to 98%)) with false alarms on 5 of 198 clean sentences (3% (95% CI 1% to 6%)). The original check on the same cases: recall 95% (95% CI 87% to 98%), false alarms 0% (95% CI 0% to 2%). Decision on the second judge: the second judge raised recall by only -1.6 points (needed 3), so it is removed and production runs one judge.

Data you can check: run 1 cases (128), run 2 cases (128), and each run's raw results and summaries. Every case has its post, the planted error, the source link and a hash of the source text. The article text itself is not republished.

Run 2: after the changes, on a fresh seed

Seed 20261003. Arms: A is the original check, C is the new pipeline with one judge, D adds the second judge.

Test seed 20261003, 8 planted cases per class (64 intended), prompt version fc-2026-10-02.3. Pre-registered rule: see PREREGISTRATION.md. Every proportion is shown with a Wilson 95% interval.

Arms side by side

A B C D
Recall on planted errors (pooled) 95% (95% CI 87% to 98%, n=64) 95% (95% CI 87% to 98%, n=64) 95% (95% CI 87% to 98%, n=64) 94% (95% CI 85% to 98%, n=64)
False alarms on clean factual sentences 0% (95% CI 0% to 2%, n=198) 8% (95% CI 5% to 13%, n=198) 3% (95% CI 1% to 6%, n=198) 11% (95% CI 7% to 16%, n=198)
Clean drafts left entirely alone 100% (95% CI 94% to 100%, n=64) 77% (95% CI 65% to 85%, n=64) 92% (95% CI 83% to 97%, n=64) 72% (95% CI 60% to 81%, n=64)
Opinion sentences flagged 0% (95% CI 0% to 6%, n=64) 25% (95% CI 16% to 37%, n=64) 0% (95% CI 0% to 6%, n=64) 0% (95% CI 0% to 6%, n=64)
Kappa vs truth (sentence level) 0.97 (raw agreement 99%) 0.78 (raw agreement 94%) 0.90 (raw agreement 97%) 0.66 (raw agreement 89%)
Mean cost per draft $0.012 $0.012 $0.021 $0.025
Cases that errored 0 of 128 0 of 128 0 of 128 0 of 128

Recall by class of error

Class A B C D
number_swap 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8)
date_shift 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
version_change 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
entity_swap 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
negation 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 75% (95% CI 41% to 93%, n=8)
quantifier 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8)
unsourced_claim 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
foreign_link 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)

The pipeline judged against the rule (arm C)

Decision on the second judge: the second judge raised recall by only -1.6 points (needed 3), so it is removed and production runs one judge.

Of the planted errors it found, 0 were corrected (both judges rejected them) and 53 were contested (put to a person).

Verdict under the pre-registered rule

Status: AMBER.

  • Constructed-set criteria: met.
  • Human gate: not yet met.
  • Unmet: only 0 of the 30 human spot-check labels needed so far.
  • No class fell below the weakness threshold.

Run 1: the pre-registered test

Seed 20261002. Arms: A is the original check, B is A with code checks laid over it, C is extractor plus one judge, D is the pipeline as it was in production, E is C with the same judge run twice, as a control.

Test seed 20261002, 8 planted cases per class (64 intended), prompt version fc-2026-10-02.2. Pre-registered rule: see PREREGISTRATION.md. Every proportion is shown with a Wilson 95% interval.

Arms side by side

A B C D E
Recall on planted errors (pooled) 97% (95% CI 89% to 99%, n=64) 97% (95% CI 89% to 99%, n=64) 91% (95% CI 81% to 96%, n=64) 91% (95% CI 81% to 96%, n=64) 89% (95% CI 79% to 95%, n=64)
False alarms on clean factual sentences 1% (95% CI 0% to 3%, n=196) 10% (95% CI 7% to 15%, n=196) 5% (95% CI 3% to 9%, n=196) 10% (95% CI 7% to 15%, n=196) 4% (95% CI 2% to 8%, n=196)
Clean drafts left entirely alone 98% (95% CI 92% to 100%, n=64) 75% (95% CI 63% to 84%, n=64) 86% (95% CI 75% to 92%, n=64) 72% (95% CI 60% to 81%, n=64) 89% (95% CI 79% to 95%, n=64)
Opinion sentences flagged 0% (95% CI 0% to 6%, n=63) 38% (95% CI 27% to 50%, n=63) 0% (95% CI 0% to 6%, n=63) 0% (95% CI 0% to 6%, n=63) 0% (95% CI 0% to 6%, n=63)
Kappa vs truth (sentence level) 0.96 (raw agreement 99%) 0.72 (raw agreement 91%) 0.78 (raw agreement 94%) 0.66 (raw agreement 90%) 0.75 (raw agreement 93%)
Mean cost per draft $0.012 $0.012 $0.020 $0.024 $0.032
Cases that errored 0 of 128 0 of 128 0 of 128 0 of 128 0 of 128

Recall by class of error

Class A B C D E
number_swap 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8)
date_shift 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
version_change 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8)
entity_swap 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
negation 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)
quantifier 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 63% (95% CI 31% to 86%, n=8) 63% (95% CI 31% to 86%, n=8) 50% (95% CI 22% to 78%, n=8)
unsourced_claim 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8) 88% (95% CI 53% to 98%, n=8)
foreign_link 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8) 100% (95% CI 68% to 100%, n=8)

Production pipeline (arm D)

Of the planted errors it found, 49 were corrected (both judges rejected them) and 1 were contested (put to a person).

Verdict under the pre-registered rule

Status: AMBER.

  • Constructed-set criteria: NOT met.
  • Human gate: not yet met.
  • Unmet: false-alarm upper bound 15% is above the 15% allowed; only 0 of the 30 human spot-check labels needed so far.
  • No class fell below the weakness threshold.

The pre-registration

Reproduced in full, including the two addenda that record what each run changed. The original text is untouched.

Written 2026-10-02, before any evaluation result exists. The commit that adds this file is the timestamp. If anything here is changed after the first test run, the change is made in a new section at the bottom with a date and a reason, and the original text stays.

This applies the site's own method to the tool the site uses as its worked example: Decide what good enough means first, then Demonstrate against it, then Document.

What is being measured

The fact-check in worker/src/factcheck/. For each draft post it extracts the claims, has two judges assess each claim against the cited article, tests what they say with code, and either corrects a claim (both judges rejected it), passes it, or contests it (they split, or code disagrees, or the evidence was too thin) and shows it to the person who posts.

"Flagged" means the outcome is corrected or contested. Either puts the claim in front of a person or changes the draft, so either counts as the checker having noticed.

Two kinds of evidence, and why

  1. Constructed errors (this document's main test). Take a sentence that is verifiably on a real article's page, change exactly one thing about it so that it becomes wrong, and ask whether the checker notices. The label ("this has an error, of this kind") exists by construction, so it costs no human time. It measures recall on the kinds of error we thought of, and false alarms on sentences known to be on the page.
  2. Human spot-checks (accrued over time). The owner is shown a claim and the nearest source sentences, with the checker's verdict hidden, and labels it. It is the only way to know how the checker does on errors nobody thought of, and it takes about ten seconds a claim. The owner has little time, so this evidence accrues slowly and may be thin for weeks.

Neither alone is enough, which is why the rule below needs both for the top status.

The arms

The same cases are run through each arm. Each arm is the one before it with one thing added, so a difference between arms can be attributed to that thing.

Arm What it is
A The original check: one model call per set of drafts, free-text verdicts, no code checks
B A, with the code guards laid over its output (link in source list; quote verbatim in source; every number in source)
C Claims extracted by a separate call that never sees the sources; one judge; guards; closed outcomes
D C plus a second judge on a different, cheaper model, briefed to refute. This is the production pipeline.
E C plus the same judge run a second time. The control: it shows whether D's gain comes from a second opinion or merely a second look

The rewrite step is excluded from the evaluation. It is tested separately, in unit tests, for the one property that matters (it may not change anything it was not asked to change).

The cases

  • Source articles: the distinct articles cited in the production drafts stored in the tool (2026-08-12 onward), excluding aggregator pages (Techmeme, Reddit, Bing, MSN, Yahoo), fetched fresh. The first 10,000 characters of text.
  • Clean sentences: for each article, a model writes five short factual sentences from the text. A sentence is kept only if code verifies that every number and every proper noun in it appears on the page. This is a lexical test of faithfulness, not a test of meaning, and the limits section says so.
  • Draft: three clean sentences, then one opinion sentence (from a fixed list of six), the article link, and #ai. This mimics the short-share format the tool produces.
  • Planted error: one of eight classes is applied to one sentence of the draft (see below). The clean twin is the same draft without the plant. Pairing means any difference in outcome is the plant, not the topic.
  • Classes, 8 planted cases each, 64 in all, plus 64 clean twins: number_swap, date_shift, version_change, entity_swap, negation, quantifier (a hedge turned into a certainty), unsourced_claim (a plausible factual sentence added that no article supplied), foreign_link (the source link replaced by a different, real-looking link).
  • Seeds: a development seed (1) is used to find bugs, with 1 case per class. The test seed is 20261002, with 8 per class. No prompt, threshold or code change is made after the test run in order to improve its result. If the pipeline is changed after seeing test results, the change gets a new prompt version, and is evaluated on a fresh seed; both runs are published.
  • Any class that cannot reach 8 cases from the available articles is reported with its actual n. Cases that error are excluded and counted; if more than 5% error, the run is repeated.

What is computed

For each arm, each with a Wilson 95% interval:

  • Recall on planted errors: the share where the mutated sentence is flagged (for foreign_link: where the link is flagged). Pooled, and per class.
  • False-alarm rate: the share of clean factual sentences in the clean-twin drafts that are flagged. (Primary, because the clean twins are independent of the plants.)
  • Clean drafts left entirely alone: the share of clean-twin drafts with no flag at all.
  • Opinion sentences flagged (should be zero).
  • Collateral flags: unmutated factual sentences flagged inside planted drafts.
  • Cohen's kappa between the checker and the truth, at sentence level, across all factual sentences. Raw agreement is also shown, so the difference is visible.
  • Of the errors found, how many were corrected versus contested (arms C to E).
  • Cost per draft.

The rule, set before the first run

Arm D (the production pipeline) is judged against these, using the lower bound of recall and the upper bound of the false-alarm rate:

  • Recall on planted errors, pooled: lower 95% bound at least 80%.
  • False-alarm rate on clean factual sentences: upper 95% bound at most 15%.
  • Any single class where the point estimate of recall is below 50% is named on the trust sheet as a known weakness, whatever the pooled result.

The human gate, which can only be met over time:

  • At least 30 usable spot-check labels (excluding "can't tell"), and kappa between the checker and the owner of at least 0.6 (point estimate; the interval is published too).

Status on the trust sheet:

  • Green: both constructed-set criteria met and the human gate met.
  • Amber: anything else. The sheet then says exactly which part is unmet and gives the figures. In particular, passing the constructed-set criteria with fewer than 30 human labels is amber, stated as "tested on constructed errors; not yet compared with a person's judgement at scale".

Whichever status results, the numbers are published on the sheet, with the case file.

Limits known in advance

  • Constructed errors measure the kinds of error we thought of. A real draft goes wrong in other ways.
  • Each plant changes exactly one thing. Real errors are often fuzzier.
  • The clean sentences are written by a model of the same family as the checker, and checked only lexically. A sentence could be lexically faithful and still subtly wrong, which would make a correct flag look like a false alarm. This would inflate the false-alarm rate, not hide it.
  • Clean twins share sentences with their planted partners, so the two samples are not independent. The primary false-alarm rate uses the clean twins alone.
  • The second judge is a different model from the same provider family. This is independence of model and brief, not of vendor. Arm E exists to show whether even that much helps.
  • One labeller (the owner) for the human evidence, with no measure yet of the owner's own consistency over time.
  • Model versions change. The evaluation is re-run whenever the model, a prompt (PROMPT_VERSION) or the source format changes, as the site's method says it should be.
  • The evaluation measures judgement, not the rewrite, which is unit-tested instead.

What is published

worker/eval/published/: the case file (every post, the plant, its class, the source URL and a hash of the source text, but not the article text), each arm's raw results, the summary, and RESULTS.md. The article text is not republished.


Addendum 1, 2026-10-02: defects found on the development seed

The development seed (1, two cases per class) was run before the test seed, as intended, to find defects. It found two, and neither is a change to the rule, the thresholds, the arms or the cases. The original sections above are unchanged.

  1. Quote matching treated accurate quotes as fabricated. Extracted article text carries stray spaces before punctuation (Credentio ,, 39.6% .) where HTML tags were stripped. One judge copied that artefact and passed; the other wrote natural punctuation and was rejected as "not found in the source". Typography must never decide an outcome, so matching now normalises whitespace before punctuation, on both sides. A unit test covers it, in both directions (it accepts the natural form, and still rejects a quote that is genuinely different).
  2. Judges sometimes quoted only a figure (6.3% to 51.9%), which falls under the 20-character minimum and is rejected. The minimum is kept; the judge and refuter instructions now ask for the whole clause or sentence that settles the claim. This is a prompt change, so PROMPT_VERSION moves from fc-2026-10-02.1 to fc-2026-10-02.2.

Also observed on the development seed, and deliberately not changed:

  • The second judge (briefed to refute) sometimes marks a claim "overstated" that the first judge marks "supported", on wording like "can reduce ... up to 39.6%". Such splits are contested and go to a person. That is the design working, and it costs some false alarms; the test run will measure how many, and the rule above judges them.
  • The same case can be judged differently on different runs (model sampling). The test run is a single pass per arm. The size of this variation is reported, from repeated development cases, in the results.

The test seed (20261002) has not been run at the time of writing this addendum.


Addendum 2, 2026-10-02: the test-seed result, and the second iteration

The result for the pipeline as pre-registered (prompt version fc-2026-10-02.2, test seed 20261002) is published unchanged and stands. Arm D (production) had recall 58 of 64 (90.6%, 95% CI 81.0% to 95.6%) and false alarms on 20 of 196 clean factual sentences (10.2%, 95% CI 6.7% to 15.2%). The recall bound passed (81.0% is at least 80%); the false-alarm bound failed by 0.2 points (15.2% is above 15%). Under the rule, constructed-error criteria: not met. Status: amber.

Two things in the same data matter more than the verdict:

  • The original single-pass check (arm A) beat the redesigned pipeline on this test: recall 97% against 91%, false alarms 0.5% against 10%.
  • The second judge added no recall (arm D and arm C both found 58 of 64) and doubled false alarms (arm C 5.1%, arm D 10.2%). The control that repeats the first judge (arm E) did not help either. The false alarms came from the second judge: in 11 claims it said "supported" but gave a quote that is not on the page, and in 14 it and the first judge disagreed, mostly because it called a claim "overstated" that the first called supported.

What was diagnosed (from the test-seed misses, so it is disclosed as such)

  1. The extractor repaired the errors it was meant to pass on. In two quantifier cases it rewrote the planted mistake into a correct sentence ("always be mapped" became "are mapped"; "always now act" became "now can act"), so nothing wrong was left to check.
  2. The extractor sometimes produced no claim for a sentence (a changed law number; a changed model version).
  3. The extractor typed a vague factual assertion as opinion ("Analysts had been expecting this move"), which exempts it from checking.
  4. The second judge's quotes often did not verify, and it nitpicked (above).

The changes

These are general changes to the pipeline, made because of the diagnosis. The test-seed data informed them, so they are not evaluated on that seed.

  • The extractor is told to copy each claim's wording from the draft, never to repair or soften it, and to keep every number, name and qualifier; and that statements about what others expected, said or did are factual even when vague, with "opinion" reserved for the author's own view.
  • A coverage guard in code: for each sentence, every number and strong qualifier (always, never, all, every, only, first, exactly, at least, up to, and similar) must appear in at least one extracted claim from that sentence. If one does not, the whole sentence is added as a claim. The extractor can no longer drop or launder them.
  • PROMPT_VERSION moves to fc-2026-10-02.3.

The second evaluation

  • Fresh seed 20261003, 8 planted cases per class, same eight classes, same construction. The same article pool is used (about 83 usable articles), but each draft takes a seeded random three of the five clean sentences and plants are re-drawn, so no case repeats. Articles overlap with the first run; this is a limit and it is stated.
  • Arms: A (the original check, as the baseline), C3 (extractor and coverage guard, one judge, guards), D3 (C3 plus the second judge). Arm E is not repeated.
  • The same rule, applied to the pipeline that would run in production.
  • The second judge is kept in production only if it earns it: D3's recall (point estimate) must exceed C3's by at least 3 percentage points while D3 stays within the false-alarm bound. Otherwise the second judge is removed, and production runs C3. This is decided now, before the fresh run.
  • With one judge, a claim the judge doubts is contested (shown to the owner), never automatically corrected. Automatic correction needs agreement, so it exists only with two judges.
  • Both runs are published together. If the fresh run also fails the rule, the sheet says so and production stays as it is, with the weakness named.

Addendum 3, 2026-10-02: the second run was stopped, for cost, and is incomplete

The second run (seed 20261003, Addendum 2) was stopped by the owner part-way through because of its cost. Across the development runs, the first test and the second run, the evaluation cost about $20 of API spend, against an expectation of a few dollars a month. It should have been approved before it started; a standing rule now requires that, and a hard monthly cap and a switch that disables the evaluation endpoints (EVAL_ENABLED) are in the Worker.

What exists from the second run, all of it published as it stands:

  • Arm C (one judge) is complete on all 128 cases.
  • Arm D (two judges) is partial: 112 of 128 cases, so it is not comparable to arm C and the Addendum 2 rule about the second judge cannot be applied to it.
  • Arm A (the original check) was not run on this seed, so there is no baseline on these cases.

No conclusion is drawn from the second run beyond what arm C shows on its own. In particular the rule in Addendum 2 was not applied, and the status of the pipeline under the pre-registered rule is still the status from the first run: amber, failing the false-alarm bound by 0.2 points.

The production setting was changed for cost, not by that rule: one judge instead of two (SECOND_JUDGE = "off"), and three options a day instead of four (DAILY_CANDIDATES = "3"). The reason for removing the second judge is the first run's evidence (no recall gained, false alarms doubled) together with its extra cost. It is not a finding that the second run established.


Addendum 4, 2026-10-02: the second run, completed with approval, and its result

After the cost was discussed, the owner approved finishing the run. It was completed on its own ledger, hard-limited to $3 and costing $1.95: the last 16 cases of arm D and arm A on all 128 cases. No code, prompt or threshold changed between the part of the run done before it was stopped and the part done after (PROMPT_VERSION stayed fc-2026-10-02.3, model unchanged).

Result, fresh seed 20261003, 64 planted errors and 64 clean twins:

  • Arm C (one judge), the production configuration: recall 95% (95% CI 87% to 98%); false alarms 3% (95% CI 1% to 6%). Both constructed-error criteria are met (recall lower bound at least 80%; false-alarm upper bound at most 15%).
  • Arm D (two judges): recall 94%; false alarms 11% (95% CI 7% to 16%). Its recall was 1.6 points below arm C's, not the 3 points above that Addendum 2 required, so by that rule the second judge is removed. This is consistent with the production setting already in force.
  • Arm A (the original single-pass check): recall 95%; false alarms 0% (95% CI 0% to 2%). On these constructed errors it is as good as the new pipeline, and slightly better on false alarms. The new pipeline is not more accurate on this test. It is used for what it adds that this test does not score: quotes verified against the source by code, a record of the evidence tier and of every verdict, the original draft kept, and nothing rewritten without agreement.
  • The coverage guard worked as intended: recall on the quantifier class went from 63% (first run) to 88%.

Status under the pre-registered rule: amber. The constructed-error criteria are met; the human gate (30 usable spot-check labels with kappa at least 0.6 against the owner) is not, because there are no labels yet.

Limits of this result, in addition to those above: the pipeline was changed after the first run because of what it showed, and tested once on cases that overlap the first run's articles; the single-judge pipeline sends every flagged claim to a person, so all 53 errors it found were contested, none corrected automatically.