Stage 3 — linking and coding
For each paper the screen admitted, answer two questions: which original study did this re-test, and what happened. This is the stage that produces the actual research output, and the only stage that reads whole papers.
Where Stage 3 sits
Stage 3 takes the works the screen passed and turns each into one permanent verdict; a separate export renders those verdicts into the CSV that leaves this repository.
The two questions
They are answered together but judged separately. A paper's target can be certain while its outcome is not, and the reverse never happens — an outcome is only ever coded against a link the pipeline accepted.
Which original? Resolved by a ladder of steps, cheapest
first. Writes doi_o, title_o,
authors_o, study_o and the
link_method that found them.
What happened? Coded from the same reading that found the link. A replication gets one axis; a reproduction gets two.
One row per replication–original pair. A paper that
re-tests three different originals becomes three rows. Several studies
of one original stay one row, with the study numbers collapsed into
study_o.
The resolution ladder
Steps in cost order. The row takes the first one that resolves and stops there, so a paper whose title states its target never costs a model call, and only the papers that resist everything cheaper reach the last step and get their full text downloaded. The counts are how many shipped rows each step actually produced, which is the honest picture of where the work happens.
Why the search steps sit at the bottom. OpenAlex bills by request shape, and the spread is three orders of magnitude: a lookup by DOI is free, a filter query is 1×, a free-text search is 10×, and a content download is 100×. So a title search is the single most expensive thing a row can do short of downloading the paper — which is exactly why it is tried after everything a cached candidate list can answer.
What the ladder has learned
The running record kept beside the ladder version — newest first. Each entry is a change to how an original is found, and why.
When a step does not end the row
A step ends the row only when it both resolved the link and settled the outcome. This is the single most confusing thing about the ladder, and it exists for a concrete reason: the abstract often names the original clearly while saying nothing about how the replication turned out.
So the ladder carries the accepted link and keeps descending towards the closing sections that state the verdict. When the full-text step does answer, its reading replaces the carried one — except that a later unsettled outcome never overwrites an earlier settled one.
A carried link outranks a withheld rule pick at every no-answer
exit. No PDF, no document, no context, an incomplete screen, a
provider failure — an outage below an accepted link no longer writes
target_pending over it. A provider failure is not an answer, so it
does not restore a withheld pick either.
Getting the document
The full-text step needs the paper. A waterfall tries every source in priority order and stops at the first that returns something that passes its own content check.
A document need not be a PDF, but every source has a content check a record page fails. Four sources hand back a sections dict instead of a file, and each is paired with the test that says whether what came back is a document at all. A repository landing page restates the abstract and adds citation chrome; coding a row from one looks like full text and is not — so the HTML check subtracts the abstract rather than thresholding the total. Measured 2026-08-07: five landing pages carried 0–1,706 characters beyond their abstracts, three real full texts carried 49,193–71,641.
A result that fails its check is no document: the row ends at
no_fulltext_available and the failure is never cached as a success.
Blank pdf_source on a shipped row is not a gap — it means the row
resolved before the full-text step was ever needed.
Coding the outcome
Two vocabularies, because a replication and a reproduction are different questions. They are different scales and are never summed.
uninformative is the authors' verdict;
cannot_be_determined is ours. The first says the study
itself could not settle the question. The second says we could not tell from
what we read.
A reproduction's outcome is its two settled axis values joined,
which is why the flat outcome list hides which axis failed. technical
failure exists on the computation axis because the most common real
reproduction outcome in economics is being defeated by the materials — no data,
no code, an unrunnable workflow — and the old grid could only record that as a
numerical disagreement nobody observed.
Checking its own work
Three independent checks run before a row is allowed to settle, each aimed at a different way a link can be wrong.
The metadata doi_o actually points to is fetched (CrossRef,
then OpenAlex) and compared with the resolved record. Mismatches are
re-resolved from title and author in three tiers, strictest first — a wrong
correction is worse than a flag. Runs once, inside the tier, and the
answer is stored on the row.
DOI verification checks the record against its own metadata, never against
the target the paper named — so the wrong entry picked from the right
list passes it. A separate cold call is shown only the study, the quoted
evidence and the record. A confident "not the named target" demotes the row
to keyed_link_disputed, keeping both readings for a human.
Every accepted pooled-search link gets one of four grades —
clearly_target, likely_target,
unlikely_target, clearly_not_target. The grade sets
link_confidence and is appended to link_evidence. It
never changes the link and drops no row: graded rather than binary, because
the binary check flagged 0 rows in 200 on this class.
Every prompt this stage sends
Read from shared/prompts.py. Click one to read it whole — the
assembly and every fragment it splices, which together are exactly what the
version hash covers, so this page and the cache key cannot disagree.
One prompt per vocabulary, asking both questions. The abstract, reference-list and full-text steps all send the same target prompt — only the evidence block differs, never the task or the acceptance rule. That is why a target found from an abstract and one found from a PDF are comparable answers.
Rows that do not ship
A verdict is permanent whatever it says, but not every verdict is a row the validation import should receive. The export partitions those into named files — each one a question someone can answer, not a bin.
Verdicts, generations and re-runs
Stage 3 runs as a claimed, budget-gated tier. Each work ends in one permanent verdict row whose payload can rebuild its CSV rows offline. The verdict row is the checkpoint — not a file, and not a position in a CSV.
Re-extract works you name. Adds to the worklist — it re-admits works the checkpoint had subtracted.
Re-extract a population. Three kinds of name: a result verdict
(the work's whole ending), a link method (one row of it), or
field=value over any column of the exported row. Named, never
inferred — each changelog entry above carries the command that reopens what
it fixed.
Editing a prompt or a model mints a new extract generation, which reopens every work at once, because it changes what the pipeline asks.
A ladder change is not a generation change, deliberately.
EXTRACT_LADDER_VERSION was in the fingerprint until 2026-08-10 and
is not any more. A ladder edit reaches a population its author already knows —
ladder 23 addressed the 105 rows ladder 22 was measured over — while reopening
every settled work costs a whole campaign's wall clock for 3,025 works. So the
reopen is named on the command line rather than inferred.
Re-asking a changed prompt over one population
Editing a prompt looks like it has two settings — reopen everything, or leave the old answers standing. There is a third, and it is the one to reach for when an edit only addresses rows you can already name. It is two declarations used together, and neither works alone.
Editing a prompt moves the generation, and that by itself reopens every
work. Declare the new generation as accepting the old one's verdicts —
_GENERATION_EQUIVALENCES[new] = (old,), keyed by the current
generation so a later edit matches nothing and reopens strictly.
--redo-status adds just the works you want back to the
worklist. Without step 1 you reopen everything; without step 2, nothing.
The two older forms of --redo-status describe what the
ladder did. The third describes what a prompt concluded, which
is the population a prompt edit actually reaches:
| Form | Names | Reach for it when |
|---|---|---|
| a result verdict | the work's whole ending | the change reaches an ending — no_original_found,
api_error |
| a link method | one row of a result | the change reaches how an original was found — a ladder bump |
| field=value | any column of the exported row | the change reaches what a model concluded |
.venv/bin/python -m extract.tier --run --redo-status outcome=cannot_be_determined,abstract_r=
Every row the outcome prompt could not settle, plus every row it had nothing to read. Values match the rendered row, case-insensitively; an empty value means blank. A work matches when any of its rows does.
--mode validation first (real verdicts the live export ignores),
read the implementation of any worklist-changing flag, and run the same command
without --run to see what it would buy. All three rules were written
after they were broken.The export
data/extracted.csv is the end of this repository. It is written
whole, once — sorted, through a temp file and one rename — by
the only writer there is. Nothing appends to it.
The export renders only the works the named release put in an admitted pile, and drops the ones the current screen discards. A verdict outlives the routing that bought it: a work today's rule book discards would otherwise keep reaching the validation import forever. The FLoRA and validated skip lists are applied at render too, so a paper that enters FLoRA after extraction stops shipping.
Where the pipeline stops. Human validation lives in the
separate flora-validation repo; its csv_to_db.py
reads this file and writes the Supabase validation tables. This repository only
ever reads those tables back, for the dashboard.