Stage 3 — linking and coding

For each paper the screen admitted, answer two questions: which original study did this re-test, and what happened. This is the stage that produces the actual research output, and the only stage that reads whole papers.

Where Stage 3 sits

Stage 3 takes the works the screen passed and turns each into one permanent verdict; a separate export renders those verdicts into the CSV that leaves this repository.

Reading the pipeline counts…

The two questions

They are answered together but judged separately. A paper's target can be certain while its outcome is not, and the reverse never happens — an outcome is only ever coded against a link the pipeline accepted.

the link

Which original? Resolved by a ladder of steps, cheapest first. Writes doi_o, title_o, authors_o, study_o and the link_method that found them.

the outcome

What happened? Coded from the same reading that found the link. A replication gets one axis; a reproduction gets two.

One row per replication–original pair. A paper that re-tests three different originals becomes three rows. Several studies of one original stay one row, with the study numbers collapsed into study_o.

Reading the model configuration…

The resolution ladder

Steps in cost order. The row takes the first one that resolves and stops there, so a paper whose title states its target never costs a model call, and only the papers that resist everything cheaper reach the last step and get their full text downloaded. The counts are how many shipped rows each step actually produced, which is the honest picture of where the work happens.

Reading link_original.py…

Why the search steps sit at the bottom. OpenAlex bills by request shape, and the spread is three orders of magnitude: a lookup by DOI is free, a filter query is 1×, a free-text search is 10×, and a content download is 100×. So a title search is the single most expensive thing a row can do short of downloading the paper — which is exactly why it is tried after everything a cached candidate list can answer.

What the ladder has learned

The running record kept beside the ladder version — newest first. Each entry is a change to how an original is found, and why.

Reading the changelog…

When a step does not end the row

A step ends the row only when it both resolved the link and settled the outcome. This is the single most confusing thing about the ladder, and it exists for a concrete reason: the abstract often names the original clearly while saying nothing about how the replication turned out.

Reading OUTCOME_DESCENT…

So the ladder carries the accepted link and keeps descending towards the closing sections that state the verdict. When the full-text step does answer, its reading replaces the carried one — except that a later unsettled outcome never overwrites an earlier settled one.

A carried link outranks a withheld rule pick at every no-answer exit. No PDF, no document, no context, an incomplete screen, a provider failure — an outage below an accepted link no longer writes target_pending over it. A provider failure is not an answer, so it does not restore a withheld pick either.

Getting the document

The full-text step needs the paper. A waterfall tries every source in priority order and stops at the first that returns something that passes its own content check.

Reading extracted.csv…

A document need not be a PDF, but every source has a content check a record page fails. Four sources hand back a sections dict instead of a file, and each is paired with the test that says whether what came back is a document at all. A repository landing page restates the abstract and adds citation chrome; coding a row from one looks like full text and is not — so the HTML check subtracts the abstract rather than thresholding the total. Measured 2026-08-07: five landing pages carried 0–1,706 characters beyond their abstracts, three real full texts carried 49,193–71,641.

A result that fails its check is no document: the row ends at no_fulltext_available and the failure is never cached as a success. Blank pdf_source on a shipped row is not a gap — it means the row resolved before the full-text step was ever needed.

Coding the outcome

Two vocabularies, because a replication and a reproduction are different questions. They are different scales and are never summed.

Reading the schema…

uninformative is the authors' verdict; cannot_be_determined is ours. The first says the study itself could not settle the question. The second says we could not tell from what we read.

A reproduction's outcome is its two settled axis values joined, which is why the flat outcome list hides which axis failed. technical failure exists on the computation axis because the most common real reproduction outcome in economics is being defeated by the materials — no data, no code, an unrunnable workflow — and the old grid could only record that as a numerical disagreement nobody observed.

Checking its own work

Three independent checks run before a row is allowed to settle, each aimed at a different way a link can be wrong.

DOI verify

The metadata doi_o actually points to is fetched (CrossRef, then OpenAlex) and compared with the resolved record. Mismatches are re-resolved from title and author in three tiers, strictest first — a wrong correction is worse than a flag. Runs once, inside the tier, and the answer is stored on the row.

keyed confirm

DOI verification checks the record against its own metadata, never against the target the paper named — so the wrong entry picked from the right list passes it. A separate cold call is shown only the study, the quoted evidence and the record. A confident "not the named target" demotes the row to keyed_link_disputed, keeping both readings for a human.

search grade

Every accepted pooled-search link gets one of four grades — clearly_target, likely_target, unlikely_target, clearly_not_target. The grade sets link_confidence and is appended to link_evidence. It never changes the link and drops no row: graded rather than binary, because the binary check flagged 0 rows in 200 on this class.

Reading extracted.csv…

Every prompt this stage sends

Read from shared/prompts.py. Click one to read it whole — the assembly and every fragment it splices, which together are exactly what the version hash covers, so this page and the cache key cannot disagree.

One prompt per vocabulary, asking both questions. The abstract, reference-list and full-text steps all send the same target prompt — only the evidence block differs, never the task or the acceptance rule. That is why a target found from an abstract and one found from a PDF are comparable answers.

Reading shared/prompts.py…

Rows that do not ship

A verdict is permanent whatever it says, but not every verdict is a row the validation import should receive. The export partitions those into named files — each one a question someone can answer, not a bin.

Reading the set-aside files…

Verdicts, generations and re-runs

Stage 3 runs as a claimed, budget-gated tier. Each work ends in one permanent verdict row whose payload can rebuild its CSV rows offline. The verdict row is the checkpoint — not a file, and not a position in a CSV.

--redo

Re-extract works you name. Adds to the worklist — it re-admits works the checkpoint had subtracted.

--redo-status

Re-extract a population. Three kinds of name: a result verdict (the work's whole ending), a link method (one row of it), or field=value over any column of the exported row. Named, never inferred — each changelog entry above carries the command that reopens what it fixed.

a new generation

Editing a prompt or a model mints a new extract generation, which reopens every work at once, because it changes what the pipeline asks.

A ladder change is not a generation change, deliberately. EXTRACT_LADDER_VERSION was in the fingerprint until 2026-08-10 and is not any more. A ladder edit reaches a population its author already knows — ladder 23 addressed the 105 rows ladder 22 was measured over — while reopening every settled work costs a whole campaign's wall clock for 3,025 works. So the reopen is named on the command line rather than inferred.

Re-asking a changed prompt over one population

Editing a prompt looks like it has two settings — reopen everything, or leave the old answers standing. There is a third, and it is the one to reach for when an edit only addresses rows you can already name. It is two declarations used together, and neither works alone.

1 · hold

Editing a prompt moves the generation, and that by itself reopens every work. Declare the new generation as accepting the old one's verdicts — _GENERATION_EQUIVALENCES[new] = (old,), keyed by the current generation so a later edit matches nothing and reopens strictly.

2 · name

--redo-status adds just the works you want back to the worklist. Without step 1 you reopen everything; without step 2, nothing.

The two older forms of --redo-status describe what the ladder did. The third describes what a prompt concluded, which is the population a prompt edit actually reaches:

FormNamesReach for it when
a result verdict the work's whole ending the change reaches an ending — no_original_found, api_error
a link method one row of a result the change reaches how an original was found — a ladder bump
field=value any column of the exported row the change reaches what a model concluded
.venv/bin/python -m extract.tier --run --redo-status outcome=cannot_be_determined,abstract_r=

Every row the outcome prompt could not settle, plus every row it had nothing to read. Values match the rendered row, case-insensitively; an empty value means blank. A work matches when any of its rows does.

A column that does not exist is refused, not silently matched against nothing — a typo that reopens zero works reads exactly like "already fixed". And the equivalence is a claim that every work you did not reopen would still get its recorded answer: make it deliberately, in a comment beside the entry.
Before any run that spends: exercise changed code in --mode validation first (real verdicts the live export ignores), read the implementation of any worklist-changing flag, and run the same command without --run to see what it would buy. All three rules were written after they were broken.

The export

data/extracted.csv is the end of this repository. It is written whole, once — sorted, through a temp file and one rename — by the only writer there is. Nothing appends to it.

Reading extracted.csv…

The export renders only the works the named release put in an admitted pile, and drops the ones the current screen discards. A verdict outlives the routing that bought it: a work today's rule book discards would otherwise keep reaching the validation import forever. The FLoRA and validated skip lists are applied at render too, so a paper that enters FLoRA after extraction stops shipping.

Where the pipeline stops. Human validation lives in the separate flora-validation repo; its csv_to_db.py reads this file and writes the Supabase validation tables. This repository only ever reads those tables back, for the dashboard.

Running it