How the pipeline works

Every rule, prompt and pile on this page is read from the code that runs, not written alongside it. Edit a prompt and this page changes with it β€” so what you read here is what the pipeline actually does.

What the pipeline does

FLoRA Extractor finds replication and reproduction studies for the FLoRA database. For each candidate paper it works out which original study the paper re-tested and what the outcome was. Four stages, each narrowing the set.

Code structure

Every module below, with the first line of its own docstring. The pipeline is five packages: four stages and the shared services they all use.

Reading the modules…

Stage 1 β€” discovery

A one-off scan of the full OpenAlex snapshot (725 GB) keeps every paper that passes a keyword gate. The result is the survivor pool β€” a few million rows of parquet, shared through a private Hugging Face repository rather than re-scanned. You almost never run this stage; you pull the pool.

The gate is a keyword filter, deliberately wide: Stage 1 is tuned for recall, because a paper it drops can never be recovered by any later stage. Precision is Stage 2's job.

Why the pool is not the corpus. The scan admits on keywords alone. Most of what survives is not a replication β€” it is everything whose title or abstract mentions one. Stage 2 exists to sort that.

The gate's actual vocabulary, what each arm admitted, where the pool and its overlay are stored, and how Stage 2 streams them: the Stage 1 tab.

Stage 2 β€” routing

A rule book of JSON specs sorts the pool into piles. The governing constraint is short and load-bearing:

Rules may only route or discard. Only an LLM screen may admit.

A rule that matches sends a work to a pile. Works nothing matches, and works with no text to match against, end in pending with a reason.

Piles

Reading conventions.json…

Why a work ends in pending

Reading conventions.json…

The rule book

Every spec the engine loads, in precedence order. Each carries its own record of what it was measured against and what was rejected while writing it β€” that is the material you need to change one responsibly.

Reading filter/spec…

The screen

The only step that may admit a paper. Two models vote on each work; a paper proceeds unless the votes say it is not a replication. Screening runs as a claimed job β€” each machine claims a batch first, so two machines never pay for the same paper.

The prompt it sends is build_classify_prompt, in full below.

Stage 3 β€” extraction

For each admitted paper, a ladder of steps tries to find the original study, cheapest step first, and returns at the first one that resolves. The same reading also codes the outcome. Each paper ends in one permanent verdict; a separate export renders those verdicts as data/extracted.csv.

Two outcome vocabularies. A replication is coded on one axis (successful, failed, mixed, and so on). A reproduction is coded on two independent axes β€” computation and robustness β€” and its outcome is the two settled values joined. They are different scales and are never summed.

The resolution ladder

The ordered list of resolution steps, and the running record of why each step behaves the way it does.

Reading link_original.py…

Every LLM prompt

Every prompt the pipeline can send, with the version hash that keys its cache. A builder assembles its prompt from spliced fragments at call time, so both the assembly and each fragment are shown β€” together they are exactly what the version hash covers, so this page and the cache key cannot disagree.

Reading shared/prompts.py…

What can be changed

Three things are meant to be edited, and each has a different blast radius.

Keywords β€” the Stage 1 gate
Widening the gate means re-scanning 725 GB, so it changes only between campaigns. A narrower gate silently loses papers no later stage can recover.
Rules β€” filter/spec/*.json
Cheap to change and cheap to measure. Editing any spec moves the bundle hash, which mints a new release β€” the store refuses to mix a release with a spec bundle that has moved. Record what you measured in the spec's own measured block; every rule above shows its predecessor's.
Prompts β€” shared/prompts.py
The most expensive lever. Editing a prompt changes its version hash, which invalidates exactly its cached answers and mints a new extraction generation β€” reopening every work at once. Do it knowingly. Where an edit provably cannot change an answer, the old and new cache keys can be declared equivalent so the existing answers keep serving.