Stage 1 — discovery
One scan of the whole OpenAlex corpus, one keyword gate, one artifact. Everything the rest of the pipeline will ever see is decided here — a paper this stage drops can never be recovered by any later stage.
Where Stage 1 sits
The whole pipeline, with the live count at each hand-off and the artifact it is read from. Stage 1 is the highlighted band: it ends at the survivor pool, and nothing downstream ever touches the snapshot again.
What Stage 1 is
OpenAlex publishes its whole corpus as column-projectable parquet on S3 —
about 725 GB. search/snapshot_scan.py streams every partition once,
applies a single vectorized gate, and writes the rows that pass into the
survivor pool. That scan takes 13–21 hours and is run once per
gate, not once per campaign.
You almost never run this stage. The pool is shared through a private Hugging Face dataset repo, so joining the project means pulling a few GB rather than re-reading 725 GB. The scan is only re-run when the gate itself changes.
Stage 1 searches and does not filter. It applies no exclusion pattern, no phrase guard and no year cut — every exclusion decision belongs to Stage 2's rule book, which reads the pool. That separation is deliberate: one rule set, in one place, over a corpus that kept everything the search found.
Stage 1 is tuned for recall. Precision is Stage 2's job.
The gate — what we search for
Two arms, unioned. A work that trips either is a survivor; nothing
else about it is judged here. Both are read live from
filter/phrase_detection.py, which is the one vocabulary Stage 1 and
Stage 2 share.
The survivor pool
What the gate kept, and what each row carries. The pool holds enough metadata to rebuild a candidate row without ever re-opening the snapshot.
What a pool row holds
The schema is _POOL_SCHEMA in search/snapshot_scan.py.
The three hit_* booleans are the gate's own record of why
each row was kept, which is what makes the per-arm counts above possible at all.
Where it is stored
Three artifacts: the parquet dataset, the provenance sidecar that says what the dataset is, and the text overlay that fills in abstracts the snapshot did not ship.
How Stage 2 reads it
Stage 2 never re-reads the snapshot and never loads the pool into memory. It streams it in batches through one function, with the overlay applied on the way past.
filter/engine/pool_reader.py
iter_pool_batches(pool_dir, overlay_dir, batch_size=50_000, aliases)
→ yields pyarrow RecordBatches, overlay text applied over the pool's own
An overlay row wins over the pool's own text, empty or not
(_apply_overlay). Every overlay row was written deliberately by a
backfill this project ran, and the rows that need replacing rather
than filling are exactly the boilerplate abstracts — "International audience",
bare keyword lists — that a fill-only overlay would leave sitting in front of
the screen's voters.
The backfill worklist is not only the no-text rows.
overlay.worklist() takes every routed row whose
pending_reason is no_text, plus every
admitted row that identifies an OSF record — text or no text. The OSF
phase does not fetch an abstract; it fetches the registration template line the
two osf-registration-* specs match on, and no abstract substitutes
for it. That is why OSF dominates the overlay's source counts above.
The pool is part of the release id
A routing release is the sha256 of six inputs. Two of them are Stage 1's, so a different pool or a different overlay mints a different release rather than silently overwriting the old one's decisions.
Running it
In practice you run the first one and never the second.