Stage 1 — discovery

One scan of the whole OpenAlex corpus, one keyword gate, one artifact. Everything the rest of the pipeline will ever see is decided here — a paper this stage drops can never be recovered by any later stage.

Where Stage 1 sits

The whole pipeline, with the live count at each hand-off and the artifact it is read from. Stage 1 is the highlighted band: it ends at the survivor pool, and nothing downstream ever touches the snapshot again.

Reading the pool and the routing store…

What Stage 1 is

OpenAlex publishes its whole corpus as column-projectable parquet on S3 — about 725 GB. search/snapshot_scan.py streams every partition once, applies a single vectorized gate, and writes the rows that pass into the survivor pool. That scan takes 13–21 hours and is run once per gate, not once per campaign.

You almost never run this stage. The pool is shared through a private Hugging Face dataset repo, so joining the project means pulling a few GB rather than re-reading 725 GB. The scan is only re-run when the gate itself changes.

Stage 1 searches and does not filter. It applies no exclusion pattern, no phrase guard and no year cut — every exclusion decision belongs to Stage 2's rule book, which reads the pool. That separation is deliberate: one rule set, in one place, over a corpus that kept everything the search found.

Stage 1 is tuned for recall. Precision is Stage 2's job.

The gate — what we search for

Two arms, unioned. A work that trips either is a survivor; nothing else about it is judged here. Both are read live from filter/phrase_detection.py, which is the one vocabulary Stage 1 and Stage 2 share.

Reading the vocabulary…

The survivor pool

What the gate kept, and what each row carries. The pool holds enough metadata to rebuild a candidate row without ever re-opening the snapshot.

Reading parquet footers…

What a pool row holds

The schema is _POOL_SCHEMA in search/snapshot_scan.py. The three hit_* booleans are the gate's own record of why each row was kept, which is what makes the per-arm counts above possible at all.

Reading _POOL_SCHEMA…

Where it is stored

Three artifacts: the parquet dataset, the provenance sidecar that says what the dataset is, and the text overlay that fills in abstracts the snapshot did not ship.

Reading the pool directory…

How Stage 2 reads it

Stage 2 never re-reads the snapshot and never loads the pool into memory. It streams it in batches through one function, with the overlay applied on the way past.

filter/engine/pool_reader.py

iter_pool_batches(pool_dir, overlay_dir, batch_size=50_000, aliases)
    → yields pyarrow RecordBatches, overlay text applied over the pool's own

An overlay row wins over the pool's own text, empty or not (_apply_overlay). Every overlay row was written deliberately by a backfill this project ran, and the rows that need replacing rather than filling are exactly the boilerplate abstracts — "International audience", bare keyword lists — that a fill-only overlay would leave sitting in front of the screen's voters.

The backfill worklist is not only the no-text rows. overlay.worklist() takes every routed row whose pending_reason is no_text, plus every admitted row that identifies an OSF record — text or no text. The OSF phase does not fetch an abstract; it fetches the registration template line the two osf-registration-* specs match on, and no abstract substitutes for it. That is why OSF dominates the overlay's source counts above.

The pool is part of the release id

A routing release is the sha256 of six inputs. Two of them are Stage 1's, so a different pool or a different overlay mints a different release rather than silently overwriting the old one's decisions.

Reading release.py…

Running it

In practice you run the first one and never the second.