Stage 2 — routing and screening
Five million keyword survivors go in; a few thousand papers worth an LLM's full attention come out. This is the stage that decides what the project spends money on, and it is built so that a cheap rule can never make that call on its own.
Where Stage 2 sits
The whole pipeline, with the live count at each hand-off. Stage 2 is the two highlighted bands — it takes the survivor pool and produces the screened rows Stage 3 reads.
The 5 million, step by step
Every row of the pool, followed to the point where it either stops or reaches Stage 3. Each step names what removed the rows and what it cost.
Nothing here is deleted. A discarded work keeps its row in the routing table with the rule that discarded it and the text that matched, so any pile can be re-read, sampled and measured. "Discard" means "not sent to an LLM", not "removed from the corpus".
The one rule about rules
Rules may only route or discard. Only an LLM screen may admit.
This is the constraint the whole stage is built around, and it is worth being precise about what it means. A keyword rule can say "this is definitely not a paper" and drop it. A keyword rule can say "this looks worth reading" and put it in a queue. What a keyword rule can never do is say "this is a replication" — that judgment costs two model calls, and it is the only judgment that lets a row travel onward.
The bundle is therefore a whitelist: nothing is screened unless a positive rule admits it to a screening pile, and nothing reaches Stage 3 unless the screen then votes it through.
The five piles
Every routed work lands in exactly one. The pile decides what the work costs and what it can still become.
The rule book
Declarative JSON in filter/spec/, evaluated with pyarrow compute
over the pool. Precedence resolves multi-match: higher wins, and
a work matching six rules is normal — the full set is kept, because overlap
between rules is what diagnostics measure.
Two columns below are the ones to read together. Matched is how many works the rule's pattern hit. Won is how many it actually claimed — the rest were taken by a higher-precedence rule. A rule with a large gap is not broken; it is being outranked, which is usually the design.
Shadow rules are evaluated but never win. A
shadow: true spec is scored against every row and its matches are
recorded, so it can be measured on real data before it is trusted to move
anything. Promoting one is a one-line spec change — and the measurement that
justifies it is why the record exists.
Matched, won, outranked
A rule can match hundreds of thousands of works and claim far fewer. That gap is not a rule failing to classify anything — it is the engine resolving overlap, which it is supposed to do. The four counts in the table above are an exact identity, so every missing row has a named cause:
Each row of the pool is scored against every rule, and multi-match is normal — that is why overlap can be measured at all. But a work lands in exactly one pile, chosen by the highest-precedence non-shadow rule that matched it. So a rule that matched a work has three possible fates for it:
This rule was the highest-precedence match, so it chose the pile. These are the rows the rule actually decided.
This rule won the row and sent it to a screening pile — but the work had
no abstract, so the engine moved it to pending/no_text rather
than screen it blind. The rule's decision was overridden by the one policy
that outranks every rule.
Another rule with a higher precedence also matched this work and claimed it first. Nothing went wrong: broad rules are meant to be outranked by specific ones, which is what precedence numbers are for.
A shadow rule shows no wins at all — it is evaluated against every row and its matches are recorded, but it never claims one. Its Matched is a forecast: what it would take if it were promoted. That number is the measurement a promotion is argued from.
The one thing this table cannot tell you is whether a rule is
right. A rule that wins 355,211 rows into discard is
confidently deciding a third of a million works, and only sampling that pile
says whether it should have. That is what
filter.engine diagnose and the domain check below are for.
Why 89% is pending
pending is the biggest pile by a wide margin, and it is
not a rejection. It is the engine saying nothing has judged these works
yet. There are exactly two reasons, both assigned by the engine and never by a
spec.
no_text is recoverable coverage, not a verdict.
A work whose abstract is empty has said nothing, and absence of evidence must
not convert into a proceed — so a screening pile is downgraded rather than
screened blind. Supply the text through the overlay and the work routes into
its screening pile on the next run.
The downgrade has one deliberate exemption. An OSF record's title is its description — "Exact Replication of Rinck & Becker (2007)" — and probing OSF showed most of the missing text does not exist to be fetched: of 40 sampled, 22% had no description, no files, no wiki and no child components anywhere. Those rows are sent to the screen on their titles, which the classify prompt is built to handle. A row with no title either is still downgraded: there is nothing to read.
What a rule failed to govern
A rule can match thousands of works and still miss most of the population it
claims to be about. A spec may declare a domain — the rows it is
about, evaluated the same way as its match but changing no routing — so
the two can be compared after a route.
The third number is the one to read. Works inside the rule's
declared domain that the rule did not match, and that some other rule
sent to a paying pile. This measurement was written on 2026-08-08, after a
campaign paid for the failure it now names: an OSF discard matched 1,308 works
and looked healthy, while 878 registrations it should have governed were
admitted by a generic text rule instead — about 450 preregistrations each bought
a two-voter screen and a full Stage 3 extraction, and settled as
cannot_be_determined, because a preregistration reports no outcome.
The screen
The only step that may admit a paper. Two models read the title and abstract and each answer a fixed schema; a gate turns the two votes into one decision.
The gate
Defined once, as screen_gate(). It is deliberately the most
conservative rule that still discards anything:
Claimed before it spends
Screening runs as a claimed job: a machine claims a batch through the state
authority before the first voter is asked, so two machines can never pay
for the same paper. Without --run nothing is claimed, fetched or
spent — the runner prints the row count, the token-length distribution of the
abstracts it would send, and what that would cost.
Each raw response is written to disk before the verdict row that names it, and a verdict is written per work, so an interrupted run re-claims only the works that have no verdict yet.
The cheap tier is dormant
A second, discard-only tier exists over the screen_cheap pile. All
three screen_cheap specs are shadow, so no live row
reaches it. It may only discard, and only on two explicit noes — one keep, an
unrecognised label, an unreadable reply or a provider failure all pass the row
through unchanged. Its proceed never admits: it means "on to the
expensive screen".
What the screen actually sends
The exact text, read from shared/prompts.py. Click a prompt to
read it whole — the assembly and every fragment it splices, which together are
exactly what the version hash covers.
The hand-off to Stage 3
Stage 3 does not screen. It reads the screen's answer off the row it is handed, which is why the verdict has to travel with the work rather than be recomputed.
There is no hand-off file. Stage 3 builds its worklist
in process: the extract tier asks the state authority which works the
screen admitted, then rebuilds each row straight off the pool with
iter_export_rows + screen_columns — the same two
functions a CSV export would use. Writing a CSV and parsing it back would be a
third representation of the same thing, and a place for the two to drift.
A row travels on a verdict, not on a routing decision. A work the rules routed into a screening pile but no live expensive-screen run ever decided is not exported. Routing says "this deserves an LLM's attention"; only the validated pair says "this reaches Stage 3". A work whose votes are still short of a gate decision reads as incomplete and belongs with the unscreened — a half-screened row is exactly what the screened-only hand-off exists to hold back.
Verdicts are read across releases, deliberately: a verdict follows the work, and the release scopes the piles rather than the evidence. But only within the current screening generation — the hash of the voter pair and the classify prompt. Swap a voter or edit the prompt and the old answers stop counting, those works become claimable again, and they no longer steer the hand-off.
Release ids and re-runs
A routing release is the sha256 of six inputs. Anything that could move a row between piles is in there, so two runs with the same id are the same routing, and a changed input mints a new id instead of silently overwriting the old decisions.
Routing is derived data — the next route run
recomputes it from the pool and the specs. That is why a live tier verdict is
never written into the routing table: it would be erased. Verdicts are applied
where the rows leave the engine instead.