Systematic reviews decide which treatments doctors use, which policies governments fund and which findings a field believes. The slowest part is screening: reading thousands of titles and abstracts against criteria written before the search. The criteria are the part hunch runs.
Try it first
A review of Wilson disease, a rare disorder in which copper builds up in the body, asked which of four drugs works best. Its rule began: “We included WD patients of any age or stage. The study drug had to be one of four established therapies, namely DPen, trientine, TTM or Zn.” It excluded, among others, “abstract-only publications”. What decides a paper is often not in its title. Whether a hospital’s twelve years of patients compared the drugs is in the full text; that a paper is a letter is in its metadata. Screening doesn’t decide what goes in the review; it decides what gets read.1. Get the papers
fetch.py adds what a screening tool shows: each paper’s type (article, review, letter, …) and journal. Abstracts can’t be republished, so the first command shows a legal note and fetch.py rebuilds them on your machine only, with PubMed filling gaps. Some papers still have only a title.
2. Paste the review’s criteria into a question
prototype/examples/screening/anxiety.yml
3. Read in hunch’s order
Sort the pile by how sure hunch is that a paper could qualify, and read from the top. Drag the line to see how many of the papers that ended up in the review you have found: In the anxiety review, all 72 are in the first 30% of the pile. The other 6,900 papers could have gone unread without losing one of them.4. Decide when to stop
A reviewer can’t see the curve; they don’t know which papers are in the review until they’ve read them. So the stopping rule has to be set in advance, and checked.act says how sure a “no” must be to set a paper aside unread, and min_recall says how many of the screeners’ papers must still be read:
prototype/examples/screening/anxiety.yml
The lesson
The tool this field already uses is active learning: a model that watches each screening decision and moves papers like the ones kept up the pile. It is a strong baseline. On the anxiety review it beats hunch everywhere: 95% of the papers in the review after 17% of the pile, against hunch’s 28%. On the diet review hunch wins everywhere: 30% against 52%. On Wilson disease it’s split. Active learning finds what the screeners kept sooner (95% after 44–51% of the pile, against hunch’s 77%); hunch finds the papers that ended up in the review sooner (95% after 28%, against 37%). The difference is what each one learns from. Active learning copies the screeners, including the papers they keep just in case. hunch reads the criteria, before anyone has screened a paper. When the two disagree, as on Wilson disease, the disagreement is worth reading: it is where the screeners’ habits and the written rule part.Use it on your data
Export your search results with titles and abstracts, paste your criteria intoinstructions, screen a few hundred papers to set the bar, and hunch test tells you how much you can safely leave unread.
How it was measured
The data
The data
Three reviews from SYNERGY (De Bruin et al. 2023, CC0): van Dis et al. 2020 on cognitive behavioural therapy for anxiety (9,883 papers, 689 kept at screening, 72 in the review); Moran et al. 2020 on whether poor nutrition makes animals take more risks (5,244; 619; 142); Appenzeller-Herzog et al. 2019 on drugs for Wilson disease (2,896; 146; 26).
gold is the screeners’ title-and-abstract decision; included whether the paper ended up in the review, which every such paper had passed at screening. Each paper also carries its type and journal from OpenAlex, which hunch sees and active learning doesn’t (it reads title and abstract). OpenAlex no longer has most abstracts; fetch.py fills gaps from PubMed, leaving 90%, 71% and 63% of the papers with one. Each spec quotes the review’s criteria as SYNERGY records them. This page shows paper titles (OpenAlex metadata, CC0) and quotes no abstract. A first run without type and journal cost $0.73; it said yes to letters and reviews the criteria exclude, which an adversarial review caught. A fourth review, on class size in schools, was dropped because PubMed covers almost none of its papers.The results, in full
The results, in full
From
measure.py; the share of the pile read, in order of hunch’s p(yes), with 95% intervals from 1,000 bootstrap resamples.Active learning’s ranges span its three runs. With few papers in each review (72, 142, 26), the last one or two decide the “all” column: active learning’s last paper on Wilson disease and on diet has no abstract, so read the “95%” and “all but one” columns as the steadier comparison. Work saved at 95% of the screeners’ papers (WSS@95) in hunch’s order: 57.7%, 50.1%, 17.7%. Papers with only a title are harder: in hunch’s order, 95% of the screeners’ title-only papers take 63%, 76% and 72% of the title-only pile, against 36%, 32% and 93% with an abstract (on Wilson disease the screeners kept many papers whose abstracts the criteria rule out).
How active learning was run
How active learning was run
ASReview 3.0.8’s simulator (
baseline.py), its default model (elas_u4), starting from one kept and one rejected paper, three times from different starting papers per review; it learns from the screeners’ decisions, hunch from none, and it stops once it has found every paper the screeners kept. Starting it instead from hunch’s ten likeliest papers, as a reviewer who screens those first would (8, 10 and 8 of them kept), changed almost nothing: 95% of the papers in the review after 16.4%, 52.2% and 36.0% of the pile, all of them after 28.5%, 70.8% and 72.3%. The runs are local and free.Wilson disease: where hunch and the screeners part
Wilson disease: where hunch and the screeners part
Reading in hunch’s order, 95% of what the Wilson disease screeners kept takes 77% of the pile, and a recall bar for them would mean reading almost everything, so that spec has none. Among the papers they kept are records of conference poster sessions and cohorts of patients followed over time, which might compare the drugs and can only be told apart by the full text. hunch mostly says no to both; among the papers in the review, the ones it was least sure of (p(yes) 0.06 to 0.11) are such cohorts and one OpenAlex types as a review. Still, 25 of the 26 are in the first 28% of the pile, and all of them in the first 39%.
The recall bars were chosen on half the papers
The recall bars were chosen on half the papers
act for “no” is the lowest bar that keeps 95% of the screeners’ papers on one half of each review (split by a hash of the paper’s id), then checked on the other half: anxiety 0.96 (95.3% kept, 40.2% read, none of the 38 papers in the review set aside), diet and risk 0.96 (96.1%, 45.8%, none of 58). The min_recall checks in the specs run on all papers, so they are regression bars; the half-split numbers are the honest ones.Limits
Limits
Three reviews, all from one collection. The criteria are those the reviews published, which can be sharper than the protocol the screeners worked from. The papers and reviews are published online, and the model may have seen some of them. Screening decisions are only as good as the screeners: a paper they wrongly skipped counts against nobody.
Cost
Cost
All three reviews, 18,023 papers, with type and journal: 0.41 in 165 seconds). Before that, 0.04 for a 300-paper trial of each: 0.