Selecting papers for a review without reading all of them
In short: export every search hit as a .bib file, let a language model screen titles and abstracts against written criteria, then hand-check a random sample to see how often it's right. The model…
- published
- read time
- 5 min
- words
- 939
- lang
- en
- filed under
- Research
In short: export every search hit as a .bib file, let a language model screen titles and abstracts against written criteria, then hand-check a random sample to see how often it's right. The model saves the reading. The sample tells you whether you can trust it.
A systematic review starts with a search that returns far more papers than anyone wants to read. The first pass, title and abstract screening, is mostly mechanical: does this paper meet the inclusion criteria or not? It's slow, it's boring, and by paper three hundred your judgement is not what it was at paper ten.
I built a small screening tool for this, and a set of extraction prompts for the stage after it, which are public as LitReviewGPT. Here's how the pipeline works, what the first version of the tool got wrong, and how I check whether the screening is any good.
The pipeline
- Export. Most databases you'd search can export to BibTeX. Merge the exports into one
.bibfile with titles and abstracts. - Screen. For each entry, send the title, the abstract and your inclusion and exclusion criteria to a model, and ask for a decision.
- Filter. Write the included entries back out as a new
.bib, so it drops straight into a reference manager or a LaTeX project. - Check. Draw a random sample from the whole pool, screen it yourself without looking at the model's answers, and compare.
- Extract. For the papers that survive, ask the same structured questions of every full text: study type, number of participants, sensors, methods, limitations. Missing answers come back as NA rather than a guess.
What the first version got wrong
The first version of the screening tool was a Gradio app. Upload a .bib, paste your criteria, paste an API key, and get back a filtered file and a count of included papers out of the total. It worked, and it taught me where this kind of tool breaks.
- It parsed the answer by looking for the word "yes". Anywhere in the reply. A model that writes "No, although yes, it uses wearables" counts as an include. Free-text answers need a strict format, not a substring check.
- It kept no reason. The decision was a boolean. When I disagreed with it, I had no way to see which criterion it had applied, so I couldn't fix the criteria.
- Papers without abstracts were judged on the title alone. The prompt said "No abstract available" and the model decided anyway. Those should go to a human, not to a guess.
- There was no measurement. It told me how many papers it kept, not how many it kept correctly.
The fixes are small: ask for JSON with a decision, the criterion that decided it, and one sentence of reason. Add a third answer, unsure, and send those to a person. Then measure.
import json
from openai import OpenAI
client = OpenAI()
def screen(title, abstract, criteria, model="gpt-4o-mini"):
if not abstract:
return {"decision": "unsure", "criterion": None, "reason": "no abstract"}
prompt = (f"Criteria:\n{criteria}\n\nTitle: {title}\nAbstract: {abstract}\n\n"
'Answer in JSON: {"decision": "include" or "exclude" or "unsure", '
'"criterion": "the rule that decided it", "reason": "one sentence"}')
reply = client.chat.completions.create(
model=model, temperature=0,
response_format={"type": "json_object"},
messages=[{"role": "user", "content": prompt}])
return json.loads(reply.choices[0].message.content)
def precision_recall(pairs):
# pairs: (model_included, i_included) for each paper in the hand-checked sample
tp = sum(m and h for m, h in pairs)
fp = sum(m and not h for m, h in pairs)
fn = sum(h and not m for m, h in pairs)
return tp / (tp + fp), tp / (tp + fn)
Measuring it on a sample
You don't need to read every paper to know whether the screen works. You need a random sample from the whole pool, screened by you, blind to what the model said. Then two numbers. Precision: of the papers the model included, how many should have been included. Recall: of the papers that should have been included, how many the model kept.
The table shows how the arithmetic works, with illustrative numbers rather than results from a specific review: a hundred papers sampled, checked against a first and a tightened set of criteria.
| Criteria | Sampled | Model included | Correct includes | Missed includes | Precision | Recall |
|---|---|---|---|---|---|---|
| First draft | 100 | 30 | 18 | 2 | 0.60 | 0.90 |
| Tightened | 100 | 22 | 18 | 2 | 0.82 | 0.90 |
In screening, the two numbers don't matter equally. A false include costs you a few minutes, because you'll read that paper later and drop it. A missed include can quietly change the conclusion of the review. So I care about recall first, and use precision to decide whether the model is saving me enough reading to be worth it. If recall is low, the criteria are usually too narrow or too vague, and the reasons the model gives tell you which.
Where it still needs a person
The model doesn't replace the second reviewer that most review protocols ask for. It replaces the first exhausted pass. Anything marked unsure, anything without an abstract, and every paper in the final set still gets read by a human. And if you change the criteria, you check a fresh sample. The old numbers no longer apply.
If you have a review coming up, try this before you read anything: write your criteria as a numbered list, screen a random fifty papers yourself, and run the model on the same fifty. If recall is high and the reasons make sense, let it screen the rest. If not, the disagreements will show you exactly which criterion to rewrite.
related