Duplicate Detection in the Screening Tool

How automatic duplicate detection works in the screening tool: DOI and title matching, measured accuracy figures, and how to review or restore any merge.

How automatic duplicate detection works

Every import runs duplicate detection automatically, across files and across import batches. Two signals drive it. First, DOI matching: DOIs are normalized (URL prefixes and doi: labels stripped) and records sharing a DOI are merged. Second, title matching: titles are normalized by lowercasing, removing punctuation and stop words, and smoothing common variants such as hyphenation differences, so that case changes, punctuation, truncated subtitles, and spellings like post-traumatic versus posttraumatic still match.

Duplicates are never silently deleted. The duplicate copy is excluded with the reason duplicate and kept in the project, linked to the record it matched, so every merge is inspectable and reversible.

Measured accuracy

The de-duplicator's accuracy is measured, not claimed. Against a 300-record benchmark of distinct same-topic studies, the hardest kind of false-merge bait, plus 115 synthetic duplicate variants of the kinds seen in real database exports, it made zero incorrect merges (100 percent precision) while catching 97.4 percent of the duplicate variants. The full methodology is published in the measured accuracy section of the tool page.

Zero wrong merges is the number that matters most for a systematic review: a missed duplicate costs you a minute of screening, but a wrong merge silently deletes a unique study from your review.

Review the duplicates

The PRISMA sidebar shows a running duplicate count and a Review duplicates button once any exist.

  1. Click the Review duplicates button in the sidebar (it shows the current count, for example Review 12 duplicates).
  2. Each row shows the removed duplicate next to the record it matched, so you can compare titles, years, sources, and DOIs side by side.
  3. If a merge is wrong, click Not a duplicate on that row. The record is restored to the screening list immediately.
  4. Close the panel when you are satisfied. The remaining duplicates stay excluded with the reason duplicate and are reported in your PRISMA counts.

Tips

  • Restoring a record is always safe: it simply returns to the undecided list for normal screening.
  • Duplicate counts flow straight into the PRISMA 2020 flow diagram, so review them before you export the figure.

Import your search export and try this workflow on your own review.

Start screening your records