How to cite the Systematic Review Screening Tool

Journals and PRISMA 2020 reporting expect you to name the software used for study selection. This page gives you copy-ready citations for the tool, methods-section paragraphs you can paste and adapt, and a citation for the published accuracy benchmark. The tool is published by Research Gold as a corporate author; there is no public version number, so none is cited.

Cite the tool

Pick the format your target journal uses. Replace nothing: publisher, year, and URL are all current.

APA 7 (software)
Research Gold. (2026). Systematic Review Screening Tool [Computer software]. https://researchgold.org/systematic-review-screening-tool
Vancouver
Research Gold. Systematic Review Screening Tool [software]. Research Gold; 2026. Available from: https://researchgold.org/systematic-review-screening-tool
BibTeX
@software{researchgold2026screening,
  author    = {{Research Gold}},
  title     = {Systematic Review Screening Tool},
  year      = {2026},
  publisher = {Research Gold},
  url       = {https://researchgold.org/systematic-review-screening-tool}
}

Methods section text you can paste

Two ready-to-adapt paragraphs, written in past tense and third person for a journal methods section. Text in square brackets, such as [two] reviewers, is a placeholder you must replace with the specifics of your review. Use the second variant if you ran AI screening suggestions or relevance ranking; it discloses the assistance transparently and makes human oversight of every decision unambiguous.

Variant 1: standard dual screening
Records identified by the database searches were imported into the Systematic Review Screening Tool (Research Gold, 2026; https://researchgold.org/systematic-review-screening-tool). Duplicate records were detected automatically by DOI and title matching, verified manually, and removed. [Two] reviewers then independently screened all titles and abstracts against the predefined inclusion and exclusion criteria. Disagreements were resolved by discussion, with [a third reviewer] adjudicating where consensus could not be reached. The full texts of all potentially eligible reports were retrieved and assessed independently by [two] reviewers, and the reason for each full-text exclusion was recorded in the tool. Record counts at each stage were exported from the tool and are reported in the PRISMA 2020 flow diagram.
Variant 2: AI-assisted screening with transparent disclosure
Records identified by the database searches were imported into the Systematic Review Screening Tool (Research Gold, 2026; https://researchgold.org/systematic-review-screening-tool). Duplicate records were detected automatically by DOI and title matching, verified manually, and removed. An artificial intelligence assistant built into the tool (a large language model) provided advisory screening suggestions and ranked records by relevance; in the tool's published benchmark evaluation across five public screening datasets the assistant kept 99% of relevant records for human review, records without abstracts are never suggested for exclusion, and every suggested exclusion is re-checked by an independent recall-focused second pass (Mitchell et al., 2026). All suggestions were advisory only: [two] reviewers independently screened every title and abstract against the predefined inclusion and exclusion criteria, every final inclusion and exclusion decision was made by the human reviewers, and no record was excluded on the basis of an artificial intelligence suggestion alone. Disagreements were resolved by discussion, with [a third reviewer] adjudicating where consensus could not be reached. Full texts of potentially eligible reports were assessed independently by [two] reviewers, with reasons for exclusion recorded, and record counts at each stage are reported in the PRISMA 2020 flow diagram.

Cite the benchmark

The accuracy figures quoted in the methods text come from a replayable benchmark on five published screening corpora: the van de Schoot 2018 post-traumatic stress disorder review (SYNERGY collection), the Smid 2020 Bayesian statistics review, and the Hall 2012, Wahono 2015, and Radjenovic 2013 software engineering reviews, more than 28,000 records in total, hydrated from OpenAlex. On ~200-record labeled samples screened against each review's eligibility criteria, the AI screening assistant kept 229 of 231 relevant records for human review (99.1% recall, 100% on three corpora), with a 3.5% average false-include rate (0% on the PTSD corpus). By design it never suggests excluding a record that has no usable abstract, and every exclude suggestion is re-read by an independent second recall pass. The active-learning ranker saved 69 to 93% of screening work at 95% recall (WSS@95) across the five corpora, with 82 to 99% of relevant records found in the top 10% of the ranked pile; records missing abstracts rank poorly, so corpora with many title-only records see lower savings. Duplicate detection scored 100% precision (zero false merges) and 97.4% recall on a 300-record same-topic hard-negative benchmark with 115 synthetic variants. Full details are in the measured accuracy section of the tool page.

Benchmark report citation
Mitchell, S., Okonkwo, D., & Vasquez, E. (2026). Benchmark report: artificial intelligence screening accuracy and ranking performance of the Research Gold screening tool. Research Gold. https://doi.org/10.5281/zenodo.21158871

The full benchmark report is archived on Zenodo: doi.org/10.5281/zenodo.21158871.