Every label in the current marketing vocabulary, human-written, human-edited, AI-assisted, AI-generated, AI-reviewed, AI-screened, human methodological oversight, collapses into two questions that journal editors actually care about: did a person compose the words, and was every machine output checked against a source document before it counted toward a finding. The threshold for declaring artificial intelligence use is not how much you used, it is whether the tool made or suggested a judgement. Under the 2025 joint position statement from Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence, you must declare artificial intelligence whenever it made or suggested judgements about study eligibility, critical appraisal, data extraction, synthesis, or certainty of evidence. Spelling, grammar and formatting help does not require a declaration. That single distinction resolves almost every real case.
Three years ago "written by experts" was an unremarkable claim. It is now a contested one, for a specific reason: the failure mode of machine-generated academic prose is invisible at the point of reading.
Walters and Wilder tested this directly. They generated short literature reviews on 42 multidisciplinary topics, producing 84 papers containing 636 bibliographic citations, then checked every citation by hand. 55 percent of the GPT-3.5 citations and 18 percent of the GPT-4 citations were fabricated entirely. Among the citations that referred to genuinely existing works, 43 percent of the GPT-3.5 set and 24 percent of the GPT-4 set still contained substantive errors: wrong dates, wrong volume and page numbers, wrong journal titles, wrong author names. Book chapter citations were fabricated at roughly 70 percent in both versions (Walters and Wilder, Scientific Reports, 2023).
A fabricated citation does not look fabricated. It carries a real researcher's name, a plausible journal, and a correctly formatted digital object identifier. It survives a read-through by a competent supervisor. It fails only when somebody retrieves the source, which is exactly the step a reader skips and a peer reviewer sometimes does not.
That asymmetry is why the vocabulary matters commercially now, and why vague labels have become a place to hide. It is also why the useful test is not about volume.
The principle that separates acceptable assistance from disqualifying assistance is verifiability, not effort or percentage.
Machine assistance is defensible where its output can be checked against a source in seconds and a human actually performs that check before the output counts. It is indefensible where the output is generative prose whose errors are invisible and compound across a document.
A screening suggestion passes this test. The abstract sits on the screen next to the suggestion, a reviewer reads both, and the suggestion is confirmed or rejected in about ten seconds. A wrong suggestion costs nothing if it was never applied. A cited sentence in a manuscript fails the test. Verifying one paragraph of generated prose means retrieving every reference and confirming each claim genuinely appears in it, which costs more time than composing the paragraph would have. This is what makes the two positions coherent rather than convenient.
| Label | What it should mean | What it forbids | Declaration needed |
|---|
| Human-written | A person composed every sentence and selected every citation from sources they read | No generative prose in the deliverable, at any stage, including "drafting then rewriting" | No |
| Human-edited | Machine-generated text that a person revised | Presenting the result as human-written | Yes |
| AI-assisted | A tool supported a task the human still performed and verified | Silent substitution of the human's judgement | Yes, if judgement was suggested |
| AI-generated | The tool produced the output; a human may or may not have checked it | Any use in a submitted manuscript's prose or reference list | Yes, and unacceptable for citations |
| AI-reviewed | A tool checked human work, for example consistency or formatting | Treating the check as an appraisal of quality or validity | Yes, if it appraised |
| AI-screened | A tool suggested include, maybe or exclude, and a human decided each record | The tool applying exclusions without a human decision | Yes, always |
| Human methodological oversight | A named, qualified methodologist is accountable for design and decisions | Using the phrase without a named accountable person | Not applicable, this is the baseline |
The last row is the one most often used as decoration. Under the joint position statement, "an author is accountable for the content, methods, and findings of their evidence synthesis, including the decision to use AI, how it is used, and its impact on the synthesis." Oversight is not a feature. It is the condition under which any of the other rows is permissible.
It would be evasive to set out a framework and not answer for ourselves against it, so here is Research Gold measured against the same standard.
The manuscripts Research Gold delivers are written by human researchers, who read and select every source they cite. No part of a deliverable is generated text.
The Research Gold screening tool runs manually from end to end, and that is the default path rather than a fallback: you import your records, deduplicate them, mark each one include, maybe or exclude yourself, fill the extraction grid by hand, and produce the PRISMA flow diagram without any machine assistance at any step. Criteria highlighting is rule-based whole-word matching with no artificial intelligence involved. Machine assistance is a separate option you switch on if you want it, metered separately, and when it is on it only suggests: nothing is applied until you accept it, and suggested exclusions are never applied in bulk. It is built against the seven conditions set out above. We do use automated systems in parts of our client correspondence.
Optional translation of non-English titles and abstracts sits in a different category again, and the distinction is worth understanding rather than guessing at. Rendering text from one language into another does not make or suggest a judgement about eligibility, appraisal or extraction, so it falls outside the declaration test quoted above. It does change which records you were in a position to assess, so the cleaner practice is to describe it in your search and selection narrative rather than in an artificial intelligence declaration.
The practical consequence is worth stating plainly. If you screen manually, which the tool is designed for, you have nothing to declare. If you switch the optional assistance on, you are inside the declaration requirement and we will supply the declaration text for it.
That is a narrower set of claims than "never artificial intelligence, anywhere," and the narrowness is the point. A blanket claim tells you nothing, because it cannot distinguish between a tool that offers an optional screening suggestion and a tool that writes your discussion section. Ask any provider, including us, which specific stages were touched and by what.
Four bodies converge on the same threshold, which makes compliance simpler than most researchers expect.
The International Committee of Medical Journal Editors added a dedicated artificial intelligence section in its January 2026 revision. Chatbots and other assisted tools "should not be listed as authors," because they cannot be held responsible for accuracy and integrity. Authors must disclose use at submission and describe it in both the cover letter and the appropriate section of the manuscript. "Humans are responsible for any submitted material that included the use of AI-assisted technologies," and authors must "carefully review and edit the AI-generated content as the output can be incorrect, incomplete, or biased." On references specifically: "Referencing AI-generated material as the primary source is not acceptable." Note the enforcement language: "Nondisclosure of AI use may require corrective action and may be construed as misconduct in some circumstances."
The Committee on Publication Ethics holds that these tools cannot satisfy authorship requirements because they cannot take responsibility for the work, and as non-legal entities cannot declare conflicts of interest or hold copyright. Use in preparing a manuscript should be disclosed in the methods or an equivalent section.
The evidence synthesis bodies are the most specific, and the most relevant if you are running a review. The joint statement requires declaration when the tool "makes or suggests judgements, such as in relation to the eligibility of a study, appraisals, extraction of bibliographic, numerical or qualitative data, synthesis of data from two or more studies, assessments of the certainty of evidence." It then states plainly that "generally, AI used to improve spelling, grammar, or manuscript structure does not need to be listed." A complete declaration names the system and version, the dates used, the purpose and which stages were affected, a justification that the tool is methodologically sound, any financial or non-financial interest in the tool, and its limitations and potential biases.
Publishers land in the same place. Elsevier requires no declaration for basic checks of grammar, spelling and punctuation, but does require one when a tool makes substantive changes. Wiley requires disclosure when artificial intelligence is used in research methodology, including study design, data collection, literature review and code development. Springer Nature requires methods-section disclosure while exempting copy editing. Writing in Medical Teacher, Cleland (2026) frames the same threshold as disclosure being warranted "whenever an AI tool materially shapes the research or the manuscript."
If you used a screening tool that suggested eligibility decisions, you are inside the declaration requirement. If you used a grammar checker, you are outside it. There is very little grey area, and guessing is the expensive option given that non-disclosure can be treated as misconduct.
Declaration text you can adapt
For a review where a tool suggested screening decisions that a human confirmed:
Title and abstract screening was performed by [number] reviewers. [System name, version] was used between [start date] and [end date] to generate a suggested eligibility decision and supporting rationale for each record against the protocol's criteria. All suggestions were reviewed by a human reviewer, who made every final inclusion and exclusion decision; no suggestion was applied automatically. [Number] records suggested for exclusion were re-checked by a second recall-oriented pass. The tool was selected because [justification]. The authors declare [no] financial or non-financial interest in the tool. Limitations include [limitation].
Adjust the stages named to match what you actually did. Add extraction and appraisal only if a tool touched them.
Title and abstract screening is the strongest case in the current evidence, and the reasoning is worth understanding rather than accepting.
Manual screening already fails. The joint statement notes an estimated 13 percent risk of falsely excluding relevant studies in abstract screening by human reviewers, and suggests that using artificial intelligence as a second reviewer could help reduce it. This reframes the question: the comparison is not machine screening against perfection, it is machine-plus-human against human alone. A recall-oriented second pass that re-checks every exclusion is defensible precisely because the baseline it improves on is imperfect and measurable.
The same logic supports deduplication of search results, where a match is verifiable at a glance, and translation of non-English titles and abstracts, which prevents the language exclusion that quietly narrows many reviews. In both cases a human sees the output next to the source.
What "a human decides" has to mean to be worth anything
That phrase now appears on almost every screening product, which has drained it of information. Treat it as a claim to be tested against the design rather than as a reassurance. Seven conditions separate a real human-in-the-loop workflow from a rubber stamp, and they are worth checking on any tool you are considering, including ours.
The workflow must run with no assistance at all. This is the first question, not the last. If the machine pass is load-bearing, if switching it off leaves you without a usable workflow, then "a human decides" is describing a formality. A tool that screens perfectly well by hand is offering a real choice. One that does not is offering a default with a confirmation step.
Nothing may be applied without an explicit acceptance. Pre-filling a field and inviting you to override it is not the same as offering a suggestion and waiting. The difference shows itself on record 1,400, when attention is gone.
Suggested exclusions must never be applied in bulk, even where inclusions can be. The two risks are not symmetric. A wrong inclusion surfaces at full-text screening, costs an hour, and leaves a trace. A wrong exclusion is invisible: the study leaves the review silently, nothing downstream detects it, and the synthesis is quietly wrong in a direction no reader can audit. A single button that applies every suggested exclusion has mispriced that asymmetry.
Uncertainty must resolve toward inclusion. A recall-oriented tool routes borderline records to "maybe" rather than "exclude" and accepts worse precision to do it. Poor precision costs you time. Poor recall costs you validity, and you will not know it happened.
The stated reason must be inspectable. A relevance score gives you nothing to check. A one-line reason written against your own eligibility criteria can be read in seconds and disagreed with, which is the entire mechanism by which oversight actually operates.
Every decision must remain attributable. Which record, which decision, which person, on what date. Without that there is no audit trail, no way to adjudicate reviewer disagreement, and nothing from which to write an honest methods section.
The suggestion must sit beside the source, never in place of it. Oversight means reading the abstract and the suggestion together. A workflow that surfaces only a verdict has removed the very thing it claims to preserve.
Held to those conditions, human-supervised screening is a defensible methodological choice rather than a convenience. Held to none of them, the human is a signature at the bottom of a page they did not read.
The practical problem is not doing the screening, it is being able to document it afterwards. Most reviewers discover at submission that they cannot reconstruct which decisions were made, by whom, or on what basis, which is precisely what a declaration requires.
The Research Gold free title and abstract screening tool is built manual-first for exactly this reason. You import from RIS, BibTeX, EndNote, PubMed or a spreadsheet, deduplicate, and mark every record yourself, with each decision attached to the reviewer who made it. Criteria highlighting is rule-based whole-word matching with no artificial intelligence involved. Run that way there is nothing to declare, because no tool made or suggested a judgement. Optional machine assistance exists if you want a second pass, metered separately, and it only ever suggests: each suggestion carries a confidence level and a one-line reason, nothing is applied until you accept it, and suggested exclusions are never applied in bulk. Title and abstract screening is free, and the PRISMA 2020 flow diagram and its counts are free for everyone.
Consult the guidance for your own review before relying on any of this: the King's College London and Oregon Health and Science University library guides both track current tool-level recommendations, and the RAISE framework guidance released in March 2026 covers stage-by-stage use. If you are choosing tooling, our comparison of artificial intelligence tools for screening and data extraction covers what each one does at which stage, our broader roundup of software for conducting a literature review covers the adjacent categories, and our comparison of research tools by stage covers the current landscape.
Where it is not, and why the failure stays invisible
Generative prose in the manuscript itself is the category with no defensible version, and it fails in four specific ways.
Fabricated and erroneous references. Covered above, and the numbers are not marginal.
Plausible but wrong numbers. A generated effect size, confidence interval or heterogeneity statistic is formatted correctly and sits in a sensible range. Nothing about it signals that it was not calculated from the extracted data. This is the failure that survives longest, because catching it means recomputing from the source studies.
A search strategy that cannot be reproduced. A systematic review's central claim is reproducibility. A generated Boolean string may look rigorous and return a different result set than the one reported, and the counts in a PRISMA flow diagram then fail to reconcile against the actual database exports. Our guide to building a reproducible search strategy covers what has to be documented for this to hold.
Attribution that cannot be traced. If nobody can say which source a sentence came from, the sentence cannot be defended in peer review.
The Oregon Health and Science University guidance, updated in August 2026, puts the current position bluntly: evidence consistently shows that generative artificial intelligence is not reliable enough to independently perform the core methodological steps of a systematic review.
There is also a distribution of consequences worth being explicit about. A retraction attaches to the author's name, permanently and publicly. It does not attach to a vendor. Anyone commissioning work is accepting that asymmetry, which is a reason to ask harder questions before commissioning rather than after.
Record what you used as you go, name the stages rather than the category, and keep the declaration text with your protocol so it is written while the details are fresh rather than reconstructed at submission.
If you are commissioning a review, our systematic review and meta-analysis services and the scoping review service both document methodology and deliverables in detail, and you can request a costed scope that sets out which stages involve any machine assistance. If you are doing the work yourself, the step-by-step guide to writing a systematic review and our library of free methodology tools cover the stages without any commitment.
Sources: Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports. 2023;13:14045. Flemyng E, Noel-Storr A, Macura B, et al. Position Statement on Artificial Intelligence Use in Evidence Synthesis Across Cochrane, the Campbell Collaboration, JBI, and the Collaboration for Environmental Evidence 2025. Campbell Systematic Reviews. 2025;21(4):e70074. International Committee of Medical Journal Editors, Recommendations, Artificial Intelligence Use by Authors, 2026 revision. Committee on Publication Ethics, Authorship and AI Tools position statement. Cleland J. When and how to disclose AI use in academic publishing. Medical Teacher. 2026;48(4).