QUADAS-2: Risk of Bias Tool for Diagnostic Accuracy Studies
QUADAS-2 explained: four risk-of-bias domains, signaling questions, applicability concerns, judgment rules, comparison to QUADAS-C and QUADAS-AI, and reporting.
Dr. Sarah Mitchell
May 14, 2026
Key Takeaways
QUADAS-2 is the consensus risk-of-bias tool for primary diagnostic accuracy studies, organized into four domains and rated per domain rather than as a composite score.
Each domain has signaling questions that guide the reviewer toward a low, high, or unclear judgment without mechanical scoring.
Three of the four domains also receive a separate applicability concern rating that asks whether the study matches the review question.
QUADAS-C extends the framework to comparative diagnostic accuracy studies and QUADAS-AI is in development for artificial intelligence diagnostic tools.
Two independent reviewers should complete the assessment with a documented adjudication process, and judgments must be reported per study and per domain in the published review.
QUADAS-2 is the standard tool for assessing risk of bias and applicability in primary diagnostic test accuracy studies included in systematic reviews. It is the second version of the Quality Assessment of Diagnostic Accuracy Studies tool, published in 2011 by Penny Whiting and colleagues on behalf of the Cochrane Screening and Diagnostic Tests Methods Group, and remains the tool endorsed by Cochrane and recommended by major medical journals for any diagnostic accuracy review.
The tool covers four risk-of-bias domains (patient selection, index test, reference standard, flow and timing) and assesses applicability concerns in the first three of those domains separately. Each domain is evaluated using signaling questions that guide the reviewer to a domain-level judgment of low, high, or unclear risk of bias. This guide unpacks the four domains, walks through the signaling questions, explains the customization step that QUADAS-2 expects every review to perform, and covers extensions such as QUADAS-C for comparative diagnostic accuracy.
The Four Domains and Why QUADAS-2 Replaced QUADAS-1
The original QUADAS tool, published in 2003 by Penny Whiting, Anne Rutjes, and colleagues, was the first widely adopted instrument for assessing diagnostic accuracy studies. It contained fourteen items scored as yes, no, or unclear. The tool quickly became standard in Cochrane diagnostic test accuracy reviews and was cited in thousands of systematic reviews over the following decade.
Practical experience with QUADAS-1 surfaced three recurring problems. First, the single composite score that some reviewers computed from the fourteen items was unreliable and not endorsed by the developers, but its use was widespread enough that interpretations diverged. Second, several items mixed risk of bias with quality of reporting, which conflated study conduct with study description and confused reviewers. Third, applicability concerns, the question of whether a study's design matches the review question, were not cleanly separated from risk of bias, leading to inconsistent ratings across reviews.
QUADAS-2 addressed all three problems with a redesign. The new tool has four risk-of-bias domains rather than fourteen items, separates risk of bias from applicability explicitly, uses signaling questions that guide rather than dictate the judgment, and explicitly recommends against any composite score. The structure is closer to what reviewers familiar with RoB 2 for randomized trials or ROBINS-I domain-level assessment expect: domain-level judgments rather than item-level scores. The quality assessment tools compared overview places QUADAS-2 in the broader landscape of risk-of-bias tools.
Domain 1: Patient Selection
The first domain asks whether the way patients were selected for the study could have introduced bias into the diagnostic accuracy estimates. Two signaling questions guide the judgment. The first asks whether a consecutive or random sample of patients was enrolled. Convenience sampling, especially when patients were selected based on prior clinical suspicion or prior test results, can inflate the apparent accuracy of the index test. The second asks whether the study avoided inappropriate exclusions, such as removing patients with difficult-to-classify disease, equivocal results, or comorbidities that complicate diagnosis.
A third signaling question, applied when the study used a case-control design, asks whether the study avoided a case-control design. Case-control studies for diagnostic accuracy almost always overestimate sensitivity and specificity because they sample only clearly positive and clearly negative patients, omitting the diagnostic uncertainty that defines real-world clinical practice. The signaling question is phrased so that "yes" indicates lower bias risk: yes, the study avoided a case-control design.
The applicability concern for patient selection asks a separate question: does the patient sample match the review question? A well-conducted study in a tertiary referral cohort may have low risk of bias but high applicability concern if the review question is about primary care. This separation is one of the most important conceptual contributions of QUADAS-2. It allows a reviewer to flag a study as well-conducted but not directly relevant, rather than penalizing it twice.
Domain 2: Index Test
The second domain covers the index test, which is the test whose accuracy the review is evaluating. The first signaling question asks whether the index test results were interpreted without knowledge of the reference standard. Knowledge of the reference standard during index test interpretation introduces interpretation bias, which can substantially inflate apparent accuracy. The blinding question is sometimes hard to answer from the published methods: many studies do not state explicitly whether the index test reader was blinded.
The second signaling question, applied when the index test has a threshold, asks whether the threshold was pre-specified. Post-hoc threshold optimization, where the study chose the threshold after seeing the data, inflates apparent accuracy because the chosen threshold maximizes performance in the study sample. Pre-specified thresholds, especially those defined in a separate derivation cohort or in published clinical guidelines, avoid this inflation. If the threshold was data-driven, the signaling question is rated "no" and the domain typically rated at high risk.
The applicability concern for the index test asks whether the test, as conducted in the study, matches the test the review is evaluating. Subtle differences (a different machine model, a different antibody, a different software version, a different reader experience level) can introduce applicability concerns even when the index test concept is the same. Reviewers should consult the diagnostic test accuracy meta-analysis framework for how to handle this kind of subtle test variation.
Domain 3: Reference Standard
The third domain covers the reference standard, which is the test or process used to determine the true disease status against which the index test is compared. The first signaling question asks whether the reference standard is likely to correctly classify the target condition. An imperfect reference standard introduces bias that propagates into the index test's estimated accuracy. A reviewer assessing a new biomarker against an established histopathology reference can usually answer this favorably; a reviewer assessing a new test against an older test with known accuracy limitations cannot.
The second signaling question asks whether the reference standard results were interpreted without knowledge of the index test. The same blinding concern that applies to the index test applies in reverse to the reference standard. If the pathologist or radiologist reading the reference standard knew the index test result, their interpretation may have been influenced, which biases the diagnostic accuracy estimate in either direction depending on the nature of the influence.
The applicability concern for the reference standard asks whether the target condition as defined by the reference standard matches the target condition in the review question. A reference standard that defines disease more broadly or more narrowly than the review intends can produce applicable-looking accuracy estimates that do not generalize. This is especially common for evolving diagnostic criteria, where reference standards in older studies may not match contemporary definitions.
Domain 4: Flow and Timing
The fourth domain covers the flow of patients through the study and the timing between the index test and the reference standard. Three signaling questions structure the judgment. The first asks whether there was an appropriate interval between the index test and the reference standard. Long intervals introduce the possibility that disease status changed between the two tests, which decouples the index test result from the reference standard truth and biases the accuracy estimates.
The second asks whether all patients received the same reference standard. Differential verification occurs when some patients receive one reference standard and others receive a different one, often based on the index test result. This introduces bias because the apparent accuracy depends on how the two reference standards classify cases differently. The third asks whether all patients were included in the analysis. Partial verification occurs when patients without a reference standard result are excluded from the analysis, which is again often correlated with the index test result. Both differential and partial verification require careful handling in the bias judgment.
Patient flow problems are the most commonly missed source of bias in diagnostic accuracy reviews. Reviewers focused on patient selection and index test blinding sometimes treat flow and timing as a checklist item rather than a substantive concern, but the bias introduced by inappropriate flow can be larger than the bias from interpretation issues. There is no separate applicability concern for flow and timing because flow is fundamentally about how the study was run rather than what the study was about.
Signaling Questions Versus Domain Judgments
A central design feature of QUADAS-2 is the distinction between signaling questions and domain judgments. Signaling questions are concrete, factual questions about study conduct (was the threshold pre-specified, were patients consecutively enrolled). Domain judgments are summary risk-of-bias ratings for the entire domain on a three-point scale: low, high, or unclear.
The signaling questions guide the reviewer toward the domain judgment but do not determine it mechanically. A reviewer assessing a study with three signaling questions answered yes and one answered unclear must apply judgment to decide whether the unclear answer is enough to push the domain to high risk, whether it is methodologically minor enough to keep the domain at low risk, or whether reporting is incomplete enough to rate the domain unclear. The tool's developers explicitly support this judgment-based approach over mechanical scoring.
This is why QUADAS-2 does not produce a numerical score. Composite scoring was a known problem with QUADAS-1 and is explicitly discouraged in QUADAS-2 documentation. Reviewers should resist the temptation to compute a weighted average of domain ratings: such scores are not validated and undermine the structured judgment the tool is designed to support. The risk of bias assessment guide discusses this design pattern across all current risk-of-bias tools.
Need expert quality assessment for your review?
Our methodologists conduct dual-reviewer risk of bias assessments, GRADE certainty ratings, and publication-ready summary tables.
QUADAS-2 is explicitly designed to be customized for the specific review question. The developers recommend that every review team review the signaling questions before using the tool and adapt them to the clinical context. Adaptation can include adding signaling questions (for example, a review of imaging tests might add a signaling question about reader experience), removing signaling questions that do not apply (a review of automated tests might remove blinding questions), or rewording signaling questions to make them more specific to the review's clinical context.
This customization is intended to be transparent and pre-specified in the protocol. A review that customizes QUADAS-2 without documenting the changes makes the resulting risk-of-bias ratings hard to compare with other reviews. A review that documents its customization in the PROSPERO record and the published protocol allows readers and other reviewers to understand exactly what was assessed.
Common customizations include adding signaling questions about specific quality issues known to plague the test under review, adding decision rules for partial verification (one common rule: treat partial verification as high risk unless the verified and unverified samples are demonstrably similar), and adding context-specific applicability concerns. A 2024 sampling of diagnostic accuracy reviews found that roughly half of reviews explicitly customized the tool, with the rest using it unmodified.
Applicability Concerns Versus Risk of Bias
The cleanest conceptual contribution of QUADAS-2 is the separation of applicability concerns from risk of bias. A study with low risk of bias may have high applicability concerns if its design (its setting, its patient population, its test variant, its target condition definition) does not match the review question. A study with high risk of bias may have low applicability concerns if its design closely matches the review question even though the study was poorly conducted.
This separation matters for two reasons. First, reviewers can present the two dimensions independently in summary risk-of-bias plots, which makes the underlying issues visible. Second, the implications for the synthesis are different: low risk of bias but high applicability concern suggests that the study is well-conducted but does not generalize to the review's target population, while high risk of bias but low applicability concern suggests that the study addresses the right question but its accuracy estimate is unreliable.
Applicability concerns apply to the first three domains (patient selection, index test, reference standard) but not to the fourth (flow and timing). Flow and timing is about how the study was run, not about whether the study matches the review question, so there is no applicability dimension. This asymmetry is sometimes confusing for first-time reviewers, who expect a fourth applicability rating but find none.
QUADAS-C, QUADAS-AI, and Future Extensions
The QUADAS team has continued to extend the tool. QUADAS-C, published in 2021, is the comparative accuracy extension. It assesses risk of bias when a primary study compares two or more index tests against a common reference standard. The four QUADAS-2 domains remain, but each gains additional signaling questions specific to comparative designs (for example, whether the two index tests were interpreted independently of each other). Comparative diagnostic accuracy reviews should use QUADAS-C rather than QUADAS-2.
QUADAS-AI, currently under development by the QUADAS team and the Methods in Medical AI working group, is the extension for studies of artificial-intelligence-based diagnostic tools. It addresses challenges specific to machine-learning models, such as separating internal validation from external validation, evaluating data leakage between training and test sets, and assessing reporting of model parameters. The extension is expected to incorporate concepts from the TRIPOD-AI reporting guideline and to be published in the next several years.
For now, AI-based diagnostic accuracy studies are typically assessed with QUADAS-2 augmented by additional reviewer-defined signaling questions about training data leakage, external validation, and reporting of model parameters. Reviewers planning an AI-focused diagnostic accuracy review should document their augmented signaling questions in the protocol and apply them consistently across studies.
QUADAS-2 Versus RoB 2 and ROBINS-I
QUADAS-2 is part of a family of risk-of-bias tools that share design principles but apply to different study types. RoB 2 assesses randomized controlled trials of interventions. ROBINS-I assesses non-randomized studies of interventions. QUADAS-2 assesses primary diagnostic accuracy studies. Newcastle-Ottawa Scale assesses observational studies and is still widely used in non-Cochrane systematic reviews although the more modern ROBINS-I family is preferred where applicable.
The tools should not be mixed within a single domain. A diagnostic accuracy review should use QUADAS-2 for all included primary studies; it should not use RoB 2 or NOS for the same studies. A review of treatment effects in randomized trials should use RoB 2 and not QUADAS-2, even if some of the trials happen to include diagnostic outcomes. The right tool for the study type produces ratings that other reviewers in the same area will recognize and accept.
There are reviews where multiple tools apply. A diagnostic accuracy review of a screening test in a randomized trial may use QUADAS-2 for the diagnostic accuracy estimates and RoB 2 for the trial-level treatment effects. An umbrella review of diagnostic accuracy systematic reviews uses neither but typically uses AMSTAR-2 for the review-level appraisal. Reviewers should map each included study to the right tool based on the study design and the outcomes used in the review. The Newcastle-Ottawa scale guide covers when older observational-study tools are still appropriate.
Reporting QUADAS-2 in PRISMA-DTA Manuscripts
The PRISMA extension for diagnostic test accuracy reviews, PRISMA-DTA (2018), specifies how QUADAS-2 ratings should be reported. The reporting expectations are: a per-study, per-domain rating of low, high, or unclear risk of bias and applicability concern; a summary plot showing the distribution of ratings across studies (the so-called traffic-light plot or weighted-bar plot); a narrative description of the most common bias concerns and how they were handled in the synthesis; and a sensitivity analysis exploring whether high-risk studies materially change the pooled accuracy estimates.
The summary plots are typically generated with the robvis package or directly within Review Manager if the review is a Cochrane Diagnostic Test Accuracy review. Robvis produces high-quality traffic-light plots and weighted summary bars in R, accepts QUADAS-2 input formatted as a spreadsheet, and exports figures suitable for manuscript submission. Recent updates to robvis have added QUADAS-C support.
The reporting expectations also include transparency about the assessment process: how many reviewers conducted independent assessments, how disagreements were resolved, whether any signaling questions were customized, and whether any specific decision rules (such as for partial verification) were applied. Most journals expect at least two independent reviewers with a documented adjudication process, mirroring the dual-screening expectation for study selection.
Common Reviewer Pitfalls
A handful of pitfalls show up in nearly every QUADAS-2 application. The first is inconsistent handling of partial verification: some reviewers treat any missing verification as high risk, others apply per-study judgment, and the same study can receive different ratings from different reviewers if no decision rule is pre-specified. Documenting the partial-verification decision rule in the protocol prevents this inconsistency.
The second is mis-rating differential verification. Differential verification is sometimes confused with partial verification, but the two are different problems. Partial verification means some patients have no reference standard; differential verification means different patients have different reference standards. Both bias the accuracy estimates, but the mechanisms and the implications for analysis are different.
The third is overlooking patient flow. Domain 4 is often the last domain to be assessed and is sometimes treated as a checklist item rather than a substantive concern. A study with consecutive enrollment, blinded interpretation, and an appropriate reference standard can still introduce substantial bias if patient flow problems mean some patients are systematically excluded.
The fourth is rating studies as unclear when they should be high risk. A reviewer who cannot determine whether the threshold was pre-specified or whether the reference standard was blinded should usually rate the domain at high risk, not unclear, because the absence of explicit reporting is itself a methodological concern. Unclear should be reserved for cases where the study cannot be assessed at all, not for cases where the relevant information was not reported.
A Worked Example
Consider a study evaluating a new biomarker against established histopathology for early detection of pancreatic cancer in a tertiary cancer center. Patient selection: the study enrolled consecutive patients with suspected pancreatic mass on imaging, applied an inappropriate exclusion of patients with biliary obstruction, and was a cohort design. Two of three signaling questions favor low bias; one (inappropriate exclusion) is concerning. Domain judgment: high risk of bias, because the excluded patients are a clinically important subgroup. Applicability concern: high, because the tertiary cancer center patient population does not match the review question about primary care screening.
Index test: the new biomarker. Blinded interpretation: yes. Pre-specified threshold: no, the threshold was optimized in the same dataset. Domain judgment: high risk of bias, because the data-driven threshold inflates accuracy. Applicability concern: low.
Reference standard: histopathology, which correctly classifies pancreatic cancer; blinded interpretation: yes. Domain judgment: low risk of bias. Applicability concern: low, because the target condition definition matches the review.
Flow and timing: interval between biomarker and histopathology was at most two weeks (appropriate); all patients received the same reference standard (yes); patients without histopathology were excluded from analysis. The exclusion is a partial verification problem. Domain judgment: high risk of bias due to partial verification.
The summary for this study: three of four risk-of-bias domains rated high, with patient selection and index test driving the rating, and one applicability concern flagged. The study should be included in the synthesis but flagged in the sensitivity analysis as a high-risk study whose contribution to the pooled estimate should be examined separately.
Frequently Asked Questions
What is QUADAS-2?
QUADAS-2 is the Quality Assessment of Diagnostic Accuracy Studies tool, version 2, published in 2011 by Penny Whiting and colleagues on behalf of the Cochrane Screening and Diagnostic Tests Methods Group. It assesses risk of bias and applicability concerns in primary diagnostic test accuracy studies included in systematic reviews. The tool covers four risk-of-bias domains and three applicability concerns, with signaling questions that guide reviewers to a low, high, or unclear judgment for each domain.
What are the four domains of QUADAS-2?
The four risk-of-bias domains are patient selection, index test, reference standard, and flow and timing. The first three also have a corresponding applicability concern that is rated separately, while flow and timing has no applicability concern because it is about study conduct rather than study scope. Each domain is assessed using signaling questions that guide the reviewer to a domain-level judgment of low, high, or unclear risk of bias.
What is the difference between QUADAS-1 and QUADAS-2?
QUADAS-1 had fourteen items scored individually, sometimes mixed risk of bias with quality of reporting, and did not separate applicability from bias. QUADAS-2 has four risk-of-bias domains and three applicability concerns, uses signaling questions that guide rather than dictate judgments, separates risk of bias from applicability explicitly, and explicitly recommends against composite scoring. QUADAS-2 is the current standard; QUADAS-1 is now considered superseded.
How do you complete a QUADAS-2 assessment?
A reviewer reads the primary study and answers the signaling questions for each domain based on what the study reports. The signaling questions guide the reviewer toward a domain-level judgment of low, high, or unclear risk of bias. Reviewers should consult the QUADAS-2 documentation for the precise wording of each signaling question, customize the tool for the specific review question, document any customizations in the protocol, and have at least two independent reviewers complete the assessment with a documented adjudication process.
When should I use QUADAS-2 versus RoB 2?
Use QUADAS-2 for primary diagnostic test accuracy studies, which evaluate the sensitivity and specificity of a test against a reference standard. Use RoB 2 for randomized controlled trials of interventions, even if those trials happen to include diagnostic outcomes. Use ROBINS-I for non-randomized studies of interventions. The choice depends on the study design and the outcomes used in the review, not on the topic alone. Some reviews use multiple tools when different included studies have different designs.
What is QUADAS-C?
QUADAS-C is the 2021 extension of QUADAS-2 for comparative diagnostic accuracy studies, where a primary study compares two or more index tests against a common reference standard. The four QUADAS-2 domains remain, but each gains additional signaling questions specific to comparative designs, such as whether the two index tests were interpreted independently of each other. Reviews of comparative diagnostic accuracy should use QUADAS-C rather than QUADAS-2.
Pro Tip
Customize the signaling questions in the protocol to reflect the specific clinical context of the review.
Pro Tip
Rate poorly reported but methodologically plausible items as unclear, but rate unreported safeguards such as blinding as high risk when the omission itself is a concern.
Pro Tip
Use the robvis-style traffic-light figure to display per-study per-domain ratings in the published review.
Pro Tip
Keep the per-domain narrative justification short: one or two sentences citing the specific reported detail that drove the rating.
Pro Tip
Pilot the customized QUADAS-2 form on three to five included studies before applying it to the full set.
Frequently Asked Questions
6
QUADAS-2 is used to assess the risk of bias and applicability of primary diagnostic accuracy studies included in systematic reviews. It evaluates each study across four domains, patient selection, index test, reference standard, and flow and timing, and produces a domain-level judgment of low, high, or unclear risk of bias. The tool is the consensus standard recommended by the Cochrane Screening and Diagnostic Tests Methods Group and is the assessment expected by editors of journals that publish diagnostic accuracy reviews.
The four domains of QUADAS-2 are patient selection, index test, reference standard, and flow and timing. Patient selection covers sampling and exclusions. Index test covers blinding and threshold pre-specification. Reference standard covers correct classification of the target condition and blinded interpretation. Flow and timing covers the interval between tests, whether all patients received the same reference standard, and whether all patients were included in the analysis.
No. The original QUADAS, published in 2003, was a 14-item checklist that produced a single quality score. QUADAS-2, published in 2011, replaced the checklist with a domain-based framework that uses signaling questions to guide judgments and explicitly recommends against composite scoring. The original QUADAS should not be used for new reviews. Always cite Whiting et al. 2011 and use QUADAS-2.
QUADAS-2 is not scored numerically. Each domain receives a qualitative judgment of low, high, or unclear risk of bias based on the signaling question responses and the reviewer's methodological assessment. Composite or summary scores are explicitly discouraged because they treat different sources of bias as numerically equivalent, which is not methodologically valid. The review reports per-study per-domain judgments rather than a single quality score.
QUADAS-2 assesses single-test diagnostic accuracy studies that compare one index test to a reference standard. QUADAS-C extends QUADAS-2 to comparative diagnostic accuracy studies in which two or more index tests are evaluated head-to-head against the same reference standard in the same patients. QUADAS-C is used alongside QUADAS-2 and adds signaling questions about between-test comparability such as identical reference standard application and blinded interpretation across both tests.
After piloting the customized form, an experienced reviewer typically spends 30 to 60 minutes per study on a single-pass QUADAS-2 assessment, with disagreement adjudication adding 10 to 20 minutes per discordant study. Faster ratings are usually a sign that the reviewer is not engaging with the signaling questions or the supplementary appendices, both of which often contain the methodological detail that drives the judgment.
Share
Found this useful? Share it with your colleagues.
Need expert quality assessment for your review?
Our methodologists conduct dual-reviewer risk of bias assessments, GRADE certainty ratings, and publication-ready summary tables.
Quality Assessment Takes Expertise. Our Team Does It Daily.
Rigorous risk of bias assessment, GRADE evaluations, and summary tables that satisfy peer reviewers. We handle the methodology so your review stands up to scrutiny.
Our promise: Free rework on synthesis or GRADE assessment if reviewers push back.
4.9 / 5Quote within a few hoursPRISMA 2020 + GRADE certaintyPhD methodologistConfidential by default
Dr. Sarah Mitchell holds a PhD in Biostatistics from Johns Hopkins Bloomberg School of Public Health and has over 15 years of experience in systematic review methodology and meta-analysis. She has authored or co-authored 40+ peer-reviewed publications in journals including the Journal of Clinical Epidemiology, BMC Medical Research Methodology, and Research Synthesis Methods. A former Cochrane Review Group statistician and current editorial board member of Systematic Reviews, Dr. Mitchell has supervised 200+ evidence synthesis projects across clinical medicine, public health, and social sciences.
Quality Assessment Takes Expertise. Our Team Does It Daily.
Rigorous risk of bias assessment, GRADE evaluations, and summary tables that satisfy peer reviewers. We handle the methodology so your review stands up to scrutiny.
A rigorous, doctoral-level guide to conducting a meta-analysis: defining the question, extracting effect sizes and their variances, choosing a between-study variance estimator, pooling, and diagnosing heterogeneity and bias.
Meta-analysis in psychology pools the effect sizes from many studies into one reliable result. Learn the definition, real examples, and how researchers run one.
Roughly 80 systematic reviews are published daily. The average takes 67.3 weeks, uses 5 authors, and costs about $141,195 in researcher time. Every figure sourced and linked.