Quality Assessment Tools Compared: RoB 2, NOS, JBI, and GRADE
Choosing the wrong quality assessment tool is one of the most common methodological errors in systematic reviews. This guide compares RoB 2, NOS, JBI, and GRADE with a clear decision framework.
Dr. Sarah Mitchell
April 16, 2026
Want to try this yourself? Use our free research tools, no sign-up required.
Key Takeaways
RoB 2 is for randomized controlled trials only. It is the Cochrane standard.
NOS is for cohort and case-control studies. Aggregate scores should be interpreted cautiously.
JBI checklists cover study designs NOS and RoB 2 do not, including qualitative research.
GRADE operates at the outcome level across all studies, not at the individual study level.
Never use GRADE as a substitute for study-level bias assessment. They measure different things.
Use all four tools as free online instruments without software installation.
Quality assessment is not optional in a systematic review. Every included study must be evaluated for methodological rigor, and the tools you choose must match the study designs in your review. Selecting the wrong instrument can undermine your entire evidence synthesis and invite major revisions from peer reviewers. This guide provides a thorough side-by-side comparison of seven quality assessment tools you are most likely to encounter: RoB 2, the Newcastle-Ottawa Scale (NOS), JBI checklists, GRADE, QUADAS-2, AMSTAR 2, and ROBINS-I.
Understanding what each tool evaluates prevents the most common mistake in quality assessment: using the wrong instrument for your study design.
RoB 2 and NOS measure study-level risk of bias, meaning how likely it is that the methods introduced systematic error. ROBINS-I also measures study-level bias but targets non-randomized studies of interventions with far more granularity than NOS.
JBI checklists measure methodological quality, especially for qualitative and mixed-methods research where bias operates differently. QUADAS-2 measures risk of bias and applicability concerns in diagnostic test accuracy studies. AMSTAR 2 evaluates the methodological quality of other systematic reviews, making it essential for umbrella reviews.
GRADE measures evidence certainty at the outcome level across a body of studies. Risk of bias findings from the other tools feed into GRADE as one of five domains. GRADE is never a replacement for study-level assessment. It builds on top of it.
Mixing these up creates serious problems. NOS is an input to GRADE, not a substitute. QUADAS-2 cannot replace RoB 2 for intervention studies. Reviewers will flag misapplication immediately.
RoB 2: The Standard for Randomized Controlled Trials
Developed by Sterne et al. (2019) and endorsed by Cochrane, RoB 2 evaluates five domains: randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. Each domain receives a judgment of Low risk, Some concerns, or High risk.
The current version uses structured signaling questions within each domain to guide assessors, improving inter-rater reliability substantially over the original Cochrane tool.
When to use RoB 2: Any systematic review that includes randomized controlled trials. Mandatory for Cochrane reviews. Use a separate tool for any non-randomized studies.
Newcastle-Ottawa Scale: Cohort and Case-Control Studies
Developed by Wells et al., NOS uses a star-rating system across three categories: Selection, Comparability, and Outcome/Exposure. Maximum score is 9 stars. Common thresholds: 7+ high quality, 5-6 moderate, below 5 low.
Key limitation: the star system creates a false sense of precision. NOS is best used for domain-level comparison rather than total-score ranking.
NOS versus ROBINS-I: NOS is faster (10 to 15 minutes) and simpler. ROBINS-I is more thorough (30 to 60 minutes) and preferred when you need granularity comparable to RoB 2.
JBI Checklists: Qualitative, Cross-Sectional, and Mixed-Methods
The Joanna Briggs Institute provides 13 critical appraisal checklists covering designs that NOS and RoB 2 do not handle: qualitative research, cross-sectional studies, case reports, case series, and prevalence studies.
Each item is rated Yes, No, Unclear, or Not applicable with no aggregate score. JBI avoids numerical scoring because a single number cannot capture the multidimensional nature of methodological quality in qualitative research.
When to use JBI: Studies outside RoB 2 and NOS scope. Especially valuable for mixed-methods systematic reviews and scoping reviews.
Our JBI Critical Appraisal tool covers all major checklist types and exports results for your appendices.
QUADAS-2: Diagnostic Test Accuracy Studies
Developed by Whiting et al. (2011), QUADAS-2 is the standard tool for diagnostic test accuracy studies. It has four domains: patient selection, index test, reference standard, and flow and timing. Each domain is assessed for risk of bias, and the first three are also assessed for applicability concerns.
Diagnostic accuracy studies have unique bias sources: verification bias, spectrum bias, and incorporation bias all require signaling questions that RoB 2 does not contain.
When to use QUADAS-2: Reviews comparing imaging modalities, laboratory tests, screening instruments, and AI-based diagnostic algorithms.
AMSTAR 2: Assessing the Quality of Systematic Reviews
Developed by Shea et al. (2017), AMSTAR 2 evaluates the methodological quality of systematic reviews themselves. It contains 16 items, of which 7 are critical domains. Failing on even one critical domain results in a Critically Low confidence rating. The four categories are High, Moderate, Low, and Critically Low.
When to use AMSTAR 2: Umbrella reviews, overviews of reviews, guideline development, and research proposals demonstrating weaknesses in prior reviews.
ROBINS-I: Non-Randomized Studies of Interventions
Developed by the Cochrane Bias Methods Group, ROBINS-I evaluates seven domains: confounding, participant selection, intervention classification, deviations from interventions, missing data, outcome measurement, and reported result selection.
Each domain is judged as Low, Moderate, Serious, Critical, or No information. The overall judgment takes the worst domain rating. Cochrane reviews now require ROBINS-I for non-randomized studies rather than NOS.
Use our free ROBINS-I Tool for structured assessments with all seven domains.
Need expert quality assessment for your review?
Our methodologists conduct dual-reviewer risk of bias assessments, GRADE certainty ratings, and publication-ready summary tables.
QUADAS-2 and AMSTAR 2 in Depth: The Domains Reviewers Check
These two tools are the ones researchers most often apply superficially, so it is worth seeing exactly what each one judges, because reviewers will check.
QUADAS-2 (diagnostic test accuracy studies) assesses four domains, and the structure has a feature people miss: the first three domains are rated twice, once for risk of bias and once for applicability.
Patient selection (risk of bias and applicability): was a consecutive or random sample enrolled, were case-control designs avoided, and were inappropriate exclusions avoided? The classic flaw here is spectrum bias, enrolling clearly diseased and clearly healthy patients while excluding the ambiguous cases that make diagnosis hard, which inflates apparent accuracy.
Index test (risk of bias and applicability): was the test interpreted without knowledge of the reference standard, and was the positivity threshold pre-specified? A threshold chosen after seeing the data is a common, serious bias.
Reference standard (risk of bias and applicability): is the reference standard likely to classify the target condition correctly, and was it interpreted blind to the index test?
Flow and timing (risk of bias only): did all patients receive a reference standard, the same reference standard, and were they all included in the analysis? Watch for partial verification bias (only test-positive patients get the reference standard) and differential verification bias (positives and negatives get different reference standards).
Each domain is judged low, high, or unclear risk. A diagnostic review that reports a single overall "quality score" rather than these domain-level judgements is using QUADAS-2 incorrectly.
AMSTAR 2 (appraising existing systematic reviews) has 16 items, but the overall rating does not come from summing them. It hinges on seven critical domains, where a flaw undermines confidence in the whole review:
Item 2: protocol registered before the review began
Item 4: adequacy of the literature search
Item 7: justification for excluding individual studies (with a list of exclusions)
Item 9: risk-of-bias assessment of included studies
Item 11: appropriateness of the meta-analytical methods
Item 13: accounting for risk of bias when interpreting results
Item 15: assessment of publication bias
The overall confidence rating follows a rule, not a tally: High = no or one non-critical weakness; Moderate = more than one non-critical weakness; Low = one critical flaw (with or without non-critical weaknesses); Critically Low = more than one critical flaw. A review can satisfy 14 of 16 items and still be rated Critically Low if the two it misses are both critical domains, which is exactly the judgement reviewers expect you to make rather than reporting "12 of 16 met."
Head-to-Head Comparison of All Seven Tools
Tool
Target Study Design
Domains
Scoring
Time per Study
Developer
RoB 2
Randomized controlled trials
5
Low / Some concerns / High
15-30 min
Sterne et al. (2019)
NOS
Cohort and case-control
3 categories
Stars (max 9)
10-15 min
Wells et al.
JBI
Qualitative, cross-sectional, case series
8-13 items
Yes / No / Unclear / NA
10-20 min
Joanna Briggs Institute
QUADAS-2
Diagnostic test accuracy
4
Low / High / Unclear
20-30 min
Whiting et al. (2011)
AMSTAR 2
Systematic reviews
16 (7 critical)
High / Moderate / Low / Critically Low
20-40 min
Shea et al. (2017)
ROBINS-I
Non-randomized interventions
7
Low / Moderate / Serious / Critical
30-60 min
Cochrane
GRADE
All designs (outcome-level)
5 down + 2 up
High / Moderate / Low / Very Low
15-30 min
Schunemann et al.
Match the rows to the study designs in your included studies and you have your answer. Most reviews will need two or three tools from this list, not just one.
Figure 1. Quality assessment tool to study-design fit matrix.
When to Use Multiple Tools in One Review
Most systematic reviews include studies with more than one design. Here are the most common multi-tool scenarios:
Mixed-design intervention reviews require RoB 2 for the trials and either NOS or ROBINS-I for the observational studies, then GRADE for each pooled outcome. Diagnostic accuracy reviews with intervention comparisons may need QUADAS-2 alongside RoB 2. Umbrella reviews always require AMSTAR 2 for included reviews. Mixed-methods reviews need JBI for qualitative studies alongside RoB 2 or NOS for quantitative studies.
Practical tip: State your planned tools in the protocol. PROSPERO registration requires you to specify which quality assessment instruments you will use for which study designs. Changing tools after data extraction raises transparency concerns.
How to Report Quality Assessment Results (PRISMA Requirements)
PRISMA 2020 (Item 13b) requires you to present risk of bias assessments for each included study and describe how they were incorporated into the synthesis.
RoB 2 and ROBINS-I: Use traffic-light plots and summary bar charts. Our Risk of Bias Tool generates publication-ready charts. NOS: Present star ratings by domain, not just totals. JBI: Checklist completion table with narrative of common concerns. QUADAS-2: Paired bar charts per domain plus applicability concerns. AMSTAR 2: Confidence rating plus which critical domains were not met. GRADE:Summary of Findings tables from our GRADE Evidence Tool. See our GRADE framework guide for details.
You must also describe how quality assessment influenced conclusions through sensitivity analyses, subgroup analysis by bias level, or GRADE downgrading. Stating that quality was assessed without connecting it to conclusions is a common reason for revision requests.
Combining Quality Scores with GRADE
Study-level tools and GRADE are complementary, not redundant. First assess each study with the appropriate tool. Then for each outcome, determine whether the evidence has a serious risk of bias problem. If most studies in a pooled estimate have high risk of bias, downgrade GRADE certainty by one or two levels. Finally, evaluate the remaining GRADE domains (inconsistency, indirectness, imprecision, publication bias).
Schunemann et al. and the GRADE Working Group provide this rule: if more than 50% of the weight in a meta-analysis comes from high risk of bias studies, downgrade by one level. If high-risk studies also show different results from low-risk studies, downgrade by two levels.
Figure 2. Study-level quality tools feed evidence into GRADE for outcome-level certainty.
GRADE ratings without underlying study-level assessments will be flagged by reviewers. Presenting RoB 2 tables without GRADE leaves readers without a clear statement about confidence in each finding.
Common Reviewer Criticisms of Quality Assessment
"Wrong tool for this study design." Using NOS for cross-sectional studies or AMSTAR 2 for primary studies leads to major revision or rejection.
"No second reviewer." Quality assessment requires at least two independent reviewers with documented disagreement resolution.
"Quality not incorporated into synthesis." Conducting assessments but ignoring them violates PRISMA. Reviewers expect sensitivity analyses or GRADE integration.
"Arbitrary quality thresholds." Categorizing NOS scores without validated thresholds is a common target. Justify cutoffs with precedent or sensitivity analysis.
"No risk of bias visualization." For RoB 2 and ROBINS-I, traffic-light plots are expected.
Worked Example: Mixed-Design Review Tool Selection
Consider a systematic review on AI-assisted screening for diabetic retinopathy with 8 randomized controlled trials, 12 cohort studies, 6 diagnostic accuracy studies, and 3 qualitative studies:
Use RoB 2 for the trials (focus on outcome measurement and blinding). Use NOS for the cohort studies (focus on comparability, since facility-level confounding is a major threat). Use QUADAS-2 for the diagnostic accuracy studies (focus on patient selection, since many AI studies use convenience samples). Use JBI Qualitative Checklist for the qualitative studies.
After individual assessments: Apply GRADE to each quantitative outcome, and the JBI ConQual approach for qualitative findings. This example illustrates why most reviews need multiple tools selected at the protocol stage.
Training Your Second Reviewer on Quality Assessment
Independent dual assessment is a requirement. Start with calibration exercises where both reviewers independently evaluate 3 to 5 pilot studies using each required tool, then compare results item by item.
Create a decision rules document for subjective judgments. Example: "Award the NOS comparability star for age if the study controlled for age in any multivariable model." These rules reduce ambiguity and make assessments reproducible.
Use structured forms rather than free text. All Research Gold tools produce structured output with predefined response options. Calculate Cohen's kappa before resolving disagreements (above 0.80 is excellent, 0.60 to 0.80 is substantial, below 0.60 requires calibration). Document your disagreement resolution process for the methods section.
Key Takeaways
RoB 2 is for randomized controlled trials only, developed by Sterne et al. (2019) for Cochrane.
NOS is for cohort and case-control studies. Report by domain rather than aggregate score.
JBI checklists cover qualitative research, cross-sectional studies, and case series.
QUADAS-2 is the only appropriate tool for diagnostic test accuracy, developed by Whiting et al. (2011).
AMSTAR 2 assesses systematic reviews themselves, developed by Shea et al. (2017). Essential for umbrella reviews.
ROBINS-I provides RoB 2-level rigor for non-randomized intervention studies.
GRADE operates at the outcome level and synthesizes study-level findings into evidence certainty ratings.
Most reviews require two or three tools. Match each to the study designs in your included studies.
Always use two independent reviewers and report inter-rater agreement.
Present results using standard visualizations and connect them to your synthesis.
Your choice of quality assessment tools depends on the review framework. Understand Cochrane versus non-Cochrane expectations to select appropriately. If your review includes mixed study designs and you need expert guidance on the right combination of tools, talk to our team from our systematic review methodologists.
Our Research Gold systematic review covers quality assessment, risk of bias, and GRADE analysis from protocol to publication.
Frequently Asked Questions
5
No. RoB 2 is designed specifically for RCTs. Using RoB 2 on a cohort study is a methodological error that reviewers will flag. Use NOS, ROBINS-I, or JBI for observational studies.
GRADE can be applied to narrative reviews. The Cochrane Handbook recommends completing GRADE and acknowledging that imprecision and inconsistency are harder to assess without quantitative synthesis.
ROBINS-I is more structured with seven domains and a four-level scale. NOS uses a star-rating system across three broad categories. ROBINS-I is preferred for Cochrane reviews; NOS is widely used in non-Cochrane journals.
There is no validated formula. Report the distribution of NOS scores and make a narrative judgment. Describe your reasoning in the GRADE footnotes.
Yes to both. A table provides transparency for reviewers. A summary bar chart provides a visual overview. GRADE findings should appear in a Summary of Findings table. Need help with your systematic review or meta-analysis? [Get a free quote](/get-a-quote) from our team of PhD researchers.
Share
Found this useful? Share it with your colleagues.
Need expert quality assessment for your review?
Our methodologists conduct dual-reviewer risk of bias assessments, GRADE certainty ratings, and publication-ready summary tables.
Quality Assessment Takes Expertise. Our Team Does It Daily.
Rigorous risk of bias assessment, GRADE evaluations, and summary tables that satisfy peer reviewers. We handle the methodology so your review stands up to scrutiny.
Our promise: Free rework on synthesis or GRADE assessment if reviewers push back.
4.9 / 5Quote within a few hoursPRISMA 2020 + GRADE certaintyPhD methodologistConfidential by default
Dr. Sarah Mitchell holds a PhD in Biostatistics from Johns Hopkins Bloomberg School of Public Health and has over 15 years of experience in systematic review methodology and meta-analysis. She has authored or co-authored 40+ peer-reviewed publications in journals including the Journal of Clinical Epidemiology, BMC Medical Research Methodology, and Research Synthesis Methods. A former Cochrane Review Group statistician and current editorial board member of Systematic Reviews, Dr. Mitchell has supervised 200+ evidence synthesis projects across clinical medicine, public health, and social sciences.
Need professional help with your systematic review or meta-analysis? Get a free quote from our team of PhD researchers.
Quality Assessment Takes Expertise. Our Team Does It Daily.
Rigorous risk of bias assessment, GRADE evaluations, and summary tables that satisfy peer reviewers. We handle the methodology so your review stands up to scrutiny.
Meta-analysis in psychology pools the effect sizes from many studies into one reliable result. Learn the definition, real examples, and how researchers run one.
Human-written, AI-assisted, AI-screened: the labels have stopped being descriptive. Here is the single threshold journals actually use, what you must disclose, and where Research Gold draws the line.