Score the methodological quality of cohort, case-control, and cross-sectional studies using the Newcastle-Ottawa Scale. Assign stars across the selection, comparability, and outcome or exposure domains, read built-in interpretation thresholds, and export a traffic-light figure as PDF, PNG, SVG, or CSV. Working with a modified scale? Build your own instrument on the Custom Instrument tab, with your own domains, items, and star values, and share it with your review team as a JSON file.
Add studies and name them. For each study, select the option that best describes the study for every item. Options that award stars are marked with a icon. Hover or tap the icon next to any item for the official coding note. Working with a co-reviewer? Turn on Dual Reviewers to score each study twice, independently, and see item-level agreement.
Each study can earn up to 9 stars. Quality ratings: 0-3 Low, 4-6 Moderate, 7-9 High. The original scale defines no validated total-star cut-off, so these bands are a common convention only. For a published thresholding approach, convert the domain stars using the Agency for Healthcare Research and Quality (AHRQ) Good / Fair / Poor standards, and set your quality threshold a priori in your protocol.
Load sample data to see how the tool works, or clear all fields to start fresh.
Representativeness of the exposed cohort
Selection of the non-exposed cohort
Ascertainment of exposure
Demonstration that outcome of interest was not present at start of study
Comparability of cohorts on the basis of the design or analysis
Comparability - additional factor
Assessment of outcome
Was follow-up long enough for outcomes to occur?
Adequacy of follow-up of cohorts
Next step
PhD reviewer mirrors your scoring across every included study, flags disagreements with kappa, and delivers a publication-ready quality table. Fixed per-study price, no full-project commitment.
Our promise: Free rework on search, screening, or synthesis if reviewers push back.
Will your review include a meta-analysis? We run the systematic review and the pooled analysis together, in one project.
Quote my review + meta-analysisTimeline
Most projects deliver in under 2 weeks. We confirm an exact date in your quote.
If reviewers push back
If reviewers question the search, screening, or synthesis, we rework the section free.
Confidentiality
NDA available on request before any project discussion. Your data, study design, and manuscript stay private either way.
Select the appropriate tab for your study design: Cohort Studies, Case-Control Studies, or Cross-Sectional Studies. Each tab presents the relevant items for that design. Working with a modified scale? Build it on the Custom Instrument tab.
Add a row for each study in your review and enter the study identifier. Click a study row to expand its assessment form.
For each NOS item, select the option that best describes the study. Options marked with a star icon contribute to the total score.
Review the summary table showing stars per domain and overall quality ratings. Export results as CSV or copy the formatted text to your clipboard.
Want a PhD methodologist to handle the whole project?
Get a full quality appraisal for all your cohort and case-control studies. Free rework on search, screening, or synthesis if reviewers push back. Pay only after you approve your quote.
Each NOS star represents a specific methodological criterion that the study meets. Stars are awarded for secure exposure ascertainment, appropriate control selection, adequate follow-up, and other quality indicators. The maximum 9 stars span selection (4), comparability (2), and outcome/exposure (3).
While 7-9 stars is commonly considered high quality, 4-6 moderate, and 0-3 low, these thresholds are conventions rather than validated cutoffs. Use NOS scores to conduct sensitivity analyses by re-running your meta-analysis restricted to high-quality studies to test whether results are robust to study quality.
A total score of 6 can arise from very different quality profiles. Always report the domain-level breakdown (selection, comparability, outcome/exposure) so readers can identify whether specific quality concerns apply to your evidence base.
NOS inter-rater reliability is moderate. Best practice requires two independent reviewers to score each study, with discrepancies resolved through discussion or a third reviewer. This calculator has a built-in Dual Reviewers mode that captures both ratings, flags every item where the reviewers differ, and reports the agreement rate with Cohen's kappa for your methods section.
Observational studies, including cohort, case-control, and cross-sectional designs, form the backbone of evidence in many systematic reviews, particularly when randomized controlled trials are impractical or unavailable. The Newcastle-Ottawa Scale calculator implements the star-based scoring system developed by Wells et al. (2000) at the Universities of Newcastle (Australia) and Ottawa (Canada) to provide a standardized, reproducible method for appraising the methodological quality of these study designs. A NOS score tool allows reviewers to assign up to 9 stars across three categories: Selection (up to 4 stars), Comparability (up to 2 stars), and Outcome for cohort studies or Exposure for case-control studies (up to 3 stars). Each star represents a specific methodological criterion that the study satisfies, and the total score serves as a summary indicator of cohort study quality assessment that can be reported alongside effect estimates in meta-analyses. For cross-sectional studies, which fall outside the original NOS scope, a modified version adapted by Herzog et al. (2013) extends the star-based framework to a maximum of 10 stars and addresses the methodological concerns specific to cross-sectional designs, including sample size justification, response rate, and standardized outcome measurement. This calculator includes that adaptation as a dedicated tab, clearly labelled as an adaptation rather than the official scale. Because a scoping review by Kelly et al. (2024) found that two of Herzog's items, sample size and statistical test, measure reporting quality and statistical precision rather than risk of bias, the cross-sectional tab also offers a risk-of-bias focus toggle that excludes those two items and rescales the total to 8 stars, so reviewers can report either the familiar quality score or a stricter internal-validity appraisal. And because many published reviews adapt the scale to their own question, the Custom Instrument tab lets a review team define its own domains, items, and star values, score studies against that definition, and share the instrument as a JSON file so every reviewer appraises with an identical item set.
The Selection category evaluates whether the exposed and non-exposed cohorts (or cases and controls) were drawn from comparable populations and whether exposure or outcome ascertainment methods were reliable. The Comparability category, the most heavily weighted per-star domain, assesses whether the study controlled for the most important confounders, typically age, sex, and other variables central to the research question. The Outcome (or Exposure) category examines how the outcome was measured, whether follow-up was sufficiently long and complete, and whether objective assessment methods were used. The Cochrane Handbook for Systematic Reviews of Interventions (Higgins et al., 2023) acknowledges NOS as one of several validated tools for observational study quality, while PRISMA 2020 (Page et al., 2021) requires all systematic reviews to present results of quality or risk of bias assessments for every included study. It is worth noting that AMSTAR 2 (Shea et al., 2017) assesses the quality of systematic reviews themselves rather than primary studies. This is a distinct purpose that should not be confused with the NOS, which evaluates individual observational study methodology. Systematic review management platforms such as Covidence and DistillerSR now include built-in quality assessment modules that support NOS scoring alongside other appraisal tools, streamlining the workflow for review teams.
A common interpretation framework classifies studies scoring 7-9 stars as high quality, 4-6 as moderate quality, and 0-3 as low quality, although these thresholds are conventions rather than empirically validated cutoffs. Best practice involves using NOS scores in sensitivity analyses, for instance restricting a meta-analysis to high-quality studies to determine whether the pooled effect remains robust. NOS scores can also serve as a covariate in meta-regression, testing whether study quality moderates the treatment effect across studies. In dose-response meta-analyses, quality-stratified pooling by NOS score is particularly important because poorly designed studies may distort the shape of the exposure-response curve. Similarly, funnel plot asymmetry can be stratified by NOS score to disentangle publication bias from small-study effects driven by lower methodological quality. Because inter-rater reliability for NOS has been reported as moderate in validation studies (Lo et al., 2014), dual independent scoring followed by consensus resolution is essential. Reviewers should calculate and report the initial agreement rate, ideally using Cohen's kappa as an inter-rater reliability measure, in their methods section.
Choosing the right quality assessment tool depends on the study design and the level of detail required. NOS provides a relatively quick, star-based scoring system well-suited for large reviews with many observational studies. For a more granular, domain-based evaluation of non-randomized comparative studies, the ROBINS-I bias assessment for non-randomized studies uses signaling questions across seven domains with four judgment levels. Randomized trials should be appraised with the Cochrane RoB 2 risk of bias tool rather than NOS, as the two instruments address fundamentally different bias mechanisms. For reviews incorporating qualitative or prevalence data, the JBI critical appraisal checklists offer design-specific item sets from the Joanna Briggs Institute (Aromataris & Munn, 2020). Regardless of which instrument you choose, documenting the complete scoring rationale in your data extraction template builder ensures transparency and reproducibility across your review team.
The Newcastle-Ottawa Scale is a widely used tool for assessing the quality of non-randomized studies (cohort and case-control) in systematic reviews and meta-analyses. Developed by Wells et al., it assigns stars across three categories: Selection, Comparability, and Outcome (for cohort studies) or Exposure (for case-control studies). A study can earn a maximum of 9 stars, with higher scores indicating better methodological quality.
While there is no universally agreed threshold, a common interpretation uses three tiers: 7-9 stars indicates high quality, 4-6 stars indicates moderate quality, and 0-3 stars indicates low quality. Some systematic reviews use the NOS as a continuous variable in meta-regression to assess whether study quality moderates the pooled effect estimate. Always report the specific domain scores alongside the total.
Both versions share the same structure with three categories and similar principles, but they differ in specific items. The cohort version evaluates outcome assessment (independent blind assessment, record linkage, follow-up adequacy), while the case-control version evaluates exposure ascertainment (secure records, structured interviews, same method for cases and controls, non-response rate). The Comparability category is identical in both versions.
Yes. The original scale was built for cohort and case-control studies, but a modified Newcastle-Ottawa Scale adapted for cross-sectional studies (Herzog et al., 2013) is widely used in published reviews. This calculator includes it as a third tab. It is an adaptation rather than the official scale, so report it as such. If you prefer a purpose-built instrument, the JBI critical appraisal checklist for analytical cross-sectional studies is a common alternative.
The cross-sectional adaptation scores out of 10 stars instead of 9. Selection rises to a maximum of 5 stars and covers representativeness of the sample, sample size justification, comparability of non-respondents, and ascertainment of the risk factor (a validated tool earns 2 stars). Comparability keeps its maximum of 2 stars. The outcome domain (maximum 3) rewards how the outcome was assessed (independent blind assessment or record linkage earns 2 stars, self-report 1) and whether an appropriate statistical test with measures of association was reported.
Strictly, no. The Herzog 2013 cross-sectional adaptation includes a sample-size item and a statistical-test item, but Kelly et al. (2024, Journal of Clinical Epidemiology 172:111408) show that these assess reporting quality and statistical precision rather than risk of bias, and recommend keeping such items out of a risk-of-bias score. This calculator flags both items and provides a Risk-of-bias focus toggle on the cross-sectional tab that excludes them, rescaling the total to a maximum of 8 stars so the score reflects internal validity only. Leave the toggle off to reproduce the standard Herzog 10-star score for comparability with published reviews, or turn it on for a bias-focused appraisal.
On the cross-sectional tab, the Risk-of-bias focus toggle removes the two reporting-quality items (sample size and statistical test) identified by Kelly et al. (2024), scoring the study out of 8 instead of 10. Every other item, the traffic-light figure, and all exports update automatically, and the export note records that reporting-quality items were excluded. This lets you present either the familiar Herzog 2013 quality score or a stricter risk-of-bias appraisal, depending on what your protocol specifies.
For the 9-star cohort and case-control scale, a common convention reads 7 to 9 stars as high quality, 4 to 6 as moderate, and 0 to 3 as low. The 10-star cross-sectional adaptation often uses Very Good for 9 to 10 stars, Good for 7 to 8, Satisfactory for 5 to 6, and Unsatisfactory for 0 to 4. These thresholds are conventions, not validated cutoffs, so define your quality threshold a priori in your protocol and always report the domain-level stars alongside the total.
Yes. Many published reviews adapt the Newcastle-Ottawa Scale to their own research question, changing items, star values, or whole domains. The Custom Instrument tab in this calculator lets you build such an adaptation directly: define your own domains, items, coding notes, and star-earning response options, and the scoring form, summary table, traffic-light figure, and all exports work against your definition. You can start from a blank instrument, from an editable copy of the cross-sectional adaptation, or from a preset that follows a published pattern for association studies (the adapted scale of Agnew-Blais and Danese, 2016, augmented with the cross-sectional risk-of-bias domains of Kelly et al., 2024). Because a custom instrument has no validated cut-offs, define any quality bands a priori in your protocol and report the full item set alongside your scores.
Yes. Turn on Dual Reviewers mode and every study carries two independent answer sets, Reviewer A and Reviewer B, with a switcher inside each assessment form. Items where the two reviewers chose different options are flagged, and an agreement panel reports how many items were double-rated, the percent agreement, and the unweighted Cohen's kappa across all double-rated items, which you can report in your methods section. Reviewer A serves as the primary or consensus rating used in the summary table, traffic-light figure, and exports, so after resolving disagreements by discussion, record the consensus as Reviewer A.
Export it as a JSON file from the instrument builder. The file contains the complete definition: domains, items, coding notes, response options, star values, and any quality bands. A co-reviewer imports the same file on the Custom Instrument tab and scores against exactly the same item set, which is essential for dual independent rating. The instrument also saves automatically in your own browser, and the exported file doubles as the supplementary documentation of your appraisal tool that journals increasingly expect.
Key limitations include: (1) the scoring is somewhat subjective, as different reviewers may interpret criteria differently; (2) the scale assigns equal weight to all items, even though some may be more important for specific research questions; (3) the threshold cutoffs for low/moderate/high quality are not evidence-based; (4) inter-rater reliability has been reported as moderate in validation studies; (5) some adaptations mix in items that measure reporting quality or statistical precision rather than risk of bias, which Kelly et al. (2024) show should be excluded from a risk-of-bias score (this calculator lets you exclude them on the cross-sectional tab). Despite these limitations, NOS remains one of the most commonly used quality assessment tools for observational studies in systematic reviews.
Scores of 7–9 stars are generally considered high quality, 4–6 moderate quality, and 0–3 low quality. However, these thresholds are conventions, not empirically validated cutoffs. Some reviews define their own thresholds based on the research question. Always report the individual domain scores alongside the total, as a study can score 7 overall while having a critical weakness in one domain.
The NOS was developed by Wells et al. at the Universities of Newcastle and Ottawa but has limited formal validation. Inter-rater reliability has been reported as moderate (Lo et al., 2014). Despite these limitations, NOS remains the most widely used quality assessment tool for observational studies in systematic reviews. The Cochrane Handbook acknowledges NOS but recommends ROBINS-I as a more rigorous alternative for non-randomized intervention studies.
Under the cross-sectional adaptation, which is scored out of 10 stars, a study scoring 9 to 10 is often labelled Very Good, 7 to 8 Good, 5 to 6 Satisfactory, and 0 to 4 Unsatisfactory. Because the maximum differs from the 9-star cohort and case-control scale, do not compare a cross-sectional total directly against a cohort total. Report the design, the maximum used, and the domain breakdown so readers can interpret the score correctly.
Including randomized trials in your review? Use our RoB 2 tool for randomized trials to create traffic-light summary tables across 5 bias domains. For non-randomized comparative studies, the ROBINS-I assessment for non-randomized studies provides a more detailed 7-domain evaluation with signaling questions. For qualitative and mixed-methods studies, explore our JBI critical appraisal checklists covering multiple study designs.
Reviewed by
Dr. Sarah Mitchell holds a PhD in Biostatistics from Johns Hopkins Bloomberg School of Public Health and has over 15 years of experience in systematic review methodology and meta-analysis. She has authored or co-authored 40+ peer-reviewed publications in journals including the Journal of Clinical Epidemiology, BMC Medical Research Methodology, and Research Synthesis Methods. A former Cochrane Review Group statistician and current editorial board member of Systematic Reviews, Dr. Mitchell has supervised 200+ evidence synthesis projects across clinical medicine, public health, and social sciences. She reviews all Research Gold tools to ensure statistical accuracy and compliance with Cochrane Handbook and PRISMA 2020 standards.
We conduct full risk of bias assessments, GRADE evaluations, and complete systematic reviews with rigorous methodology that satisfies peer reviewers. Most projects deliver in under 2 weeks.
Our promise: Free rework on search, screening, or synthesis if reviewers push back.
Will your review include a meta-analysis? Quote my systematic review and meta-analysis
Your project is led by a named PhD methodologist with real credentials and published work.
4.9 / 5 across 1,194+ delivered projects