Back to Blog
Methodology
8 min read

STARD Checklist: Reporting Diagnostic Accuracy

STARD 2015 checklist explained: all 30 items, the reference standard and blinding problems in diagnostic accuracy studies, and how STARD differs from QUADAS-2.

Research Gold Team

September 11, 2026

Calculating accuracy measures? Our diagnostic accuracy calculator derives sensitivity, specificity, predictive values and likelihood ratios with confidence intervals.

Key Takeaways

STARD 2015 is the reporting guideline for diagnostic and prognostic accuracy studies and carries 30 items

Its central concern is the reference standard: what the new test was compared against, and whether that comparison was fair

Sensitivity and specificity are not properties of a test alone; they depend on the population and the threshold

STARD reports a study. QUADAS-2 appraises one, and the two are frequently confused

A flow diagram accounting for every participant, including indeterminate results, is expected

The STARD checklist is the reporting guideline for diagnostic accuracy studies, and the current version is STARD 2015, which carries 30 items. Its name stands for Standards for Reporting Diagnostic accuracy studies. It was first published in 2003 and revised in 2015 by Bossuyt and colleagues after evidence that the methodological details most likely to bias an accuracy estimate were still routinely missing from published reports.

The guideline's centre of gravity is the reference standard, meaning whatever the new test was judged against. That comparison is the whole study: an accuracy estimate is only as meaningful as the standard used to define who really had the condition, and most of the ways an accuracy study goes wrong are ways that comparison becomes unfair.

Accuracy is not a property of a test

This is the conceptual point behind most of the items, and it is worth stating before the list. Sensitivity and specificity are frequently quoted as if they belonged to a test, like a serial number. They do not. They are properties of a test applied to a particular population at a particular threshold, and they move when either changes.

A test evaluated in a specialist clinic, where patients have already been filtered by a referring clinician, will look more accurate than the same test used in primary care on undifferentiated patients. This is spectrum bias, and it is why an accuracy study has to describe its population precisely rather than by diagnosis alone. Predictive values additionally depend on prevalence, so a positive predictive value from a high-prevalence setting cannot be transported to a screening context at all.

The practical consequence is that reporting the participant flow, the eligibility criteria, the setting and the threshold is not bureaucratic detail. Without them the numbers cannot be applied anywhere.

The 30 items in outline

Described in our own words, grouped as the guideline arranges them. The authoritative wording sits with the STARD group and is indexed on EQUATOR's guideline index.

Title and abstract. That the article reports a study of diagnostic accuracy, using at least one measure such as sensitivity or specificity; and a structured abstract covering design, methods, results and conclusions.

Introduction. The scientific and clinical background including the intended use and clinical role of the index test, and the objectives or hypotheses.

Methods, design and participants. Whether data collection was prospective or retrospective relative to the index test and reference standard; the eligibility criteria; how participants were identified, whether by presenting symptoms, prior test results or from a registry; whether participants formed a consecutive, random or convenience series; where and when the study took place; and whether participants underwent the index test and reference standard before or after being assigned to groups.

Methods, test methods. The index test described in sufficient detail to permit replication, and the reference standard likewise with a rationale for its choice; the definition and rationale for the cut-offs or categories of both, distinguishing prespecified from exploratory; whether assessors of the index test and the reference standard were blind to the other's results and to clinical information.

Methods, analysis. The methods for estimating or comparing measures of accuracy; how indeterminate, missing or outlying results were handled; any analysis of variability in accuracy across subgroups, readers or centres, distinguishing prespecified from exploratory; and the intended sample size with how it was determined.

Results. A flow of participants, preferably as a diagram; the baseline demographic and clinical characteristics; the distribution of severity in those with the target condition and of alternative diagnoses in those without; the time interval and any clinical interventions between index test and reference standard; the cross-tabulation of index test results against the reference standard; the estimates of accuracy with their precision, such as confidence intervals; any adverse events from either test; and the results of any variability analyses.

Discussion. The study limitations, including sources of potential bias, statistical uncertainty and generalisability; and the implications for practice, including the intended use and clinical role of the test.

Other information. The registration number and registry, where the full protocol can be accessed, and sources of funding and other support with the role of funders.

Need professional help with your research?

Our PhD methodologists deliver complete systematic reviews and meta-analyses, from protocol to manuscript.

The four ways a reference standard comparison goes wrong

Four recurring problems account for most of the bias in accuracy studies, and STARD's items exist to expose each one.

An imperfect reference standard. If the standard itself misclassifies some patients, the index test is penalised for being right when the standard was wrong. Where no perfect standard exists, some studies use a composite standard or expert panel adjudication, both of which are legitimate if described. What is not legitimate is presenting an imperfect standard as definitive.

Incorporation bias. If the index test result is used, even partly, in deciding the reference standard diagnosis, the two are no longer independent and accuracy is inflated. This happens easily with expert panels who have access to the full chart.

Partial or differential verification. If only some participants receive the reference standard, and whether they do depends on their index test result, the resulting estimates are biased. This is common in practice, because clinicians reasonably do not order invasive confirmation on patients who tested negative.

A long or unequal interval between tests. If the reference standard is applied weeks after the index test, the patient's condition may have changed, so a disagreement is not necessarily an error. The item asking for the time interval exists precisely to let a reader judge this.

Blinding matters for the same reasons. An assessor reading the index test who already knows the reference result, or vice versa, cannot be relied on to be independent, and the item asks who knew what at each stage rather than for a general assurance.

Applying a result to a patient? The Fagan nomogram converts a likelihood ratio into a post-test probability.

STARD reports, QUADAS-2 appraises

These two are confused more often than any other pair in this literature, so it is worth stating the difference flatly. STARD is a reporting checklist used by authors writing up their own accuracy study; it asks whether the study is adequately described. QUADAS-2 is a risk of bias tool used by reviewers appraising somebody else's accuracy study; it asks whether the study was well conducted, across four domains of patient selection, index test, reference standard, and flow and timing.

A study can be reported in full compliance with all 30 STARD items and be judged at high risk of bias on QUADAS-2, because complete reporting of a flawed design simply makes the flaw visible. Our QUADAS-2 tool handles the appraisal side, and the same reporting-versus-appraisal distinction applies across designs, with the observational-study equivalent, STROBE and CONSORT likewise being reporting guidelines rather than quality scores.

For related designs, note that TRIPOD rather than STARD applies to multivariable prediction models, and that a systematic review of accuracy studies follows the PRISMA extension for diagnostic test accuracy rather than STARD itself.

Presenting the numbers so they can be reused

Report the two-by-two table of raw counts. Everything else, including sensitivity, specificity, predictive values, likelihood ratios and diagnostic odds ratios, can be recomputed from four numbers, and a study that gives only percentages cannot be included in a meta-analysis without the authors being contacted. Our diagnostic accuracy calculator derives the full set with confidence intervals from those counts.

Give precision for every estimate. A sensitivity of 90 percent from 20 patients with the condition has a confidence interval wide enough to include values that would make the test useless, and a bare percentage hides that.

If the index test produces a continuous result, report accuracy across a range of thresholds rather than at one, and say whether your chosen cut-off was prespecified. A threshold selected to maximise accuracy in your own sample is fitted to that sample and will usually perform worse elsewhere. For syntheses across thresholds, the summary receiver operating characteristic approach applies, and our summary ROC curve generator plots it.

Finally, for readers who need to apply your result to an individual patient, likelihood ratios are the most directly useful measure because they combine sensitivity and specificity and can be used to update a pre-test probability. The Fagan nomogram performs that conversion.

Completing the checklist

Work through the 30 items against the final manuscript, recording page and section rather than ticking, and name STARD 2015 in your methods. The items most often unanswerable are the time interval between tests, how indeterminate results were handled, and whether the threshold was prespecified. All three are cheap to state and each one materially changes how a reader reads your accuracy estimate.

Pro Tip

Report the two-by-two table

Give the raw counts of true positives, false positives, false negatives and true negatives. Every other accuracy measure can be recalculated from them, and a synthesis cannot include your study without them.

Pro Tip

Say when the threshold was chosen

A cut-off selected after seeing the data to maximise accuracy will not perform as well elsewhere. State whether the threshold was prespecified or derived, and if derived, say so plainly.

Pro Tip

Account for indeterminate and missing results

Excluding uninterpretable index test results inflates accuracy. Report how many there were and how they were handled, because in practice an uninterpretable test is a clinical outcome too.

Frequently Asked Questions

6
It is the reporting checklist accompanying the STARD statement, which specifies the minimum information a diagnostic or prognostic accuracy study should report. The current version, STARD 2015, carries 30 items. Authors complete it by recording where each item is addressed in the manuscript and submit it alongside the paper.
STARD stands for Standards for Reporting Diagnostic accuracy studies. It was first published in 2003 and substantially revised in 2015 by Bossuyt and colleagues, who expanded and reorganised the items in response to evidence that key methodological details were still being omitted.
That a reader should be able to see exactly which patients were tested, what the index test and the reference standard were, how and in what order they were applied, who knew what results at each stage, how many participants produced each possible combination of results, and how accuracy estimates and their precision were calculated. The underlying principle is that accuracy is a property of a test applied to a particular population at a particular threshold, not of the test alone.
STARD 2015 is the primary guideline for a single accuracy study. For a systematic review of accuracy studies, the relevant guidance is the PRISMA extension for diagnostic test accuracy, and the appraisal instrument for the included studies is QUADAS-2. For prediction and prognostic models rather than single tests, the TRIPOD statement applies.
Studies that measure how well a test identifies a target condition by comparing its results against a reference standard applied to the same participants. Their outputs are measures such as sensitivity, specificity, predictive values and likelihood ratios. They answer how well a test classifies, not whether using it improves patient outcomes, which requires a different design.
From the two-by-two table of index test results against the reference standard. Sensitivity is true positives divided by all those with the condition. Specificity is true negatives divided by all those without it. Positive and negative predictive values divide by the numbers testing positive and negative respectively, and therefore change with prevalence. Likelihood ratios combine sensitivity and specificity and are the most useful for updating probability in an individual patient.
Share

Found this useful? Share it with your colleagues.

Need professional help with your research?

Our PhD methodologists deliver complete systematic reviews and meta-analyses, from protocol to manuscript.

Explore our Systematic Review Service, handled end-to-end by a PhD methodologist.

Professional Support

Let a PhD Expert Handle Your Research

From protocol to publication-ready manuscript. Our PhD-level methodologists handle systematic reviews, meta-analyses, scoping reviews, and more. Most projects deliver in under 2 weeks.

Our promise: Free rework on search, screening, or synthesis if reviewers push back.

4.9 / 5Quote within a few hoursPRISMA 2020 + Cochrane HandbookPhD methodologistConfidential by default
Chat on WhatsApp now
RG

Written by

Research Gold Team

PhD-Level Methodologists

Our team comprises PhD-level methodologists with peer-reviewed publication records in systematic reviews, meta-analyses, and biostatistics. Every article is fact-checked against current Cochrane Handbook, PRISMA 2020, and JBI guidelines to ensure the advice you read reflects the latest evidence synthesis standards.

Research Gold designs, analyses and reports diagnostic accuracy studies and their syntheses. See our biostatistics service or request a quote.

Let a PhD Expert Handle Your Research

From protocol to publication-ready manuscript. Our PhD-level methodologists handle systematic reviews, meta-analyses, scoping reviews, and more. Most projects deliver in under 2 weeks.

Quote within a few hours. Pay only after you approve your quote. Unlimited revisions within your agreed scope. Confidential by default.