The STARD checklist is the reporting guideline for diagnostic accuracy studies, and the current version is STARD 2015, which carries 30 items. Its name stands for Standards for Reporting Diagnostic accuracy studies. It was first published in 2003 and revised in 2015 by Bossuyt and colleagues after evidence that the methodological details most likely to bias an accuracy estimate were still routinely missing from published reports.
The guideline's centre of gravity is the reference standard, meaning whatever the new test was judged against. That comparison is the whole study: an accuracy estimate is only as meaningful as the standard used to define who really had the condition, and most of the ways an accuracy study goes wrong are ways that comparison becomes unfair.
Accuracy is not a property of a test
This is the conceptual point behind most of the items, and it is worth stating before the list. Sensitivity and specificity are frequently quoted as if they belonged to a test, like a serial number. They do not. They are properties of a test applied to a particular population at a particular threshold, and they move when either changes.
A test evaluated in a specialist clinic, where patients have already been filtered by a referring clinician, will look more accurate than the same test used in primary care on undifferentiated patients. This is spectrum bias, and it is why an accuracy study has to describe its population precisely rather than by diagnosis alone. Predictive values additionally depend on prevalence, so a positive predictive value from a high-prevalence setting cannot be transported to a screening context at all.
The practical consequence is that reporting the participant flow, the eligibility criteria, the setting and the threshold is not bureaucratic detail. Without them the numbers cannot be applied anywhere.
The 30 items in outline
Described in our own words, grouped as the guideline arranges them. The authoritative wording sits with the STARD group and is indexed on EQUATOR's guideline index.
Title and abstract. That the article reports a study of diagnostic accuracy, using at least one measure such as sensitivity or specificity; and a structured abstract covering design, methods, results and conclusions.
Introduction. The scientific and clinical background including the intended use and clinical role of the index test, and the objectives or hypotheses.
Methods, design and participants. Whether data collection was prospective or retrospective relative to the index test and reference standard; the eligibility criteria; how participants were identified, whether by presenting symptoms, prior test results or from a registry; whether participants formed a consecutive, random or convenience series; where and when the study took place; and whether participants underwent the index test and reference standard before or after being assigned to groups.
Methods, test methods. The index test described in sufficient detail to permit replication, and the reference standard likewise with a rationale for its choice; the definition and rationale for the cut-offs or categories of both, distinguishing prespecified from exploratory; whether assessors of the index test and the reference standard were blind to the other's results and to clinical information.
Methods, analysis. The methods for estimating or comparing measures of accuracy; how indeterminate, missing or outlying results were handled; any analysis of variability in accuracy across subgroups, readers or centres, distinguishing prespecified from exploratory; and the intended sample size with how it was determined.
Results. A flow of participants, preferably as a diagram; the baseline demographic and clinical characteristics; the distribution of severity in those with the target condition and of alternative diagnoses in those without; the time interval and any clinical interventions between index test and reference standard; the cross-tabulation of index test results against the reference standard; the estimates of accuracy with their precision, such as confidence intervals; any adverse events from either test; and the results of any variability analyses.
Discussion. The study limitations, including sources of potential bias, statistical uncertainty and generalisability; and the implications for practice, including the intended use and clinical role of the test.
Other information. The registration number and registry, where the full protocol can be accessed, and sources of funding and other support with the role of funders.