Why "themes emerged" and a second coder can both be wrong
In reflexive thematic analysis two habits imported from quantitative research are not just unnecessary but conceptually mismatched. The first is the passive phrase that themes emerged from the data: Braun and Clarke argue this denies the interpretive work you did and misrepresents themes as objective facts you merely transcribed. The second is using inter-rater reliability, a second coder and a Cohen's kappa, as a validity check. If a theme is a meaning you constructed through engagement with the data and your theoretical lens, then two coders agreeing does not make it more true; it only shows two people were trained to see the same thing. A second analyst can still enrich a reflexive analysis by deepening interpretation, but as a collaborator, not as a reliability instrument. By contrast, in a coding-reliability or codebook design, intercoder agreement is exactly the right credibility anchor. The point is to match the quality practice to the approach, not to apply one checklist everywhere.
What a theme actually is, and the domain-summary trap
The earlier contrast between a topic and a theme has a more precise diagnostic. A fully realised theme is organised around a central organising concept, a single idea that gives the theme its coherence and that you could state in a sentence. The common failure is the domain summary: a "theme" that is really just a bucket for everything participants said about a question (for example a theme called "barriers to care" that simply collects every barrier mentioned). Domain summaries describe the data; themes interpret it. A quick test is whether your theme names a shared meaning or merely a question from your interview guide; if it mirrors the guide, you have summarised a domain rather than built a theme.
Sample size, saturation, and the reflexive position
Reviewers often ask how you justified your sample size or whether you reached saturation, and for reflexive thematic analysis the honest answer is more nuanced than a number. Braun and Clarke have argued that data saturation is conceptually incoherent for an approach where meaning is generated rather than discovered, because there is always more interpretation to be made. The better framing is information power (Malterud and colleagues): the more relevant information your sample holds for the specific study (a narrow aim, a dense sample, strong dialogue, a clear theory), the fewer participants you need. State your sample size in those terms, and report your theoretical positioning explicitly, whether you analysed from a realist or essentialist stance that takes accounts at face value, or a constructionist one that treats them as produced in a context, because that stance governs what your semantic or latent codes are entitled to claim. Documenting reflexivity across its personal, interpersonal, methodological, and contextual dimensions is what turns a defensible position into a transparent one.
Thematic analysis is learnable, but a defensible analysis of forty in-depth interviews is weeks of disciplined work, and supervisors and reviewers scrutinize qualitative rigor closely. If your timeline is tight or the analysis has to withstand a methods-savvy examiner, expert support on coding structure, intercoder reliability, and the analytic write-up is often the difference between a thin description and a publishable contribution.