Qualitative Analysis/ Thematic Coding
Thematic Coding
How to develop thematic codes for qualitative research
Coding is the engine of qualitative analysis. This guide covers what codes are and why they matter, how your analytical approach shapes the codes you create, how to build a rigorous codebook, and how to assign and review codes in AnalyZ Solutions.
16 min read
Thematic Coding
Intermediate
01Why coding is the foundation of qualitative analysis
Qualitative data arrives as language: thousands of words spoken by participants in interviews, written in field notes, or captured in open-ended survey responses. Before any pattern can be identified, any comparison made, or any finding reported, that language must be organised. Coding is how that organisation happens.
A code is a short label applied to a passage of text that captures what is happening in that passage. Codes transform unstructured language into a structured dataset of labelled excerpts that can be sorted, counted, compared, and queried. Without coding, a set of transcripts is just text. With a well-developed codebook applied consistently, it becomes evidence.
Consider a study of barriers to healthcare access with twelve participants. One participant says: "I knew the clinic existed, but I had no idea what documents I needed to bring or whether I would be turned away." Another says: "Nobody explained the process to me. I showed up and was sent home." Both passages describe different surface experiences, but a code like "Lack of procedural information" captures what they share. Applied consistently across all twelve transcripts, that code produces a retrievable body of evidence on a specific barrier.
Coding also creates an audit trail. A well-documented codebook with definitions, analytical memos, and examples produces a transparent record of the analytical decisions that led to the findings. That transparency is one of the primary bases for claiming rigour in qualitative research.
02How codes become findings
Codes are not findings in themselves. A finding emerges when coded excerpts are read, compared, and interpreted in relation to the research question.
Retrieval: all segments assigned to a code are read together. This lets the researcher examine the full body of evidence on a theme across all participants and see variation within it.
Within-code variation: even within a single code, passages vary. Some participants express a theme with urgency; others mention it in passing. Reading all segments together reveals the internal texture of a theme and prevents the researcher from treating a single vivid quote as representative.
Cross-code relationships: findings often emerge from the relationship between codes. A researcher might notice that "Lack of procedural information" frequently appears alongside "Previous negative experience" in the same segments, suggesting a compound barrier that is more than the sum of its parts.
Interpretation: the researcher brings theoretical knowledge and contextual understanding to the coded excerpts and argues for what the pattern means. Codes organise the evidence; the researcher makes the argument.
03What makes a good code
Descriptive rather than evaluative. A code should describe what is happening in the passage, not judge it. "Perceived barriers to uptake" is descriptive. "Failure to engage" imports the researcher's assessment. Descriptive codes stay close to what participants actually said.
Mutually distinct. Each code should capture something different from every other. If two codes frequently appear on the same passage and you cannot articulate the difference between them, they likely represent a single idea and should be merged.
Consistently applicable. A code should be applicable across all transcripts and by any researcher working with the data. Write a definition for every code before you begin coding.
Grounded in the data. Where possible, use participant language in code names. A code called "Sending the child away" (drawn from participants' own words) is more grounded than "Temporary separation strategy."
Keep code names short but specific. "Awareness" is too broad when there are multiple types of awareness in the data. "Awareness of programme eligibility criteria" is precise. In AnalyZ Solutions you can add a longer description and analytical memo to each code, so the name can be concise while the definition is thorough.
04How your analytical approach shapes your codes
The codes you create and how you create them should reflect the analytical approach your research design requires. A thematic analysis coded inductively will produce a different codebook than a framework analysis coded deductively from the same data, even when both are conducted rigorously.
Inductive coding
In inductive coding, codes emerge from the data. The researcher reads through transcripts without a predetermined framework and creates codes as new ideas and patterns become visible. The process is iterative: early codes are often refined or renamed as more of the data is read.
Inductive coding is suited to exploratory research where no adequate theoretical framework exists, or where the researcher wants to ensure participants' own categories drive the analysis. It is common in grounded theory and in the early stages of thematic analysis.
The risk is drift: without a fixed framework, the researcher may code inconsistently across transcripts. Frequent review of the codebook and early passages is essential.
Deductive coding
In deductive coding, the codebook is developed before coding begins, based on a theoretical framework, a programme theory of change, or a set of research questions. The researcher reads through the transcripts looking for evidence of each pre-defined code.
Deductive coding is well suited to evaluation research, where the questions are fixed in advance. It is also more reliable across multiple coders because definitions are established before any coding begins. The risk is closure: passages that do not fit the predetermined codes may be overlooked.
Hybrid coding (most common in practice)
Most applied qualitative research uses a hybrid approach: a starting set of deductive codes from the research questions, supplemented by inductive codes that emerge during analysis. The deductive codes provide structure; the inductive codes capture what the framework did not anticipate.
An evaluation of a maternal health programme might begin with deductive codes for each component of the programme theory: "Awareness of services", "Access to services", "Quality of care". During coding, the researcher repeatedly encounters passages about husbands' involvement in healthcare decisions, which was not in the original framework. An inductive code "Spousal influence on health decisions" is added. At the end of analysis, that inductive code is one of the most significant findings.
Match your coding approach to your methodology. A phenomenological study should not begin with a deductive codebook: the methodology requires the researcher to bracket prior assumptions and let the data speak. A programme evaluation with a fixed theory of change needs to demonstrate coverage of the intended outcomes. Before creating codes, be clear about what your research design requires.
05Building a codebook
A codebook is a structured document that defines every code used in the analysis. It is the single most important safeguard for analytical consistency, and the primary evidence that the analysis was conducted systematically rather than impressionistically.
Each entry should include: a code name; a definition of one to three sentences; inclusion criteria (what qualifies as a passage for this code); exclusion criteria (what should not be coded here, especially passages easily confused with this code); one or two example passages from the data; and an analytical memo recording the code's theoretical significance and how it developed.
As a concrete example, a code called "Cost as a barrier" in a healthcare access study might be defined as: "Any passage in which a participant identifies financial cost as a reason for not seeking, accessing, or continuing with healthcare. Includes direct costs (fees, medicines, transport) and indirect costs (lost income, childcare). Does not include passages where cost is mentioned in passing without being identified as a barrier." An example passage: "I would have gone to the doctor but the consultation fee alone would have taken a week's wages."
Write codebook entries before you code, not after. Researchers who create codes during coding and write definitions afterwards often find the definition does not match how they actually applied the code. Write the definition first, code to it, and update it only through a deliberate revision process with a dated note in the analytical memo.
06How to build and manage your codebook in AnalyZ Solutions
Open Thematic Coding from the qualitative home page and select Manage Codes. This is where you create and maintain your codebook throughout the analysis.
For each code, enter a name, a description (the definition), an analytical memo, keywords for auto-suggest, and a colour. Colours distinguish codes in the transcript view, coding stripes, and analysis charts. To edit an existing code, click the edit button on any code card. Changes to the definition apply immediately across all segments, so update the analytical memo with a dated note whenever you revise a definition mid-analysis.
Did you forgot to add a code? Don't worry! You can always add new codes later. Also, while assigning code you can add new codes as and when they come up.
07How to assign codes to transcript segments in AnalyZ Solutions
Select Assign Codes from the Thematic Coding cards. The screen shows a transcript on the left and the codebook panel on the right.
- Read the transcript before coding it
Read through each new transcript once without coding. Take note of what is relevant, surprising, or unclear. This reading pass reduces over-coding and gives you a sense of the transcript's overall shape before labelling individual passages.
- Select a passage
Click and drag to highlight the text you want to code. Choose the unit that captures one complete thought or experience. A passage addressing two separate themes should usually be split into two shorter segments.
- Click a code in the codebook panel
With text selected, click the code name. The segment is saved immediately. The passage receives a coloured underline in the transcript view. Coding stripes in the right margin show coding density: bright sections have been coded; dark sections have not.
- Assign multiple codes if needed
Keep the selection active and click additional codes. For example, a passage reading "I heard about the programme from a neighbour, and I worried it would cost too much" might receive both "Source of programme awareness" and "Cost as a barrier."
- Remove a misapplied code
Already-assigned codes show a small badge at the top-right corner. Click the badge to remove that code from the segment without affecting other codes on the same passage.
- Use keyword auto-suggest
Click Auto-assign codes in the toolbar. AnalyZ Solutions scans the transcript for the keywords defined in each code and generates suggestions for review. Always review auto-suggestions before accepting: keyword matching finds passages containing your terms but cannot assess relevance in context.
Watch a quick demonstration of developing and assigning codes in AnalyZ Solutions
Use the Uncoded passages toggle after finishing a transcript to highlight sections that have not been assigned any code. Long stretches of uncoded text in a transcript covering your research topic may indicate missed passages worth reviewing.
08Reviewing coded segments
Select Review Codes to audit your coding at any point. Two filter modes are available.
Simple filter: use the transcript and code dropdowns to restrict the segment list. Filtering by a single code and reading all its segments together is the most direct way to assess whether the code has been applied consistently.
Query builder: build Boolean retrieval conditions across codes. AND returns segments coded with both Code A and Code B simultaneously. OR returns segments carrying either code. NOT returns segments with Code A but not Code B. For example: a researcher wants to find all passages about nutrition knowledge that do not also mention food insecurity, to examine knowledge themes separately from access barriers. She builds the query "Nutrition knowledge" NOT "Food insecurity."
09A note on inter-rater reliability
When more than one researcher is coding the same transcripts, consistency between coders is a quality concern that needs to be addressed systematically. Two researchers applying the same codebook will not always code the same passages or apply the same code to passages they both identify as relevant.
The standard approach is to have both researchers independently code the same subset of transcripts, calculate Cohen's Kappa to measure agreement, and hold a calibration session to discuss and resolve disagreements before coding the remaining data. A Kappa of 0.60 or above is the widely accepted minimum for qualitative research.
AnalyZ Solutions supports this through the Compare Coding module. Each coder exports their coding as a JSON file from Review Codes; those files are loaded into Compare Coding, which calculates Kappa overall and per code. See How to calculate inter-rater reliability for the full process.
10Common mistakes to avoid
- Coding too broadly. A code called "Challenges" containing passages about financial barriers, distance, cultural stigma, and bureaucratic complexity is analytically useless. Split it into specific codes.
- Coding too narrowly. If "worried about the cost", "concerned about fees", and "cannot afford it" are each receiving separate codes, they should be merged under "Cost as a barrier."
- Changing code definitions without recoding earlier transcripts. When a definition changes, go back to already-coded transcripts and verify earlier passages still fit the revised definition.
- Skipping the analytical memo. The memo captures your thinking about what the code means and how it developed. Researchers who skip memos often struggle during write-up to articulate what a theme actually means.
- Treating frequency as importance. A code appearing in every transcript is not necessarily more significant than one appearing in two. A theme raised briefly by many participants may matter less than one explored at length by a few.
Frequently Asked Questions
How many codes should I have in my codebook?
There is no right number. A focused study might produce 10 to 15 codes; a broader study might produce 30 or more. A codebook with more than 50 codes usually indicates some codes are too narrow and could be grouped under broader themes. Fewer than 8 codes for a multi-topic study may not capture enough of the data's complexity.
Should I code the interviewer's speech as well as the participant's?
Usually not. Coding is typically applied to participant responses. The exception is research studying interview dynamics, or where the interviewer's contributions are themselves analytically relevant. In most applied research, code participant speech only and note any relevant interviewer interventions in the analytical memo.
Can the same passage receive more than one code?
Yes. Multiple coding is common and analytically valuable. A passage addressing two themes should receive both codes. This also enables co-occurrence analysis later, revealing which themes are discussed together.
When should I add a new code mid-analysis?
Add a new code when you encounter a passage that is clearly relevant to your research question but does not fit any existing code, and when you are confident it reflects a genuinely distinct theme. Write a clear definition immediately and return to already-coded transcripts to check whether earlier passages should have received the new code.
My colleague and I are reaching different codes. Is that a problem?
Different code names for the same concept can be resolved through calibration. Different codes representing genuinely different interpretations of the same passages are a sign the codebook definitions need sharpening. A calibration session, in which both coders read the same passages together and discuss their reasoning, usually resolves disagreement. If disagreement persists, consider whether both interpretations are valid and whether both codes should be retained.
How do I decide whether two codes should be merged or kept separate?
Pull all segments for both codes and read them side by side. If you cannot clearly articulate what distinguishes a passage in Code A from a passage in Code B, the codes are probably capturing the same idea and should be merged. If the passages feel different but you struggle to put the difference into words, that difficulty is usually a sign that the code definitions need sharpening rather than that the codes need merging. A useful test: could a colleague, reading only the code definitions, correctly assign new passages to one and not the other? If not, revise the definitions before merging.
Is there a difference between a code and a category?
In most qualitative software and methodology texts, the two terms are used interchangeably, but some traditions distinguish them. In grounded theory, a category is a higher-level concept that groups related codes. In framework analysis, categories are the predetermined analytical domains within which codes are organised. In practice, what matters is that you are consistent within your own study. If you use both terms, define each clearly and explain the relationship between them in your methods section.
How detailed should my code definitions be?
Detailed enough that a second researcher, reading only the definition and not having seen the data, could apply the code consistently. A good definition includes what the code covers, what it does not cover, and at least one concrete example of a qualifying passage. A definition that says only "passages about access" is too vague. A definition that says "passages in which a participant identifies a specific obstacle to physically reaching or entering a health facility, including distance, transport cost, road conditions, and opening hours, but not passages about the cost of services once inside the facility" gives a coder what they need to work with.
Should I develop my codebook before or after reading the transcripts?
It depends on your analytical approach. For deductive coding, the codebook is developed before reading the transcripts, based on the research questions or theoretical framework. For inductive coding, you read the transcripts first and let codes emerge from the data. In hybrid approaches, which are most common in applied research, you develop an initial set of codes from the research questions before reading, then add inductive codes as you encounter material that does not fit the initial framework. Reading the transcripts before finalising any codebook, even a deductive one, is good practice: it allows you to refine definitions with concrete examples in mind.
What is the difference between descriptive and interpretive coding?
Descriptive coding labels what a passage is about at face value: the topic, the event, the behaviour described. "Participant describes transport difficulty" is descriptive. Interpretive coding labels what a passage means or implies at a deeper level: the assumption, the emotion, the social mechanism. "Structural exclusion through mobility constraints" applied to the same passage is interpretive. Most applied qualitative research works primarily at the descriptive level, occasionally moving to the interpretive level for analytically significant passages. Academic research, particularly in sociology and critical studies, tends to work at the interpretive level throughout.
How should I handle a passage that contradicts my emerging themes?
Code it and take it seriously. Disconfirming passages are analytically productive. A passage that does not fit the dominant pattern may reveal that the pattern is less universal than it appeared, that a theme has important exceptions, or that the code definition needs revision. Qualitative findings are strengthened, not weakened, by explicitly acknowledging contradictions and explaining them. A finding that says "most participants described X, with the exception of two participants who reported Y because of Z" is more credible than one that ignores the exceptions.
Can I code data from different sources, such as interviews and documents, using the same codebook?
Yes, and doing so can strengthen a study through triangulation. However, some codes may apply to one source type but not another. Programme documents may produce evidence for codes about stated intentions or official rationale that interviews would not, while interviews produce evidence for lived experience that documents cannot capture. Note which source each segment comes from, and consider whether the nature of the evidence is comparable before combining counts from different source types in a frequency analysis.
How do I know when my coding is finished?
Coding is complete when every transcript has been coded, every segment has been reviewed for consistency, and no new passages are prompting the creation of new codes. A practical check: re-read the first transcript you coded after finishing the last one. If the first transcript now looks under-coded relative to the last, there has been coding drift and earlier transcripts need revisiting. If the coding looks consistent, the analysis is ready to move to the theme development stage.
Start developing your codebook and assign them to transcripts
Open Thematic Coding in AnalyZ Solutions and begin building your codes. Free to sign up, free to analyze.
Get started
Related Guides