Qualitative Analysis/ Thematic Relationships
Analyze relationships

How to uncover relationships between themes in qualitative data

Listing the themes in your data is the beginning of qualitative analysis, not the end. The most interesting findings often come from understanding how themes connect: which ones are discussed together, which are conceptually distinct, which cluster into broader patterns, and how they sequence across the arc of an interview.

12 min read Explore Themes & Analyze Relationships Intermediate

01Why relationships between themes matter

A list of themes from a qualitative study describes what was said. A map of how those themes relate to each other begins to explain why. The difference matters enormously for the practical usefulness of qualitative findings.

Consider a study of dropout from a vocational training programme. The researcher has coded twelve transcripts and identified themes including "Financial pressure", "Family obligations", "Low confidence", "Perceived irrelevance of training content", and "Supportive peers". A list of these themes tells a commissioner that multiple factors contributed to dropout. But it does not tell them which factors tend to co-occur, which are likely to be experienced first, or which are closely associated and which are independent. That deeper picture is what relationship analysis provides.

If "Financial pressure" and "Family obligations" always appear together in the same coded segments, they may not be two separate barriers but two expressions of the same underlying constraint: household resource scarcity. If "Low confidence" consistently appears after "Perceived irrelevance of training content" in the sequence of interview, it may be that perceived irrelevance contributes to loss of confidence rather than the other way around. These relational insights transform a list of themes into an analytical argument.

02Types of relationships qualitative data can reveal

Qualitative analysis can reveal several fundamentally different types of thematic relationship. It is worth being clear about which type you are looking for before choosing an analytical approach.

Co-occurrence: themes discussed together

Co-occurrence is the simplest type of relationship. Two themes co-occur when they appear in the same coded segment or the same interview. High co-occurrence suggests that participants think about and discuss these themes as part of the same experience. It is association, not causation, but it is an important starting point for interpretation.

Example: in a study of barriers to girls' education, "Distance to school" and "Safety concerns" consistently appear in the same segments. This suggests they are not independent barriers but that distance is experienced specifically as a safety problem. A programme that reduces distance may therefore also reduce safety concerns, and vice versa.

Distinctiveness: themes that are conceptually separate

Distinctiveness is the complement of co-occurrence. Two themes are distinct when they almost never appear together. This can indicate that participants experience them as genuinely separate topics, or that they apply to different groups or contexts in the data.

Example: in the same girls' education study, "Quality of teaching" almost never appears alongside "Distance to school". Parents discussing distance are not simultaneously discussing teaching quality: they are different problems for different families. Understanding which themes are distinct is as important as understanding which co-occur.

Hierarchy: themes that are instances of broader themes

Sometimes several specific codes are all instances of a broader underlying concept. "Distance to school", "Cost of transport", and "Road safety" might all be instances of a broader theme called "Physical access". Identifying this hierarchical structure allows the analyst to report findings at the appropriate level of abstraction.

Sequence: themes that follow each other

Within a single interview, themes often appear in a characteristic order that reflects the narrative structure of participants' accounts. If "Initial enthusiasm" consistently appears before "Disillusionment" in interview transcripts, that sequence is itself a finding: it describes a trajectory, not just a set of states.

Centrality: themes that connect many others

In a network of thematic relationships, some themes sit at the centre, connected to many others. Central themes often represent the core organising experience of a phenomenon. In the vocational training dropout study, "Financial pressure" might be central not because it is mentioned most often, but because it connects to nearly every other theme in the data.

03Methodological approaches to relationship analysis

Several established approaches exist for analysing thematic relationships in qualitative data. The approach you choose should match your analytical framework and the type of relationships you are looking for.

Constant comparison

The core method of grounded theory, constant comparison involves comparing every new passage of data with all previously coded data, asking: how is this similar to or different from what I have already coded? Through this iterative process, relationships between concepts emerge naturally, because the researcher is continuously examining how new data relates to existing codes. Constant comparison is particularly suited to inductive analyses where the relationships are not anticipated in advance.

Matrix analysis

Framework analysis uses a matrix to organise coded data systematically before comparing across themes and groups. Relationship analysis can be conducted on a matrix by examining which cells consistently contain similar content, which are consistently empty, and which show contrasting patterns. A matrix makes thematic relationships visible and auditable.

Network and mapping approaches

Some qualitative researchers construct conceptual maps or mind maps of how themes relate to each other, as part of the analytical process rather than as a presentation tool. These maps are typically produced through reflection on the coded data rather than through formal calculation, and they are revised as the analysis develops. The network diagrams produced by quantitative co-occurrence analysis are one formalisation of this approach.

Narrative analysis

When the research question concerns how experiences unfold over time, narrative analysis examines the structure of participants' accounts: the sequence of events, the turning points, and the resolution (or lack of one). Sequence relationships are best captured through narrative analysis, because they require attention to the temporal structure of participants' accounts rather than just the presence or absence of themes.

04Common pitfalls when interpreting thematic relationships

Confusing association with causation. Two themes that frequently co-occur in the same segments are associated, not causally linked. The association warrants further investigation and interpretation, but it does not establish that one theme caused the other. Present co-occurrence findings as patterns that require explanation, not as evidence of causal mechanisms.

Over-interpreting small counts. A co-occurrence count of two means only two passages in the entire dataset had both codes. That is usually not enough to support a claim about thematic connection. Patterns worth reporting should appear in at least three to four segments across at least two or three different transcripts.

Treating all co-occurrence as meaningful. Some codes co-occur frequently simply because they are both common. "Financial pressure" and "Family obligations" may appear together often simply because both appear in nearly every transcript, not because they are conceptually linked. Adjust your interpretation for the base rates of each code.

Ignoring the content of co-occurring segments. Two codes appearing together does not tell you how they are related. Always read the actual segments where both codes appear before interpreting the relationship. The content of those segments determines whether the co-occurrence is conceptually meaningful.

05How AnalyZ Solutions supports relationship analysis

Relationship analysis is spread across two modules. Explore Themes handles the quantitative face of relationships, showing co-occurrence, distinctiveness, and segment length through charts and tables. Analyze Relationships handles the structural face, showing how codes connect visually, how they cluster across transcripts, and how they sequence within interviews. Run both after coding is complete: relationship analysis on partial data produces patterns that shift as more coding is added.

In Explore Themes: Theme Connectedness

The Theme Connectedness card has two views toggled within it. The Co-occurrence view shows a heatmap of how many coded segments carry both Code A and Code B simultaneously, alongside a ranked table of the strongest code pairs. To follow up a high co-occurrence finding, use the Boolean query builder in Review Codes with the condition "Code A AND Code B" to retrieve all segments where both codes appear. Read these segments to understand the nature of the relationship before interpreting it.

The Distinctiveness view identifies code pairs that almost never appear together. A distinctiveness score of 1.0 means two codes never co-occur; 0.0 means they always appear together. High distinctiveness between two codes that should conceptually overlap may indicate a coding inconsistency worth investigating.

In Explore Themes: Segment Length

The Segment Length card shows the average word count of coded segments for each code. Themes with longer average segments are topics participants elaborated on at length. Longer segments may indicate emotional significance, complexity, or ambivalence. Shorter segments may indicate topics mentioned in passing or taken for granted.

In Analyze Relationships: Network Map

The Network Map visualises co-occurrence data as a draggable graph. Each code is a node; edges between nodes represent co-occurrence, with thicker edges indicating higher counts. Node size reflects how frequently the code appears overall. Central nodes are analytically significant: they represent themes connected to many others. Drag nodes to explore the structure and export the map as an image for reporting.

In Analyze Relationships: Thematic Clusters

Clustering groups codes by how similarly they are distributed across transcripts. Codes that tend to appear in the same interviews cluster together, even if they do not always appear in the same segment. Clusters are candidates for higher-level themes in the write-up, but always verify that a cluster makes substantive sense by reading the coded segments before accepting it as analytically meaningful.

In Analyze Relationships: Code Sequence

The Code Sequence analysis uses a Sankey diagram to show which codes tend to follow which others within transcripts. Wide bands between codes indicate that Code A is frequently followed by Code B. Self-transitions, where the same code appears in consecutive segments, are filtered out so the diagram shows genuine movement between themes. This can reveal implicit narrative structures: how participants move from topic to topic, and what tends to lead to what in their accounts.

In Analyze Relationships: Configurational Analysis

Configurational Analysis groups transcripts by their thematic profile: which codes are present and which are absent in each transcript. Transcripts sharing the same configuration form a participant type. Codes are classified as Universal (present in every transcript), Differentiating (present in some but not all), or Rare (present in only one transcript). The differentiating codes are where the most analytically useful between-participant comparison happens.

Uncovering relationships between themes using AnalyZ Solutions

06Interpreting the results: from patterns to findings

The outputs of relationship analysis are patterns, not findings. Every pattern requires interpretation before it can be reported as a finding. The following questions guide the move from pattern to finding.

For co-occurrence patterns

What do these two themes have in common? Is one a cause, consequence, or context of the other? Are they two expressions of the same underlying experience? Read the segments where both codes appear and answer these questions from the data. In the vocational training example, reading segments where "Financial pressure" and "Family obligations" co-occur might reveal that in most cases, the family obligation described is specifically financial: a participant had to take on paid work to support the family. This reading suggests "Family obligations" is often a financial obligation, not a separate category.

For network centrality

What does it mean that this theme connects to many others? Does it represent a root cause, a cross-cutting concern, or an organising framework through which participants make sense of other experiences? In the girls' education study, if "Safety concerns" is the most central node, connected to distance, cost, family attitudes, and community norms, that centrality is a substantive finding: safety is not one barrier among many but the underlying concern that organises all the others.

For clustering patterns

Do the codes in a cluster tell a coherent story when read together? Would it make sense to group them under a single higher-level theme in the report? A cluster of "Distrust of institutions", "Prior negative experience", and "Preference for informal channels" might be interpretable as a higher-level theme called "Institutional distance." But verify this by reading the segments: if the distrust is specifically about health institutions and the informal channels are about money lending, they may not belong together despite appearing in the same transcripts.

For sequence patterns

Does the sequence reflect a causal process, a temporal progression, or simply the structure of the interview guide? If "Low confidence" consistently follows "Perceived irrelevance of training content" in the sequence diagram, does reading those segments confirm that participants described losing confidence because they found the content irrelevant? Or does "Low confidence" appear later simply because it was discussed later in the interview, after a question that prompted reflection on personal development?

For distinctiveness patterns

Is the distinctiveness a substantive finding (these two themes are genuinely separate in participants' experience) or an analytical artefact (the two codes were applied inconsistently by different coders)? Check the code definitions and the coded segments before interpreting high distinctiveness as a finding.

For configurational patterns

Do the participant types identified in the configurational analysis correspond to meaningful groups in the data? A configuration shared by two or more participants may reflect a genuine type, or it may reflect a coincidence of coding. Read the transcripts within each configuration before naming the type. Pay particular attention to the differentiating codes: these are what separates one participant type from another, and reading the segments for these codes across types reveals what the variation actually means in participants' own terms.

07Reporting thematic relationships

Thematic relationships are typically reported within the findings narrative, not as separate quantitative outputs. The co-occurrence data provides evidence; the interpretive text provides the argument.

A typical reporting sentence might be: "Financial pressure and family obligations were frequently discussed together by participants who had left the programme, suggesting they represent a compound barrier rooted in household resource scarcity rather than two independent factors." This integrates the quantitative pattern (co-occurrence) with the interpretive finding (compound barrier) and its implications (not two independent factors).

Network diagrams and co-occurrence heatmaps can be included as figures in a report when they add clarity. They are particularly useful in presentations and stakeholder reports, where a visual of the thematic structure is more accessible than a prose description.

Frequently Asked Questions
How many coded segments do I need for co-occurrence analysis to be meaningful?
There is no formal minimum, but co-occurrence counts based on very few segments are fragile. A co-occurrence count of two means only two passages in the entire dataset had both codes, which is usually not enough to support a claim about thematic connection. Patterns worth reporting should appear in at least three to four segments and involve at least two or three different transcripts.
Can co-occurrence be used as evidence that one theme causes another?
No. Co-occurrence shows association. Two themes appear together frequently: that is all the analysis shows. The interpretation of what that association means, and whether one might lead to the other, requires reading the segments and applying your theoretical framework. Report co-occurrence as a pattern that warrants interpretive attention, not as evidence of a causal relationship.
The network diagram is very dense. How do I read it?
A dense network with many connections usually indicates that codes are not sufficiently distinct from each other. Consider whether some codes can be merged, or whether boundaries between codes need sharpening. A well-defined codebook typically produces a network with some central themes connected to many others and some more peripheral themes with fewer connections, not a uniform web where everything connects to everything.
What does it mean if clustering groups codes that seem conceptually unrelated?
It means those codes tend to appear in the same interviews, even if they do not directly relate. This could reflect a participant characteristic driving both: if rural participants tend to discuss both transportation and seasonal income, those two codes might cluster even though they are not conceptually linked. Check whether the cluster corresponds to a subgroup in your data before interpreting it as a thematic relationship.
When should I run relationship analysis?
After coding is complete across all transcripts. Running relationship analysis on partial data produces patterns that will change as more coding is added, which can send the analysis in unproductive directions. The one exception is exploratory use during analysis to get a preliminary sense of the data's structure, with the understanding that results are provisional.

Explore how your themes connect

Open Explore Themes or Analyze Relationships in AnalyZ Solutions. Free to sign up, free to analyze.

Get started
Related Guides