Participant personas turn analytical types into profiles that programme designers, policy makers, and field teams can actually use. This guide covers what personas are, when they are useful, how to derive them rigorously from qualitative data, and how to build them in AnalyZ Solutions with a worked example.
Personas are composite profiles of participant types. In programme design and evaluation, they help teams move from abstract analytical findings to concrete understanding of who their participants are, what distinguishes one group from another, and what each group needs. A well-constructed persona is not a character sketch, it is a structured summary of evidence about a distinct type of participant in the data.
This guide covers how to construct personas rigorously from qualitative data, when they add value, and what distinguishes an analytically grounded persona from an invented one.
A persona is a composite representation of a group of participants who share a meaningful pattern of characteristics. In qualitative research, that pattern comes from the data: the themes participants discussed, the perspectives they expressed, the experiences they described. A persona gives that pattern a name, a profile, and a voice, usually in the form of a representative quote.
Personas are widely used in UX design, where they are sometimes constructed from assumptions, demographic generalisations, or small convenience samples. In qualitative evaluation research, that approach is insufficient. A persona that is not grounded in systematic analysis of the data is a hypothesis at best and a stereotype at worst. The validity of a persona depends entirely on the rigour of the analysis that produced the types it represents.
Personas are most useful when the research has produced analytically meaningful types and when those types need to be communicated to audiences who will not read the full analysis. They are particularly appropriate for:
Personas are less useful when the research has not produced meaningful types, when all participants share a very similar profile, or when the sample is too small or too homogeneous to sustain differentiation. In those cases, the analytical finding is the homogeneity itself, and forcing a persona typology onto it misrepresents the data.
The critique of personas in social research is usually directed at invented personas: profiles constructed from assumptions, demographic averages, or researcher intuition rather than from systematic analysis of participant data. These have been rightly criticised for reinforcing stereotypes, missing within-group variation, and giving false precision to what are essentially guesses.
Grounded personas avoid these problems by deriving the type structure from systematic analysis. The groups are not decided in advance, they emerge from the data. The profiles are not written from intuition, they are built from coded segments, differentiating themes, and participant attributes that the researcher has documented. The representative quote is not invented, it is selected from actual transcripts.
This does not make grounded personas infallible. The quality of the persona depends on the quality of the coding, the appropriateness of the analytical method used to derive the types, and the care with which the researcher interprets what the type means. But a persona built on systematic configurational analysis of coded qualitative data is a fundamentally different product from one built on assumptions, and the difference should be made explicit when presenting personas to stakeholders.
The starting point for qualitative personas is not the transcript but the coded segment. A coded segment is the actual passage of text where a participant engages with a theme. Across all your transcripts, coded segments cluster into patterns: certain code combinations recur together, certain ways of framing an issue appear repeatedly. These recurring patterns are thematic positions, and they are the raw material from which personas are built.
This approach works for both individual interviews and focus group discussions. Because the unit of analysis is the segment rather than the transcript or the speaker, it does not require speaker attribution in FGDs. A focus group may contain participants who hold very different thematic positions, the analysis surfaces those positions from the content itself, regardless of who said what.
A thematic position is a coherent way of engaging with a topic that recurs across coded segments. It is defined by which codes co-occur within segments and how much of the coded content is devoted to each theme. Two participants in the same focus group may both discuss environmental behaviour, but one consistently frames it in terms of structural barriers while the other frames it in terms of personal agency. These are different thematic positions, and they warrant different personas.
Thematic positions differ from configurational types in an important way: they are derived from the relational structure of the themes within segments, not from a binary profile of which codes appeared anywhere in a transcript. This makes them more analytically precise and more suitable for data from group discussions.
The move from a thematic position to a persona involves three steps. First, the position is identified through clustering: coded segments are grouped by how similar their code combinations are. Second, the position is characterised: which codes are dominant in this cluster, what do the segments actually say, and what participant attributes are associated with this cluster. Third, the position is named and presented as a communicable persona.
The following example uses data from the environmental behaviour study used throughout this article series. Three codes were selected for persona building: Environmental awareness, Eco-socially responsible behaviour, and Perceived behavioural control. These three codes capture meaningfully different orientations toward environmental action.
The analysis clustered all coded segments carrying at least one of these codes by their code combination profiles. Silhouette scoring suggested three thematic personas as the optimal solution, with a score of 0.71 indicating strong separation between positions.
20 segments
Segments in this position are dominated by environmental awareness (present in 100% of segments) but eco-socially responsible behaviour appears in only 25% of segments. This position captures a way of engaging with the topic where knowledge and concern are well developed but behaviour change is not yet reflected in the account. The gap between awareness and action is the defining analytical feature.
"The knowledge is there. But knowing something and actually changing your behavior because of it, these are two very different things."
FGD_01_Urban_YoungAdults
Thematic profile
18 segments
This position clusters segments where eco-socially responsible behaviour (100%) and perceived behavioural control (100%) appear together. Participants expressing this position not only describe behaviour change but articulate a sense of personal agency over that change. The co-occurrence of these two codes defines a qualitatively different relationship with the topic from The Aware Non-Actor.
"That caring about the environment and caring about economic development are not opposites. The framing that you have to choose one or the other is a false choice."
IDI_04_Omar_AlNatour
Thematic profile
7 segments
The third position is defined by high perceived behavioural control (100%) with more variable presence of the other two codes. These segments describe participants with a strong sense of personal agency but whose accounts do not consistently reflect behaviour change. This position is analytically important because it surfaces a group whose sense of control is not yet translating into action, a different programme challenge from The Aware Non-Actor where agency is the missing ingredient.
"We are not careless or ignorant. We live in harder conditions with fewer options. When I see clean streets in West Amman and compare them to where I live, I think about what is actually possible."
FGD_03_MiddleAged_Men
Thematic profile
The Persona Builder is in the Analyze Relationships module. It identifies thematic personas from the co-occurrence patterns in your coded segments and suggests the most analytically meaningful number of personas using silhouette scoring. Run it after coding is complete across all transcripts.
Choose which codes define the thematic space for persona building. These should be codes that capture meaningfully different orientations toward your research topic, not every code in the codebook. Select three to six codes that vary meaningfully across participants. As you select codes, the interface shows how many segments carry at least one of the selected codes.
The analysis runs hierarchical clustering on all segments for solutions from 2 to 8 personas and scores each using the silhouette score. A score above 0.7 indicates strong separation between personas; between 0.5 and 0.7 is reasonable; below 0.5 suggests the positions overlap and the code selection should be reconsidered. The score table shows all solutions with colour-coded bars. Click any row to select a different number of personas if the suggested solution does not fit the analytical needs of the study.
Persona cards appear in a grid. Each card shows an auto-generated descriptive name derived from the frequency profile, the dominant codes with their percentage of segments, participant characteristics from transcript tags (Predominantly female, Mostly rural), and a sample segment. The name field is editable. Expand any card to see the full thematic frequency profile. Click a theme bar to open a segment panel showing all segments in that persona where that code appears. Use the Merge into dropdown to combine two personas if they are too similar to be analytically distinct.
Add to AnalyZense sends all persona profiles with their frequency data, silhouette scores, and a plain-language interpretation to a Sense Making project. Add to Canvas sends a frequency profile table showing what percentage of segments in each persona carry each selected code.
Naming personas before reading the data. The name should come after reading the coded segments for the differentiating codes in that type's transcripts, not before. A name decided in advance shapes how the researcher reads the data, which introduces confirmation bias into the profile description.
Writing profiles from memory rather than from the segments. The profile of a persona should be written while looking at the coded segments, not reconstructed from memory afterward. The segments are the evidence; the profile is the interpretation. Without the evidence in front of you, the profile will drift toward generalisation.
Presenting singleton types as personas. A type with one transcript is a unique case, not a pattern. It can be noted as an outlier or an interesting individual account, but presenting it as a persona implies it represents a group when it does not. Consider merging singletons with the most similar type, or acknowledging them as individual cases in the report rather than as a persona.
Ignoring internal variation in merged types. When two types are merged, the codes that differ between them create internal variation in the merged persona. These are the dimensions on which participants within the persona disagree. A profile that ignores this variation produces a misleadingly homogeneous persona. Flag the variable codes explicitly and note what the variation means when presenting the persona.
Treating personas as fixed across time. Personas describe the participant population as it appears in the data collected at a particular point. If the programme context changes, if new participant groups are reached, or if a longitudinal study adds new rounds of data, the personas should be revisited. A persona built on baseline data may not accurately represent participants at endline.
Using personas as the only output. Personas are a communication tool, not a replacement for the full analysis. A report that presents only personas without the underlying thematic analysis, frequency data, or configurational matrix is not giving its audience the evidence they need to scrutinise the findings. Personas work best as a layer on top of a full qualitative report, not as a substitute for one.
Open Analyze Relationships and select Persona Builder in AnalyZ Solutions.
Get started