Qualitative Analysis/ Developing Personas
Analyze Relationships, Personas

How to create personas in qualitative research

Participant personas turn analytical types into profiles that programme designers, policy makers, and field teams can actually use. This guide covers what personas are, when they are useful, how to derive them rigorously from qualitative data, and how to build them in AnalyZ Solutions with a worked example.

11 min read Practical guide Intermediate

Personas are composite profiles of participant types. In programme design and evaluation, they help teams move from abstract analytical findings to concrete understanding of who their participants are, what distinguishes one group from another, and what each group needs. A well-constructed persona is not a character sketch, it is a structured summary of evidence about a distinct type of participant in the data.

This guide covers how to construct personas rigorously from qualitative data, when they add value, and what distinguishes an analytically grounded persona from an invented one.

01What a persona is and what it is not

A persona is a composite representation of a group of participants who share a meaningful pattern of characteristics. In qualitative research, that pattern comes from the data: the themes participants discussed, the perspectives they expressed, the experiences they described. A persona gives that pattern a name, a profile, and a voice, usually in the form of a representative quote.

Personas are widely used in UX design, where they are sometimes constructed from assumptions, demographic generalisations, or small convenience samples. In qualitative evaluation research, that approach is insufficient. A persona that is not grounded in systematic analysis of the data is a hypothesis at best and a stereotype at worst. The validity of a persona depends entirely on the rigour of the analysis that produced the types it represents.

Personas describe types, not individuals. A persona named "The Reluctant Adopter" is not a description of any one participant. It is a composite of the characteristics shared by a group of participants whose coded data shows a similar thematic profile. Individual participants in that group may differ on some dimensions; the persona captures what they share.

02When personas are useful in qualitative research

Personas are most useful when the research has produced analytically meaningful types and when those types need to be communicated to audiences who will not read the full analysis. They are particularly appropriate for:

Personas are less useful when the research has not produced meaningful types, when all participants share a very similar profile, or when the sample is too small or too homogeneous to sustain differentiation. In those cases, the analytical finding is the homogeneity itself, and forcing a persona typology onto it misrepresents the data.

03The evidence base for personas: why grounded profiles work better

The critique of personas in social research is usually directed at invented personas: profiles constructed from assumptions, demographic averages, or researcher intuition rather than from systematic analysis of participant data. These have been rightly criticised for reinforcing stereotypes, missing within-group variation, and giving false precision to what are essentially guesses.

Grounded personas avoid these problems by deriving the type structure from systematic analysis. The groups are not decided in advance, they emerge from the data. The profiles are not written from intuition, they are built from coded segments, differentiating themes, and participant attributes that the researcher has documented. The representative quote is not invented, it is selected from actual transcripts.

This does not make grounded personas infallible. The quality of the persona depends on the quality of the coding, the appropriateness of the analytical method used to derive the types, and the care with which the researcher interprets what the type means. But a persona built on systematic configurational analysis of coded qualitative data is a fundamentally different product from one built on assumptions, and the difference should be made explicit when presenting personas to stakeholders.

04From thematic positions to personas: the analytical foundation

The starting point for qualitative personas is not the transcript but the coded segment. A coded segment is the actual passage of text where a participant engages with a theme. Across all your transcripts, coded segments cluster into patterns: certain code combinations recur together, certain ways of framing an issue appear repeatedly. These recurring patterns are thematic positions, and they are the raw material from which personas are built.

This approach works for both individual interviews and focus group discussions. Because the unit of analysis is the segment rather than the transcript or the speaker, it does not require speaker attribution in FGDs. A focus group may contain participants who hold very different thematic positions, the analysis surfaces those positions from the content itself, regardless of who said what.

What a thematic position is

A thematic position is a coherent way of engaging with a topic that recurs across coded segments. It is defined by which codes co-occur within segments and how much of the coded content is devoted to each theme. Two participants in the same focus group may both discuss environmental behaviour, but one consistently frames it in terms of structural barriers while the other frames it in terms of personal agency. These are different thematic positions, and they warrant different personas.

Thematic positions differ from configurational types in an important way: they are derived from the relational structure of the themes within segments, not from a binary profile of which codes appeared anywhere in a transcript. This makes them more analytically precise and more suitable for data from group discussions.

How positions become personas

The move from a thematic position to a persona involves three steps. First, the position is identified through clustering: coded segments are grouped by how similar their code combinations are. Second, the position is characterised: which codes are dominant in this cluster, what do the segments actually say, and what participant attributes are associated with this cluster. Third, the position is named and presented as a communicable persona.

05Worked example: environmental behaviour study

The following example uses data from the environmental behaviour study used throughout this article series. Three codes were selected for persona building: Environmental awareness, Eco-socially responsible behaviour, and Perceived behavioural control. These three codes capture meaningfully different orientations toward environmental action.

The analysis clustered all coded segments carrying at least one of these codes by their code combination profiles. Silhouette scoring suggested three thematic personas as the optimal solution, with a score of 0.71 indicating strong separation between positions.

The Aware Non-Actor

20 segments

Predominantly female · Mostly rural

Segments in this position are dominated by environmental awareness (present in 100% of segments) but eco-socially responsible behaviour appears in only 25% of segments. This position captures a way of engaging with the topic where knowledge and concern are well developed but behaviour change is not yet reflected in the account. The gap between awareness and action is the defining analytical feature.

"The knowledge is there. But knowing something and actually changing your behavior because of it, these are two very different things."

FGD_01_Urban_YoungAdults

Thematic profile

Environmental awareness 100% Eco-responsible behaviour 25%

The Committed Actor

18 segments

Mostly female · Mostly rural

This position clusters segments where eco-socially responsible behaviour (100%) and perceived behavioural control (100%) appear together. Participants expressing this position not only describe behaviour change but articulate a sense of personal agency over that change. The co-occurrence of these two codes defines a qualitatively different relationship with the topic from The Aware Non-Actor.

"That caring about the environment and caring about economic development are not opposites. The framing that you have to choose one or the other is a false choice."

IDI_04_Omar_AlNatour

Thematic profile

Eco-responsible behaviour 100% Perceived behavioural control 100%

The Constrained Agent

7 segments

Mostly male · Mostly urban

The third position is defined by high perceived behavioural control (100%) with more variable presence of the other two codes. These segments describe participants with a strong sense of personal agency but whose accounts do not consistently reflect behaviour change. This position is analytically important because it surfaces a group whose sense of control is not yet translating into action, a different programme challenge from The Aware Non-Actor where agency is the missing ingredient.

"We are not careless or ignorant. We live in harder conditions with fewer options. When I see clean streets in West Amman and compare them to where I live, I think about what is actually possible."

FGD_03_MiddleAged_Men

Thematic profile

Perceived behavioural control 100% Environmental awareness 57% Eco-responsible behaviour 43%
Note on FGD data. Two of the three personas above draw on segments from FGD transcripts. The personas are not attributed to individual speakers. They represent thematic positions that appeared in the group discussion, not profiles of specific participants.

06How to do it in AnalyZ Solutions

The Persona Builder is in the Analyze Relationships module. It identifies thematic personas from the co-occurrence patterns in your coded segments and suggests the most analytically meaningful number of personas using silhouette scoring. Run it after coding is complete across all transcripts.

Step 1: Select themes

Choose which codes define the thematic space for persona building. These should be codes that capture meaningfully different orientations toward your research topic, not every code in the codebook. Select three to six codes that vary meaningfully across participants. As you select codes, the interface shows how many segments carry at least one of the selected codes.

Step 2: Review the suggestion

The analysis runs hierarchical clustering on all segments for solutions from 2 to 8 personas and scores each using the silhouette score. A score above 0.7 indicates strong separation between personas; between 0.5 and 0.7 is reasonable; below 0.5 suggests the positions overlap and the code selection should be reconsidered. The score table shows all solutions with colour-coded bars. Click any row to select a different number of personas if the suggested solution does not fit the analytical needs of the study.

Step 3: Explore and name personas

Persona cards appear in a grid. Each card shows an auto-generated descriptive name derived from the frequency profile, the dominant codes with their percentage of segments, participant characteristics from transcript tags (Predominantly female, Mostly rural), and a sample segment. The name field is editable. Expand any card to see the full thematic frequency profile. Click a theme bar to open a segment panel showing all segments in that persona where that code appears. Use the Merge into dropdown to combine two personas if they are too similar to be analytically distinct.

Add to AnalyZense sends all persona profiles with their frequency data, silhouette scores, and a plain-language interpretation to a Sense Making project. Add to Canvas sends a frequency profile table showing what percentage of segments in each persona carry each selected code.

07Common mistakes to avoid

Naming personas before reading the data. The name should come after reading the coded segments for the differentiating codes in that type's transcripts, not before. A name decided in advance shapes how the researcher reads the data, which introduces confirmation bias into the profile description.

Writing profiles from memory rather than from the segments. The profile of a persona should be written while looking at the coded segments, not reconstructed from memory afterward. The segments are the evidence; the profile is the interpretation. Without the evidence in front of you, the profile will drift toward generalisation.

Presenting singleton types as personas. A type with one transcript is a unique case, not a pattern. It can be noted as an outlier or an interesting individual account, but presenting it as a persona implies it represents a group when it does not. Consider merging singletons with the most similar type, or acknowledging them as individual cases in the report rather than as a persona.

Ignoring internal variation in merged types. When two types are merged, the codes that differ between them create internal variation in the merged persona. These are the dimensions on which participants within the persona disagree. A profile that ignores this variation produces a misleadingly homogeneous persona. Flag the variable codes explicitly and note what the variation means when presenting the persona.

Treating personas as fixed across time. Personas describe the participant population as it appears in the data collected at a particular point. If the programme context changes, if new participant groups are reached, or if a longitudinal study adds new rounds of data, the personas should be revisited. A persona built on baseline data may not accurately represent participants at endline.

Using personas as the only output. Personas are a communication tool, not a replacement for the full analysis. A report that presents only personas without the underlying thematic analysis, frequency data, or configurational matrix is not giving its audience the evidence they need to scrutinise the findings. Personas work best as a layer on top of a full qualitative report, not as a substitute for one.

Frequently Asked Questions
How many personas is the right number?
There is no fixed number, but three to five is typically the range within which personas remain distinct and communicable. Below three, the personas may not capture meaningful variation. Above six, they become difficult to remember and communicate, and the distinctions between them may be too fine-grained to be actionable. If the configurational analysis produces more than six types, consider merging the most similar ones, particularly singletons and types that differ on only one or two differentiating codes.
Can I develop personas without running a formal configurational analysis?
Yes, though the result will be less systematic. Some researchers develop personas through close reading and iterative grouping, reading all transcripts, identifying recurring profiles, and grouping cases by hand. This approach is valid but more susceptible to confirmation bias and harder to audit. Configurational analysis provides a structured, transparent basis for the types that makes the grouping decisions visible and reproducible.
What is the difference between a persona and a case study?
A case study describes one specific individual or site in depth, typically including context, history, and a narrative account. A persona is a composite of multiple cases sharing a similar profile, it does not describe any one individual and is not a narrative. Case studies are selected purposively to represent interesting or typical cases; personas emerge from systematic grouping of the full dataset. Both are useful, and a well-structured qualitative report may include both: personas to communicate the typology across the full sample, and case studies to illustrate specific types in depth.
How do I handle a participant who does not fit any persona well?
Most participants will fit their assigned type reasonably well, but outliers exist in every dataset. If a participant shares only some characteristics of their assigned type, note this when presenting the persona rather than ignoring the fit. A participant who is borderline between two types is analytically interesting, they may reveal something about the boundary between the types or about dimensions of variation the typology does not fully capture. Do not force a clean fit where the data does not support one.
Should personas be shared with programme participants?
Sometimes, with care. Sharing personas with community members or programme staff can be a powerful participatory validation exercise, participants recognise themselves and others in the profiles and can comment on whether the types ring true. However, personas should always be anonymised, and care should be taken that no persona is presented in a way that could be stigmatising or that participants would find reductive. If you plan to share personas with participants, consider involving them in naming and describing the types as part of the research process.
How do I cite the source of a persona in a report?
Cite the analysis that produced the types: the configurational analysis, the number of transcripts, the codes used as the basis for grouping, and the method used to derive the types. A footnote such as "Personas derived from configurational analysis of 8 interview transcripts coded across 9 themes; types identified on the basis of 5 differentiating codes" gives readers enough information to assess the basis for the typology without requiring them to read the full technical appendix.

Build your participant personas

Open Analyze Relationships and select Persona Builder in AnalyZ Solutions.

Get started
Related Guides