Structural Coding Explained (With Examples)
How – and why – to index your qualitative data by research question
The 30-second summary
Structural coding labels each segment of your data with the research question it answers. It is an indexing pass that makes a pile of transcripts searchable. It is not, on its own, your analysis.
- Your codes come from your interview guide, written before you read a single transcript. Usually one per research question, sometimes finer.
- The code says where a passage belongs, not what it says. Two participants who contradict each other get the same structural code.
- It earns its keep on volume. With twenty transcripts and two coders, it is what lets you read everyone's answer to one question side by side.
- The real coding happens afterwards, inside each indexed segment. Treating the index as your findings is the mistake this technique invites.
On a small study, you'll likely be better off skipping it and going directly to in vivo or descriptive coding.
Navigate
A structural code is not a description of what someone said. It is an address: the research question that segment answers, attached to every chunk of every transcript, so the whole dataset becomes searchable by question rather than only readable end to end.
That is the whole payoff, and it is why the technique exists. But coding by question feels like analysis, and it isn’t. What you have at the end is a filing system, and the filing is the easy half.
So this post covers what a structural code actually labels, when the pass is worth the effort, what one looks like in a real transcript, the four steps, and the mistake I see most often when someone brings me a structurally coded dataset.
Notice that the orange segments sit at a different depth in every transcript. That scatter is the problem structural coding solves, and the panel on the right is what it gives you.
What is structural coding?
Structural coding is a first-cycle technique in which your codes are your research questions. You write them before you read a transcript, and you apply each one to a whole chunk of data: a question and the answer it produced.
Rhode Island College’s open textbook on social data analysis puts it plainly. In structural coding, “codes indicate which specific research question, part of a research question, or hypothesis is being addressed by a particular segment of text”.
The distinguishing feature is that the code does not describe the content, which trips people up because every other first-cycle technique does. Descriptive coding names the topic of a passage. In vivo coding uses the participant’s own words. Process coding names the action. A structural code says only where the passage belongs. Two participants who flatly contradict each other get the same code, because they were answering the same question.
That makes the pass deductive in shape: the codes are fixed in advance, drawn from the constructs your questions are built on rather than from your data (inductive vs deductive coding covers the distinction properly). It is a bookkeeping decision, not a theoretical commitment. Nothing about a structural pass stops the coding that follows it from being entirely inductive.
When should you use it?
Structural coding costs you a full pass through the data and gives you nothing to write about. That is a poor trade on a small study and an excellent one on a large one, so the decision comes down to four conditions.
- You have a lot of transcripts. Around fifteen and up is where it starts paying. The real test isn’t a number though: it’s whether you can still hold the shape of the dataset in your head. If you can, you don’t need an index.
- Your interviews were semi-structured. The codes have to come from somewhere, and a guide is that somewhere. Unstructured interviews wander by design, so there’s nothing to index against. Our guide to structured, semi-structured and unstructured interviews covers the difference.
- More than one person is coding. A structural pass is the easiest thing in qualitative analysis to agree on, which makes it a good place for a team to calibrate before the interpretive work starts.
- Your plan involves comparing across cases. If your findings need to say “most participants described X”, you need to be able to gather everyone’s X in one place.
Miss those conditions and the technique turns into busywork. On eight unstructured interviews, or a single narrative case where the sequence inside one person’s account is the point, you’re likely better off skipping the index and coding properly straight away. Narrative analysis treats that whole account as the unit, and slicing it up by question destroys what you’re studying.
What it looks like in a transcript
Here is a short stretch of an interview from a study of secondary school teachers and a new online grading platform. The study asked three research questions: how the platform was introduced, what changed in day-to-day work, and what made it hard to sustain.
| Interview extract (abridged) | Structural code |
|---|---|
| “We got a half-day of training in the last week of term, then it went live in January.” | RQ1: How it was introduced |
| “Nobody asked us first. We heard about it in a staff meeting.” | RQ1: How it was introduced |
| “Marking takes longer now, honestly, but the parents can see everything, which they like.” | RQ2: What changed day to day |
| “And that’s the thing. Once parents can see it, they email you about it at ten at night.” | RQ3: What made it hard to sustain |
| “Sorry, going back, the training itself was fine. It was the timing that killed it.” | RQ1: How it was introduced |
Notice the last two rows. She drifted from what changed into what made it hard, then doubled back to the training half a page later. That is how people talk, and it is exactly what the pass is for: her answer to the first research question is sitting in three separate places, and after the pass you can read all three together.
Notice too what the code column doesn’t tell you. Nothing in it says marking takes longer, or that parents email at night. The content is still in the extract, uncoded. Getting it out is the next pass, and that is where the analysis lives. If you want to see what that next pass looks like, we have worked coding examples using several techniques on the same kind of data.
How to do structural coding
The pass runs in four steps.
- Write the code list from your interview guide, before you read anything. Your codes come from the constructs each research question is about. One per research question is the usual starting point, and where a single question probes two or three separate things, the interview questions underneath it give you finer codes. Either level works, as long as you pick one and hold to it. The count follows your guide, so most studies land somewhere under fifteen. If you’re heading for forty, you have started descriptive coding by accident.
- Give every code a definition and an inclusion rule. A code name on its own drifts within a week, and between two coders it drifts immediately. CUNY’s open chapter on coding in qualitative analysis sets out what a codebook entry needs: the code, a definition, an example, and criteria for what it excludes.
- Code whole exchanges, not lines, and don’t code your own questions. The unit is a question and everything the participant said in response to it, digressions included. Your prompts are not data. Keep the question number beside every extract as you go: it costs nothing, and it lets you go finer than the code when you need to.
- Retrieve by code and read across participants. Pull every segment carrying one structural code into one place and read them together. This is the first moment the pass has given you anything, and it is worth doing before you code a single line further.
Mechanically, a spreadsheet handles this fine. One tab per transcript, with columns for the participant, the question number, the extract and the structural code, then a master sheet that stacks every tab so you can filter the lot by code. It’s dull work. It is also what makes step four take ten minutes instead of a weekend.
Why it is not your analysis
A structurally coded dataset looks finished. Every extract carries a label, the spreadsheet is tidy, and the labels line up neatly with the research questions. That tidiness is the trap.
“But surely my themes should answer my research questions?”
In the write-up, yes. They shouldn’t be built that way.
The sequence that works has three stages: index the data by question, analyze it on its own terms, then map what you found back onto your questions when you write the findings chapter. It’s the middle stage people drop, and structural coding makes dropping it feel reasonable, because the data is already sorted into question-shaped piles.
I see the results of that regularly. One theme per research question, named in slightly different words from the question. Codes pre-assigned to a question before the transcript was read. Separate analyses run question by question, so nothing is ever compared across them.
The cost is specific rather than abstract. Your participants did not organize their experience around your research questions, so the most valuable finding in a qualitative study is usually the one sitting across two of them. Something that showed up under question one and question three, and means something different once you see both. An analysis run question by question is structurally incapable of noticing it.
So after the structural pass, code inside each segment with a technique that actually labels content: descriptive, in vivo, process or values coding, depending on what your question is about. Build your categories and themes across the whole dataset, with the structural codes ignored while you do it. If you are working in grounded theory, that building step is axial coding. Then bring them back at the very end as a coverage check.
That check is worth more than it sounds. Does every research question have something under it? A question with nothing to say is either a finding in itself or a sign the interview guide didn’t work, and you want to know which before your thematic analysis is written up rather than after.
The other meaning of the term
One warning, because you’ll run into it within about two searches. Some sources use “structural coding” to mean something else entirely: labeling how a piece of text is built rather than which question it answers. Codes like “opening greeting” or “call to action” name the moves in a text, which sits much closer to discourse analysis than to anything in this post.
Both usages are in circulation. The one described here, coding by research question, is the standard usage in qualitative methods texts and almost certainly what your methodology chapter means. If a source starts talking about genre, form or how something is written, it’s using the other sense. The two are not interchangeable, so check which one your reading list means.
Still have questions?
Is structural coding the same as deductive coding?
Not quite. The codes are fixed in advance, which makes the pass deductive in shape, but deductive coding usually means applying a theory or framework to see whether the data supports it. Structural codes come from your own interview guide and test nothing. You can run a structural pass and then code inductively inside it with no contradiction at all.
What if one answer belongs under two research questions?
Give it both codes. Structural codes are addresses rather than categories, and nothing says a segment can only have one. If a lot of your data ends up double-coded, take note: it usually means two of your research questions overlap more than you thought, and that’s a design issue you would rather find now than at the defense.
Can I use structural coding on open-ended survey responses?
Yes, and it’s quick, because each response already sits under the question that produced it. The honest check is whether it adds anything. If your survey has one box per question, you already have your index and the pass is bookkeeping you don’t need. It earns its place where a single open box invites people to cover several of your questions at once.
What if a participant never answers one of my questions?
Code the absence rather than skipping past it. A structural code with three participants under it and nine silences is telling you something, either about the topic or about how the question was worded, and both belong in your write-up. What you shouldn’t do is quietly narrow the research question later so that the gap disappears.
Can’t find your answer here? Ask a Grad Coach directly – the initial chat is free.
Speak with a friendly coach →