Skip to content
Grad Coach
Qualitative Data Coding

Open, Axial & Selective Coding

All three grounded theory phases, with examples from one dataset

By Derek Jansen (MBA) · Reviewed by Eunice Rautenbach (DTech)

Pick your reference style

  • APA Social sciences

    Jansen, D. (2026). Open, Axial & Selective Coding. Grad Coach. https://gradcoach.com/axial-coding/

  • MLA Humanities

    Jansen, Derek. "Open, Axial & Selective Coding." Grad Coach, Sep. 1, 2026, https://gradcoach.com/axial-coding/. Accessed 3 Sep. 2026.

  • Chicago History and arts

    Jansen, Derek. "Open, Axial & Selective Coding." Grad Coach. September 1, 2026. https://gradcoach.com/axial-coding/.

  • Harvard Author–date, general use

    Jansen, D. (2026) 'Open, Axial & Selective Coding', Grad Coach. Available at: https://gradcoach.com/axial-coding/ (Accessed: 3 September 2026).

  • Vancouver Medicine and science

    Jansen D. Open, Axial & Selective Coding [Internet]. Grad Coach; 2026 [cited 2026 Sep 3]. Available from: https://gradcoach.com/axial-coding/

  • IEEE Engineering and tech

    D. Jansen, "Open, Axial & Selective Coding," Grad Coach, 2026. [Online]. Available: https://gradcoach.com/axial-coding/. [Accessed: Sep. 3, 2026].

The 30-second summary

Grounded theory codes data in three phases. Open coding breaks the data apart, axial coding puts it back together, and selective coding picks the one account that holds it all. Each phase takes the previous phase's output as its input.

  • Open coding works on your transcripts, naming what each segment is about. Expect a long, unruly list.
  • Axial coding works on your code list, not your data, grouping codes into categories and stating how those categories relate.
  • Selective coding picks the core category and ties everything else to it. Plenty of postgrad projects stop before this, and say so.
  • Your code count should collapse at each pass. If a pass doesn't reduce, it renamed things rather than relating them.

Not sure grounded theory is your method? The second section is a two-minute check, and it points you toward thematic analysis if the answer is no.

Navigate

Grounded theory doesn’t code data once. It codes it three times, and each pass works on what the last one produced: open coding names what is in your transcripts, axial coding relates those names to each other, and selective coding picks the single account that holds the whole thing together.

That chain is what gets lost when the three are taught as separate techniques rather than as one sequence. Skipping its middle is the most common way a grounded theory chapter comes apart, because axial coding is the phase people find hardest and the one an examiner is most likely to ask about. Every qualitative study shows its codes and shows its themes; the invisible middle between two hundred of the first and five of the second is where the questions land.

In this post, I’ll work one dataset through all three phases, then cover the arithmetic of how many codes you should end up with, and whether the method’s best-known framework is compulsory. It isn’t.

Three panels side by side showing how the shape of a dataset changes across the coding phases. Panel one, open coding, transcripts to codes: twenty small identical gray chips in an even grid, captioned about 150 loose codes. An arrow leads to panel two, axial coding, codes to categories: the same chips gathered into four bordered clusters of four, with short orange lines drawn between the clusters to show the relationships, captioned about 20 categories, related. A second arrow leads to panel three, selective coding, categories to one account: four small empty boxes arranged around a single orange box labeled Core category, each joined to it by a line, captioned 1 core category.

Notice that what changes is the shape of the data, not just its volume. Open coding gives you pieces, axial coding gives you groups with connections between them, and selective coding gives you one thing everything else hangs off.


The three phases, side by side

The three get taught as a sequence and then described in language that makes them hard to tell apart. Here is what actually differs.

PhaseWhat it works onWhat you doWhat you end with
OpenYour transcriptsBreak the data into segments and name each oneA long, unruly list of codes
AxialYour code listGroup codes into categories, then state how the categories relateA handful of categories with the links between them named
SelectiveYour categoriesPick the core category and tie everything else to itOne account the whole dataset supports

The What it works on column is the one worth memorizing. Open coding is the only phase that touches your raw data. Axial coding barely looks at a transcript: its input is the code list open coding produced. That single fact resolves most of the confusion between them.

Two things the table hides, and both matter. The phases are not cleanly sequential; they’re iterative. Axial coding routinely sends you back to re-code data you thought you had finished with, and that’s the method working rather than failing. And the boundary between axial and selective is genuinely blurry, with plenty of published studies doing both in one pass and describing them separately afterwards.


Do you actually need grounded theory?

Worth settling before you learn three phases you may not need, because this will save many of you a lot of time.

Most students who tell me they’re doing axial coding aren’t building a theory. They’re analyzing interview data against a framework they chose in their literature review, which is a perfectly good study and isn’t grounded theory. Grounded theory means starting without a framework, generating one from the data, and continuing to collect data until the categories stop changing, which is what theoretical saturation means. If you already know which model you’re testing, that decision has been made.

The confusion usually shows up in one of two ways. Either the framework was picked as background reading and then never used as an analytic instrument, or a theory is expected to dictate the analysis rather than inform it. Both are worth catching early, because both are visible in the finished chapter.

So: if your aims are exploratory, there’s little existing theory, and you can still collect data, the three phases below fit. If you’re applying or testing something that already exists, you want deductive coding and probably thematic analysis, whose coding step is thematic coding. You can still borrow axial coding’s habit of stating relationships without claiming the whole method. Our guide to choosing a qualitative analysis method walks through the alternatives.


Open coding: breaking the data apart

Open coding is the first pass. You read through your transcripts and name what each segment is about, without a codebook and without worrying yet about how the codes relate.

You’ll also see this called initial coding, and the two mean the same thing. It’s a phase rather than a technique, which is the distinction that trips people up: in vivo coding, process coding and descriptive coding are all techniques you can use during open coding. Our qualitative data coding guide walks through each of them, and there are worked examples of several applied to the same kind of data.

The University of Connecticut’s research basics guide makes that relationship explicit, and it’s worth being clear about in your methodology chapter: “in vivo coding during the open coding phase” says something precise, where “I used open coding” doesn’t.

Three habits that make the phase work:

  • Stay close to the data. The temptation is to name codes at the level of your research questions. Resist it: that’s structural coding, it’s a different job, and doing it here means your categories were decided before you read anything.
  • Let the list get untidy. A hundred to two hundred codes is normal and isn’t a sign you’ve done it wrong. Reduction is the next phase’s work.
  • Code the surprises. Anything that doesn’t fit your expectations is the material grounded theory exists to find. A code applied once, to something that startled you, may end up mattering more than one applied forty times.

The mistake that ends the analysis early is treating that list as your findings. Ninety codes isn’t ninety findings; it’s ninety labels waiting to be related to each other. That process of relating is the next phase.


Axial coding: putting it back together

Axial coding takes the code list and works out how the codes relate. The name comes from working around an axis: you take one category and examine everything that bears on it.

Its output isn’t more codes. It’s categories, plus statements about how those categories relate. A category with no stated relationship to anything else is a heading, not a finding, and that distinction is the whole test of whether you’ve actually done this phase.

Strauss and Corbin’s coding paradigm is the best-known way of running it. It’s a set of six questions you put to one category at a time.

Six labeled boxes showing the coding paradigm. A horizontal spine runs from "Causal conditions: what led to this?" into "The phenomenon: what is going on here?", highlighted in orange, then on to "Strategies: what did people do about it?" and "Consequences: and what followed?". Two further boxes feed into the link between the phenomenon and the strategies: "Context: under what circumstances?" from above, and "Intervening conditions: what helped or hindered?" from below. Captioned after Strauss and Corbin, 1998.

Notice that every element is a question you put to one category, not a box you have to fill. That distinction matters, and the section on the paradigm below explains why.

Applied to the five codes above, three of them turn out to be things people do, one is the reason they do them, and one is the circumstance that makes it necessary.

The output is a sentence you could not have written before: people manage the condition invisibly because being seen as unreliable costs more than a missed dose does, and the office makes invisibility expensive.

Four steps, and the third is the one people skip.

  1. Sort your codes into candidate categories. Print them, or put them in a column and sort. Group the ones that seem to belong together and give each group a working name. Expect to move things twice.
  2. Put the paradigm questions to one category at a time. You aren’t filling in a form: if a category has no meaningful “consequences”, say so and move on.
  3. Write the relationship down as a sentence. Not a diagram, not an arrow: a sentence with a verb in it. “Fear of being seen as unreliable drives concealment strategies” is a finding. A box labeled “concealment” connected to a box labeled “fear” is a picture of one.
  4. Go back to the data and check. Take the sentence you just wrote and look for the extract that contradicts it. If there isn’t one, you have a defensible category. If there is, you have a better one waiting.

Step three is where the audit trail comes from, and the audit trail is what turns an examiner’s question into a two-minute answer. Keep a running sheet: the category, the codes underneath it, the relationship sentence, and one exemplar extract. It costs ten minutes a category and it’s close to impossible to reconstruct afterwards, which is exactly the situation students arrive in when they have intuited their themes and can’t show the working.


Selective coding: finding the core

Selective coding is the last phase. You pick the core category (the one that the rest of your categories can be related to) and write the account that holds the whole dataset together.

The test for a core category is blunt: can you tell the story of your entire study through it, without leaving major categories stranded? If two of your categories have nothing to do with your candidate core, it isn’t the core, or you have two studies.

Two honest notes about this phase.

Most postgrad research projects stop before it, and that’s often the right call. Selective coding produces a theory, and generating a defensible theory usually needs more data than a postgrad project collects, gathered until the categories stop changing. Running two phases well and saying plainly that you didn’t attempt the third is far better than claiming a theory your sample can’t support.

It overlaps heavily with axial coding. If you find yourself doing both at once, you’re in good company. Name the phases in your methodology chapter, describe what you actually did, and don’t manufacture a clean separation the work didn’t have.


Is the coding paradigm compulsory?

No, and this is worth knowing before you build a methodology chapter on it.

The six-element paradigm comes from Strauss and Corbin, and it isn’t the only version of grounded theory. Glaser, who co-founded the method with Strauss in 1967, objected to it directly. Susan Gasson’s review of rigor in grounded theory research sets out the split: Glaser saw the paradigm’s categories as “forcing” theoretical constructs onto data, and argued the resulting theories came out more descriptive than explanatory. Strauss’s position was that novice researchers need a structure to work within.

Both positions are reasonable, and you don’t have to resolve a forty-year methodological argument in your own study. What you do have to do is say which one you followed and why. A methodology chapter that describes Straussian axial coding and then presents purely emergent categories is describing one study while reporting another, and that’s the version an examiner will notice.

In practice, for most applied projects: use the paradigm as a prompt sheet rather than a template. Ask its questions of each category, use the ones that produce something, and note in your write-up that you did it that way.


How many codes should you end up with?

This is the question I get asked most about this sequence, and most guides refuse to answer it. So here is the arithmetic I give students. It’s a rule of thumb rather than a law, and it’s worth having something to check yourself against.

StageRoughly how many
After open coding100 to 200 codes
After merging the obvious duplicates50 to 80
After axial coding15 to 30 categories
Grouped into themes3 to 5

The pattern is that each pass roughly halves what came before. If a pass doesn’t reduce, it didn’t do any work: renaming forty codes leaves you with forty codes.

Three to five final themes is the range that tends to survive a defense (a viva, in the UK), ideally with two or three sub-themes each. Fewer than three usually means the analysis stopped too early and the themes are really categories. More than about seven and nobody, including you, will be able to hold the argument in their head.

“But my data genuinely has eight distinct themes.”

Sometimes it does, and then you keep eight and defend them. Far more often, two of the eight are the same theme wearing different words, and the test is whether you can state what separates them in one sentence. If you can’t, they’re one theme.

Still have questions?

What are the three types of grounded theory?

Classic (Glaserian), Straussian, and constructivist. Glaser’s version insists categories emerge with no imposed structure; Strauss and Corbin’s adds the coding paradigm and a more prescribed procedure; Charmaz’s constructivist version treats the theory as co-produced by the researcher rather than discovered in the data. Pick one, name it in your methodology chapter, and follow its rules rather than mixing all three.

Can I use axial coding without doing grounded theory?

You can use the move without claiming the method. Grouping codes into categories and stating how they relate is useful in almost any qualitative analysis, and plenty of thematic analyses do exactly that without the label. What you should not do is describe your study as grounded theory on the strength of having done one axial pass, because grounded theory makes commitments about sampling and about where your framework came from.

Is open coding the same as initial coding?

Yes, in practice. Both name the first pass through your data, before any codebook exists. You’ll see “initial coding” more often outside grounded theory and “open coding” more often inside it, so match your term to the method you’re claiming and use it consistently.

Do I have to do all three phases?

No, and saying so is better than pretending. Selective coding produces a theory, which needs enough data to reach the point where categories stop changing, and many student projects don’t have that. Running open and axial coding well and stating plainly that you didn’t attempt selective coding is a defensible position. Claiming a theory your sample can’t support isn’t.

What if my categories don’t seem to relate to each other?

That’s a finding, not a failure, and it usually means one of two things. Either you have several unrelated phenomena in one dataset, which happens when the interview guide covered too much ground, or the relationships exist at a level you haven’t reached yet and another pass will surface them. Say which you think it’s, and show the check you ran. An honest “these two categories didn’t connect” reads far better than a forced arrow.

Can’t find your answer here? Ask a Grad Coach directly – the initial chat is free.

Speak with a friendly coach →
David Phair, Grad Coach research coachEthar Al-Saraf, Grad Coach research coachKerryn Warren, Grad Coach research coachNichole Moore, Grad Coach research coachBrandon Simmons, Grad Coach research coachMatthew Courtney, Grad Coach research coachLani Malcolm, Grad Coach research coach

Pick your reference style

  • APA Social sciences

    Jansen, D. (2026). Open, Axial & Selective Coding. Grad Coach. https://gradcoach.com/axial-coding/

  • MLA Humanities

    Jansen, Derek. "Open, Axial & Selective Coding." Grad Coach, Sep. 1, 2026, https://gradcoach.com/axial-coding/. Accessed 3 Sep. 2026.

  • Chicago History and arts

    Jansen, Derek. "Open, Axial & Selective Coding." Grad Coach. September 1, 2026. https://gradcoach.com/axial-coding/.

  • Harvard Author–date, general use

    Jansen, D. (2026) 'Open, Axial & Selective Coding', Grad Coach. Available at: https://gradcoach.com/axial-coding/ (Accessed: 3 September 2026).

  • Vancouver Medicine and science

    Jansen D. Open, Axial & Selective Coding [Internet]. Grad Coach; 2026 [cited 2026 Sep 3]. Available from: https://gradcoach.com/axial-coding/

  • IEEE Engineering and tech

    D. Jansen, "Open, Axial & Selective Coding," Grad Coach, 2026. [Online]. Available: https://gradcoach.com/axial-coding/. [Accessed: Sep. 3, 2026].

Let our specialists code your data so you can focus on what really matters — analysis.

  • Coded by hand, never automated
  • Doctoral-level coding specialists
  • Matched to your methodology