Skip to content
Grad Coach
Qualitative Data Coding

Qualitative Data Coding 101

How to code your qualitative data – the smart way

By Eunice Rautenbach (DTech) · Reviewed by Derek Jansen (MBA)

Updated

Pick your reference style

  • APA Social sciences

    Rautenbach, E. (2020). Qualitative Data Coding 101. Grad Coach. https://gradcoach.com/qualitative-data-coding-101/

  • MLA Humanities

    Rautenbach, Eunice. "Qualitative Data Coding 101." Grad Coach, Dec. 24, 2020, https://gradcoach.com/qualitative-data-coding-101/. Accessed 8 Sep. 2026.

  • Chicago History and arts

    Rautenbach, Eunice. "Qualitative Data Coding 101." Grad Coach. December 24, 2020. https://gradcoach.com/qualitative-data-coding-101/.

  • Harvard Author–date, general use

    Rautenbach, E. (2020) 'Qualitative Data Coding 101', Grad Coach. Available at: https://gradcoach.com/qualitative-data-coding-101/ (Accessed: 8 September 2026).

  • Vancouver Medicine and science

    Rautenbach E. Qualitative Data Coding 101 [Internet]. Grad Coach; 2020 [cited 2026 Sep 8]. Available from: https://gradcoach.com/qualitative-data-coding-101/

  • IEEE Engineering and tech

    E. Rautenbach, "Qualitative Data Coding 101," Grad Coach, 2020. [Online]. Available: https://gradcoach.com/qualitative-data-coding-101/. [Accessed: Sep. 8, 2026].

The 30-second summary

Qualitative data coding is the first step in making sense of non-numerical data like interviews, documents or videos.

  • A code is simply a label that captures the meaning of a piece of data.
  • Coding can be deductive (using pre-set codes), inductive (letting codes emerge from the data), or a mix of both.
  • The process usually involves two stages: initial coding for broad labels, then line-by-line coding for detail.
  • Good coding lays the foundation for spotting themes, building categories and creating a transparent analysis.

By coding systematically, you’ll make your qualitative data manageable and set yourself up for strong analysis.

Navigate

As we’ve discussed previously, qualitative research makes use of non-numerical data – for example, words, phrases or even images and video. To analyze this kind of data, the first dragon you’ll need to slay is qualitative data coding (or just “coding” if you want to sound cool). But what exactly is coding and how do you do it?


What is qualitative data coding?

Let’s start by understanding what a code is. At the simplest level, a code is a label that describes the content of a piece of text. For example, in the sentence:

“My supervisor only replied to my emails after about three weeks.”

You could use “supervisor communication” as a code. That code simply says what the sentence is about.

So, building onto this, qualitative data coding is the process of creating and assigning codes to categorize data extracts. You’ll then use these codes later down the road to derive themes and patterns for your qualitative analysis (for example, thematic analysis).

Coding and analysis can take place simultaneously, but coding does not necessarily involve identifying themes (depending on which textbook you’re reading, of course). It generally refers to the process of labeling and grouping similar types of data, which makes generating themes and analyzing the data more manageable.

Makes sense? Great. But why should you bother with coding at all? Why not just look for themes from the outset? Well, coding is a way of making sure your data is valid. It forces your analysis to be undertaken systematically and leaves a trail another researcher can review (in the world of research, we call this transparency). Good coding is the foundation of high-quality analysis.

Definition of qualitative coding

What are the different types of coding?

Now that we’ve got a plain-language definition of coding on the table, the next step is to understand what overarching types of coding exist, or coding approaches. Let’s start with the two main approaches, inductive and deductive.

With deductive coding, you, as the researcher, begin with a set of pre-established codes and apply them to your data set (for example, a set of interview transcripts). Inductive coding on the other hand, works in reverse, as you create the set of codes based on the data itself, so the codes emerge from the data. Let’s take a closer look at both.

Deductive coding 101

With deductive coding, we make use of pre-established codes, which are developed before you interact with the present data. This usually involves drawing up a set of codes based on a research question or previous research. You could also use a code set from the codebook of a previous study.

Deductive coding allows you to approach your analysis with a very tightly focused lens and quickly identify relevant data. Of course, the downside is that you could miss out on some very valuable insights as a result of this tight, predetermined focus.

Deductive coding of data

Inductive coding 101

But what about inductive coding? As we touched on earlier, this type of coding involves jumping right into the data and then developing the codes based on what you find within the data.

For example, if you were to analyze a set of open-ended interviews, you wouldn’t necessarily know which direction the conversation would flow. If a conversation begins with a discussion of cats, it may go on to include other animals too, and so you’d add these codes as you progress with your analysis. Simply put, with inductive coding, you “go with the flow” of the data.

Inductive coding is great when you’re researching something that isn’t yet well understood because the coding derived from the data helps you explore the subject. Therefore, this type of coding is usually used when researchers want to investigate new ideas or concepts, or when they want to create new theories.

Inductive coding definition

A little bit of both… hybrid coding approaches

If you’ve got a set of codes you’ve derived from a research topic, literature review or a previous study (i.e. a deductive approach), but you still don’t have a rich enough set to capture the depth of your qualitative data, you can combine deductive and inductive methods – this is called a hybrid coding approach.

To adopt a hybrid approach, you’ll begin your analysis with a set of a priori codes (deductive) and then add new codes (inductive) as you work your way through the data. Essentially, the hybrid coding approach provides the best of both worlds, which is why it’s pretty common to see this in research.


How to code qualitative data

Now that we’ve looked at the main approaches to coding, the next question you’re probably asking is “how do I actually do it?”. Let’s take a look at the coding process, step by step.

Both inductive and deductive methods of coding typically occur in two stages: initial coding and line by line coding.

In the initial coding stage, the objective is to get a general overview of the data by reading through and understanding it. If you’re using an inductive approach, this is also where you’ll develop an initial set of codes. Then, in the second stage (line by line coding), you’ll delve deeper into the data and (re)organize it according to (potentially new) codes.

Let’s take a look at these two stages of coding in more detail.

Step 1 – Initial coding

The first step of the coding process is to identify the essence of the text and code it accordingly. While there are various qualitative analysis software packages available, you can just as easily code textual data using Microsoft Word’s “comments” feature.

Let’s take a look at a practical example. Assume you had the following interview data from a study on postgraduate supervision:

How did you find the first few months of your project?

Honestly, I felt quite lost. My supervisor was on sabbatical, so I worked from the proposal document and hoped I was heading in the right direction.

And once she was back?

Better, but the feedback arrived in big batches. I’d wait three weeks and then get four pages of comments all at once.

In the initial stage, you are after broad, rough codes rather than precise ones:

Interview extractInitial code
I felt quite lost … hoped I was heading in the right directionUncertainty
My supervisor was on sabbaticalSupervisor availability
I worked from the proposal documentWorking from documents
the feedback arrived in big batches … four pages of comments all at onceFeedback

These are just a starting point that you will develop and refine later. Broad codes are fine here; you will build onto them in the second stage.

For fully worked examples of individual techniques, with the coded extracts set out side by side, see our guide to qualitative coding examples.

How to decide which codes to use

But how exactly do you decide what codes to use when there are many ways to read and interpret any given sentence? Well, there are a few different approaches you can adopt. The main approaches to initial coding include:

  • In vivo coding
  • Process coding
  • Descriptive coding
  • Structural coding
  • Values coding

One term you’ll meet elsewhere is open coding. It is not a sixth technique on this list, it is another name for the broad first pass you’re doing right now, so you can read “open coding” and “initial coding” as the same stage. In grounded theory it is also the first of three named phases, followed by axial and selective coding.

Let’s take a look at each of these:

In vivo coding

When you use in vivo coding, you make use of a participants’ own words, rather than your interpretation of the data. In other words, you use direct quotes from participants as your codes. By doing this, you’ll avoid trying to infer meaning, rather staying as close to the original phrases and words as possible.

In vivo coding is particularly useful when your data are derived from participants who speak different languages or come from different cultures. In these cases, it’s often difficult to accurately infer meaning due to linguistic or cultural differences.

Process coding

Next up, there’s process coding, which makes use of action-based codes. Action-based codes are codes that indicate a movement or procedure. These actions are often indicated by gerunds (words ending in “-ing”) – for example, running, jumping or singing.

Process coding is useful as it allows you to code parts of data that aren’t necessarily spoken, but that are still imperative to understanding the meaning of the texts.

An example here would be if a participant were to say something like, “I have no idea where she is”. A sentence like this can be interpreted in many different ways depending on the context and movements of the participant. The participant could shrug their shoulders, which would indicate that they genuinely don’t know where she is; however, they could also wink, showing that they do know exactly where she is.

Simply put, process coding is useful as it allows you to, in a concise manner, identify the main occurrences in a set of data and provide a dynamic account of events. For example, you may have action codes such as “describing a setback”, “explaining a method choice”, or “negotiating with a supervisor”.

Descriptive coding

Descriptive coding aims to summarize extracts by using a single word or noun that encapsulates the general idea of the data. These words will typically describe the data in a highly condensed manner, which allows the researcher to quickly refer to the content.

Descriptive coding is very useful when dealing with data that appear in forms other than traditional text – i.e. video clips, sound recordings or images. For example, a descriptive code could be “food” when coding a video clip that involves a group of people discussing what they ate throughout the day, or “cooking” when coding an image showing the steps of a recipe.

Structural coding

Structural coding involves labeling and describing specific structural attributes of the data, rather than the actual topics expressed in it. Most often the labels come straight from your research questions or interview guide, so that each segment is tagged with the question it answers; structural coding covers that version in full. More loosely, it can mean coding according to answers to the questions of “who”, “what”, “where” and “how”. Either way, it is useful when you want to access segments of data quickly, and it helps tremendously when you’re dealing with large data sets.

For example, if you were coding a collection of theses or dissertations (which would be quite a large data set), structural coding could be useful as you could code according to different sections within each of these documents – i.e. according to the standard dissertation structure. What-centric labels such as “hypothesis”, “literature review”, and “methodology” would help you to efficiently refer to sections and navigate without having to work through sections of data all over again.

Structural coding is also useful for data from open-ended surveys. This data may initially be difficult to code as they lack the set structure of other forms of data (such as an interview with a strict set of questions to be answered). In this case, it would be useful to code sections of data that answer certain questions such as “who?”, “what?”, “where?” and “how?”.

Values coding

Finally, values coding involves coding that relates to the participant’s worldviews. Typically, this type of coding focuses on excerpts that reflect the values, attitudes, and beliefs of the participants. Values coding is therefore very useful for research exploring cultural values, and the experiences and actions that flow from them.

To recap, the aim of initial coding is to understand and familiarize yourself with your data, to develop an initial code set (if you’re taking an inductive approach) and to take the first shot at coding your data. The coding approaches above allow you to arrange your data so that it’s easier to navigate during the next stage, line by line coding (we’ll get to this soon).

While these approaches can all be used individually, it’s important to remember that it’s possible, and potentially beneficial, to combine them. For example, when conducting initial coding with interviews, you could begin by using structural coding to indicate who speaks when. Then, as a next step, you could apply descriptive coding so that you can navigate to, and between, conversation topics easily.

Step 2 – Line by line coding

Once you’ve got an overall idea of your data, are comfortable navigating it and have applied some initial codes, you can move on to line by line coding. Line by line coding is pretty much exactly what it sounds like – reviewing your data, line by line, digging deeper and assigning additional codes to each line.

With line-by-line coding, the objective is to pay close attention to your data to add detail to your codes. For example, if you have a discussion of beverages and you previously just coded this as “beverages”, you could now go deeper and code more specifically, such as “coffee”, “tea”, and “orange juice”. The aim here is to scratch below the surface. This is the time to get detailed and specific so as to capture as much richness from the data as possible.

In the line-by-line coding process, it’s useful to code everything that could conceivably matter, even if you don’t think you’re going to use it (you may just end up needing it). The obvious exception is material that is plainly not data, like the scheduling chat at the start of an interview. As you go through this process, your coding will become more thorough and detailed, and you’ll have a much better understanding of your data as a result of this, which will be incredibly valuable in the analysis phase.


Moving from coding to analysis

Once you’ve completed your initial coding and line by line coding, the next step is to start your analysis. Of course, the coding process itself will get you in “analysis mode” and you’ll probably already have some insights and ideas as a result of it, so you should always keep notes of your thoughts as you work through the coding.

There are many different types of analyses (we discuss some of the most popular ones here) and the type of analysis you adopt will depend heavily on your research aims, objectives and questions. Therefore, we’re not going to go down that rabbit hole here, but we’ll cover the important first steps that build the bridge from qualitative data coding to qualitative analysis. Harvard Library’s guide to qualitative analysis is a good map of the territory beyond them.

When starting to think about your analysis, it’s useful to ask yourself the following questions to get the wheels turning:

  • What actions are shown in the data?
  • What are the aims of these interactions and excerpts? What are the participants potentially trying to achieve?
  • How do participants interpret what is happening, and how do they speak about it? What does their language reveal?
  • What are the assumptions made by the participants?
  • What are the participants doing? What is going on?
  • Why do I want to learn about this? What am I trying to find out?
  • Why did I include this particular excerpt? What does it represent and how?

The type of qualitative analysis you adopt will depend heavily on your research aims, objectives and research questions.

As with the initial coding and line by line coding, your qualitative analysis can follow certain steps. The first two steps are code categorization and theme identification.

Code categorization

Categorization is simply the process of reviewing everything you’ve coded and then creating code categories that can be used to guide your future analysis. In other words, it’s about creating categories for your code set. Let’s take a look at a practical example.

From this categorization, you can move onto the next step, which is to identify the themes in your data.

Theme identification

From the coding and categorization processes, you’ll naturally start noticing themes, so the logical next step is to identify and clearly articulate them. You take what you have learned from the coding and categorization and group it together.

This is the part where you draw meaning from your data and start to produce a narrative. The nature of that narrative depends on your research aims and objectives, your research questions (sounds familiar?) and the qualitative data analysis method you’ve chosen, so keep those front of mind as you scan for themes.

Themes help you develop a narrative in your qualitative analysis

Tips & tricks for quality coding

Before we wrap up, let’s quickly look at some general advice, tips and suggestions to ensure your qualitative data coding is top-notch.

  • Before you begin coding, plan out the steps you will take and the coding approach and technique(s) you will follow to avoid inconsistencies.
  • When adopting deductive coding, it’s useful to use a codebook from the start of the coding process. This will keep your work organized and will ensure that you don’t forget any of your codes. George Washington University’s guide to the coding process sets out what a codebook entry should contain.
  • Whether you’re adopting an inductive or deductive approach, keep track of the meanings of your codes and remember to revisit these as you go along.
  • Avoid using synonyms for codes that are similar, if not the same. This will allow you to have a more uniform and accurate coded dataset and will also help you to not get overwhelmed by your data.
  • While coding, make sure that you remind yourself of your aims and coding method. This will help you to avoid directional drift, which happens when coding is not kept consistent.
  • If you are working in a team, make sure that everyone has been trained and understands how codes need to be assigned.

Still have questions?

How do I code qualitative data in Excel?

Put one data extract per row, then use a column for your code and a second column for any note about why you applied it. Filtering and sorting on the code column then does most of the work of grouping, which is really all that dedicated software gives you for a small project. Keep your code definitions on a separate sheet so the list stays consistent as it grows.

What software do I need for qualitative coding?

None, for a typical dissertation. Word’s comments feature or a spreadsheet handles 10 to 20 interviews perfectly well, and the learning curve on a dedicated package rarely pays for itself at that scale. If your dataset is much larger or several people are coding it, our guide to qualitative analysis software compares the options.

What is the difference between codes, categories and themes?

They are three levels of the same process. A code labels one extract (“supervisor availability”). A category groups related codes together (“supervision”). A theme is the interpretive claim you build from those categories (“students describe supervision as unpredictable rather than absent”). Codes are descriptive, themes are analytical, and categories are the step between them. Thematic coding works through the same ladder with a full example.

Do I have to code every line?

In the second pass, yes, code more than you think you need. It is far easier to ignore a code later than to go back through forty transcripts hunting for something you decided was irrelevant at the time. The exception is genuinely irrelevant material, like scheduling chat at the start of an interview.

Can’t find your answer here? Ask a Grad Coach directly – the initial chat is free.

Speak with a friendly coach →
David Phair, Grad Coach research coachEthar Al-Saraf, Grad Coach research coachKerryn Warren, Grad Coach research coachNichole Moore, Grad Coach research coachBrandon Simmons, Grad Coach research coachMatthew Courtney, Grad Coach research coachLani Malcolm, Grad Coach research coach

Pick your reference style

  • APA Social sciences

    Rautenbach, E. (2020). Qualitative Data Coding 101. Grad Coach. https://gradcoach.com/qualitative-data-coding-101/

  • MLA Humanities

    Rautenbach, Eunice. "Qualitative Data Coding 101." Grad Coach, Dec. 24, 2020, https://gradcoach.com/qualitative-data-coding-101/. Accessed 8 Sep. 2026.

  • Chicago History and arts

    Rautenbach, Eunice. "Qualitative Data Coding 101." Grad Coach. December 24, 2020. https://gradcoach.com/qualitative-data-coding-101/.

  • Harvard Author–date, general use

    Rautenbach, E. (2020) 'Qualitative Data Coding 101', Grad Coach. Available at: https://gradcoach.com/qualitative-data-coding-101/ (Accessed: 8 September 2026).

  • Vancouver Medicine and science

    Rautenbach E. Qualitative Data Coding 101 [Internet]. Grad Coach; 2020 [cited 2026 Sep 8]. Available from: https://gradcoach.com/qualitative-data-coding-101/

  • IEEE Engineering and tech

    E. Rautenbach, "Qualitative Data Coding 101," Grad Coach, 2020. [Online]. Available: https://gradcoach.com/qualitative-data-coding-101/. [Accessed: Sep. 8, 2026].

Let our specialists code your data so you can focus on what really matters — analysis.

  • Coded by hand, never automated
  • Doctoral-level coding specialists
  • Matched to your methodology