Data Science Research Topics & Ideas
Here are 50 research question ideas in data science and analytics, each paired with a real dissertation or thesis on something similar.
By Derek Jansen (MBA) · Reviewed by Eunice Rautenbach (DTech)
Updated
- 01
Which methodological choices decide what a bedside monitor's data appears to say?
- 02
Can outcomes after pediatric cardiac surgery be predicted well enough to change a decision?
- 03
How far in advance can the risk of injury during childbirth be predicted?
- 04
Do social determinants of health explain who survives sepsis?
- 05
Whose outcomes does a clinical risk score get wrong, and can that be corrected?
- 06
Which capabilities does a supply chain need before AI is worth integrating?
- 07
What makes a farmer switch a brand they have bought for a decade?
- 08
Which signals predict that an employee is about to leave, and how early?
- 09
Does a real-time information system change where trucks end up parking?
- 10
Can a refrigeration unit learn what normal looks like without anyone labeling it?
- 11
Where should the threshold sit on an anomaly detector that has to run continuously?
- 12
How do you find the odd event in a smart building that logs everything?
- 13
What counts as unusual on a campus, and who decides the label?
- 14
What representation lets a model recognize something it has never seen?
- 15
Can augmented data teach a model about events that almost never happen?
- 16
What has to change before causal inference works on data nobody designed?
A sample dissertation
New solutions for real-world causal inference X Lin · University of Oxford · 2025 - 17
Can the design of an experiment carry the inference, without a model of the outcome?
- 18
Which causal questions in political economy can an instrument actually answer?
- 19
How many different quantities could a study be estimating without ever saying which?
- 20
Where does a treatment effect change, and how would you show that it did?
- 21
Where did this table come from, and can the answer be reconstructed after the fact?
- 22
What would a system have to do to find the dataset worth joining to yours?
- 23
Can a pipeline diagnose its own data quality problems before the model sees them?
- 24
Which additions to a training set improve a model, and which only enlarge it?
- 25
How much does a co-authorship network reveal about how a research field works?
- 26
Which color and shape combinations let a reader tell categories apart reliably?
- 27
What does a person need to see to keep control of a system that decides faster than they do?
- 28
What is the right way to show a robot's decisions over space and time?
- 29
How much context can a chart carry before it stops helping the decision?
- 30
Can visual analytics surface a material property relationship a model missed?
- 31
Would an operator trust an autoscaler more if it explained its decisions?
- 32
What makes an explanation of a recommendation good enough to act on?
- 33
Can you get more out of the data you have without trusting the model less?
- 34
Which interpretability method survives contact with an actual clinician?
- 35
Does an offline evaluation predict how a recommender performs in front of people?
- 36
Why does a learned representation fail on data that looks the same to a person?
- 37
How much can a pretrained model be compressed before adaptation stops working?
- 38
Which assumptions have to be built in before a model learns structure rather than correlation?
- 39
What do you see in a complex system when you look at it one agent at a time?
- 40
What does predictive analytics look like across devices that cannot pool their data?
- 41
How should uncertainty be handled when a model is chosen by its own predictions?
- 42
What breaks in statistical learning when biomedical data arrive in three shapes at once?
- 43
Can generated data stand in for measurements you were never going to get?
- 44
When is a Bayesian model worth its cost on healthcare data?
- 45
Does a battery model built in a lab survive field testing?
- 46
Can synthetic data make a recommender private without making it useless?
- 47
What does a language model add to a recommender that matrix factorization cannot?
- 48
How do you learn from clicks without also learning the bias in what got shown?
- 49
Which biases in a recommender come from the users and which from the system?
- 50
Should a hashtag be recommended from the content, or from who else used it?
Need a helping hand?
Book a free 15-minute chat and talk it through with a doctoral-qualified Grad Coach® who’s been through the topic ideation process hundreds of times.
15-minute chat. No cost. No pressure.







