How to Read a Medical Paper: P-Values, Forest Plots and Trial Design for Interviews
Dr Akash Gandhi·NHS GP and Medicine Admissions ExpertPublished 30 July 2026 14 min read
Photo: National Museum of Health and Medicine (CC BY 2.0)
Photo: National Museum of Health and Medicine (CC BY 2.0)
How to read a medical paper: ask three questions in order. What was the study design? How big was the effect? And how precise is that estimate? A p-value tells you only how surprising a result would be if the treatment did nothing. The confidence interval tells you how large the real effect could plausibly be.
I am Dr Akash Gandhi, an NHS GP who has prepared applicants for medical school interviews at TheUKCATPeople since 2012. Oxbridge academic stations, and a growing number of MMI stations, hand you numbers and watch what you do with them. Candidates who memorised definitions freeze; candidates who say one honest sentence about the data stand out in ten seconds. Our MMI interview preparation guide covers the format; this page is the statistics, taught from one real trial and four diagrams.
Why an interviewer hands you a paper or a graph
They are not testing whether you can do the maths. They are testing whether you can tell what a study found from what it proves, and whether you can be uncertain out loud without going silent. Three things earn credit:
Describing before interpreting. Jumping to "so the drug works" skips the only step fully in your control.
Separating association from cause. Two things moving together is an observation. One causing the other is a claim, and needs a different design behind it.
Noticing what is missing. No control group, no sample size, no funder. Absences are easier to spot than errors and score as well.
In the mock interviews I run, the commonest failure is not a wrong number. It is a candidate who reads the data perfectly and then, in the next breath, states a conclusion it cannot support.
One thing to separate early. Describing a graph or table under time pressure is a different skill, covered in our MMI data interpretation guide. This page is research papers: study design, and what the statistics may claim. Oxbridge panels lean hard on that, so our Oxbridge medicine mentoring programme drills it, with the habits in our top 10 MMI tips on top.
Key Takeaway: You are marked on how carefully you move from what the page says to what it means, not on arithmetic.
The four parts of a paper, and where the truth hides
Every paper has the same four parts in the same order: introduction, methods, results, discussion. The table below is most of what you need. The abstract sells the paper; the methods decide whether it is worth anything.
Part
What it is for
What to check
Abstract
A summary written to be quoted
Whether its confident last sentence matches the numbers below
Introduction
Why the authors think it matters
The one question the study set out to answer
Methods
How the study was done
Design, who was included, what was measured, how many people
Results
What happened
The effect size and its interval, not the p-value alone
Discussion
What the authors claim
Whether the claim grew, plus the limitations paragraph
Two habits are worth stealing. Given two minutes, read the methods before the results, so you know what weight the numbers can bear. Then read the limitations paragraph, where authors quietly say what is wrong with their own work.
The hierarchy of evidence, and what it does not mean
The hierarchy of evidence ranks study designs by how easy it is for something other than the treatment to have produced the result. At the top, almost nothing else can explain the finding. At the bottom, almost anything can.
The hierarchy of evidence
Higher up the pyramid means fewer ways for something other than the treatment to have produced the result.
The hierarchy of evidence
Level (strongest first)
What it is
1. Systematic reviews and meta-analyses
Every trial found, pooled and weighed
2. Randomised controlled trials
Chance decides who gets the treatment
3. Cohort studies
Follow groups forward over time
4. Case-control studies
Start with the disease, look back
5. Case series and reports
One patient, or a handful
6. Expert opinion
Someone senior thinks so
The notes beside each tier are the definitions. What they cannot show is that a case report can raise a question but never answer one, or that the cohort design is how we learned smoking causes lung cancer, and how the Whitehall studies of health inequalities were done.
The nuance separating a good answer from a rehearsed one: the pyramid ranks designs, not papers. A badly run trial of 40 patients is worse evidence than a large, careful cohort study. And some questions cannot be randomised at all, because you cannot allocate teenagers to smoke.
Key Takeaway: Say why the top of the pyramid is stronger, not just that it is: randomisation removes explanations you cannot otherwise rule out.
Interview coaching
Choose your 1-1 interview coaching
Rated 5.0 from 550+ reviews. Practise with experienced interview experts: mock MMI and panel interviews, scored with feedback.
What makes a randomised controlled trial trustworthy
Four things: chance decides the allocation, there is a control group, people are kept unaware of who got what, and everybody is analysed in the group they were randomised to.
Randomisation. A computer, not a clinician, decides who gets the treatment, so the groups differ only by the treatment and by luck.
A control group. Without one you cannot know what would have happened anyway: most people with a cold recover in a week regardless.
Blinding. Single-blind means the patient does not know their group, double-blind means the researchers do not either. It matters most when the outcome involves judgement, like a pain score.
Intention to treat. Everyone is counted in the group they were randomised to, even if they stopped the drug. Analyse only those who finished and you have deleted the people it did not suit.
RECOVERY beats any definition. It randomised 6,425 patients admitted with COVID-19 to dexamethasone, a cheap steroid, or usual care, by intention to treat. It was open-label, so nobody was blinded, which matters less than usual because the outcome was death. Saying why a flaw does or does not matter beats spotting it, and the CONSORT checklist is what trials report against.
Bias and confounding, the ones worth naming
Confounding. A third factor explains both things you are looking at. Coffee drinkers once looked more likely to get lung cancer, but heavy coffee drinkers were also more likely to smoke.
Selection bias. The people studied are not the people who will get the drug. Tested on fit 40-year-olds, it tells you little about an 85-year-old on nine other medicines.
Publication bias. Trials finding nothing are less likely to be written up, so the literature is more positive than reality.
Research ethics, in three lines
UK research on people needs ethics approval first, and participants give informed consent and may leave at any time, as the NHS explains. Randomising is only defensible under equipoise: real uncertainty about which arm is better. The case every panel expects is Andrew Wakefield, a case series of twelve selectively recruited children with undisclosed payments behind it. Where that sits on the pyramid tells you the rest, alongside the four pillars.
Key Takeaway: Randomisation, a control, blinding and intention to treat are the four to check. Always name the one a trial did not have.
What a p-value actually means, and the three things it is not
A p-value is the probability of getting a result at least as extreme as the one you got, in a world where the treatment does nothing. That is the whole definition, and everything difficult about p-values comes from that last phrase.
Take an invented trial of a painkiller. Sixty patients in 100 improve on the drug, 45 in 100 on placebo, and the p-value comes out at 0.03.
What a p-value of 0.03 actually measures
Illustrative trial, invented to show what the shaded area means. It is not a real study.
The curve is every result you could get by luck if the drug did nothing at all. The shaded tails are results at least as far from the centre as the one this trial got. Their area is the p-value.
What a p-value of 0.03 actually measures
Element
Meaning
The curve
Every result you could get by chance alone if the treatment did nothing
The centre
No difference between the groups
The shaded tails
Results at least as extreme as the one observed, 2.12 standard deviations from no effect
The shaded area
p = 0.03, the probability of a result this extreme if there were no real effect
Imagine the drug is useless and you ran that trial 100 times anyway. Luck alone would throw up a gap this big about three times. That is all p = 0.03 says: how often chance could fake your result.
Now the three things it is not. Saying these marks a candidate out.
It is not the probability that the drug works. It is calculated by assuming there is no real effect, so it cannot be read backwards as the chance there is none. Doctors make this mistake too.
It is not the probability the result is a fluke. The 3% belongs to the data, not to the claim.
It is not a measure of how big the effect is. Test enough people and a difference too small to matter comes out "significant". The p-value moves with sample size; the effect does not.
And 0.05 is a convention: nothing meaningful separates 0.049 from 0.051. The Cochrane Handbook now tells its authors not to call results "statistically significant" at all, and to report the effect with its interval.
Key Takeaway: A p-value answers "could chance alone have produced this?" and nothing else. It cannot tell you whether a treatment works, or by how much.
Interview coaching
Get interview-ready, 1-1
Mock MMI and panel interviews with personalised feedback.
1-1 coaching with experienced interview tutors, never a salesperson
Mock MMI and panel interviews, scored with honest feedback
A free Ultimate Interview Q&A Guide (worth £349) with every package
A 95% confidence interval is the range of true effects your data are reasonably consistent with. It answers both questions a p-value cannot: is there an effect, and how big is it?
Use the real numbers. In RECOVERY, dexamethasone gave a rate ratio for death of 0.83, interval 0.75 to 0.93. It is a ratio, so 1 means no difference and 0.83 means the death rate was 17% lower on the steroid. Every value from 0.75 to 0.93 sits below 1, so the true benefit is between about 7% and 25% fewer deaths, which a p-value could never have told you. Narrow intervals come from large studies, wide ones from small studies or rare events.
What you see written
What it is telling you
What to say out loud
0.83 (0.75 to 0.93)
The whole interval sits below 1
A real reduction, between about 7% and 25% fewer deaths
1.19 (0.91 to 1.55)
The interval crosses 1
This study cannot tell us: a benefit and a harm both fit
Now the sentence that changes how you read everything else. When an interval crosses the line of no effect, the study has not shown the treatment does nothing. It has shown the study cannot tell. Patients in that trial needing no oxygen came out at 1.19, interval 0.91 to 1.55, which fits anything from a 9% fall in deaths to a 55% rise. "We do not know" is not "there is no difference".
Key Takeaway: An interval crossing the line of no effect means uncertainty, not absence of effect. Never say "it was not significant, so it does not work". Say "this study cannot separate a real benefit from a real harm".
Absolute versus relative risk, and how headlines mislead
Relative risk tells you how much a risk changed. Absolute risk tells you how much risk there was to begin with. A headline giving only the relative figure is unreadable, because the same percentage means something different depending on where it started.
Deaths within 28 days, patients on a ventilator
Real figures from the RECOVERY trial, ventilated patients, rounded to the nearest whole person in 100 (29.3% and 41.4%).
The same result stated three ways: 12 fewer deaths per 100 patients, a 29% relative reduction, and about 8 patients treated for one life saved.
Deaths within 28 days, patients on a ventilator
Group
How many out of the total
Usual care
died within 28 days
Dexamethasone
died within 28 days
Those are real numbers, from the sickest patients. Three ways of saying the same thing:
Absolute risk reduction. 41 in 100 died on usual care, 29 in 100 on dexamethasone: 12 fewer deaths per 100 patients treated.
Relative risk reduction. Those 12 are 12 out of the original 41: a reduction of about 29%. Same result, larger sounding number.
Number needed to treat. 100 divided by 12 is about 8. Treat eight ventilated patients and one lives who otherwise would not.
Here is why it bites. Apply that same 29% reduction to a condition killing one person in 100 rather than 41. The absolute benefit falls to under a third of a person per 100, and the number needed to treat rises above 300. The relative figure has not moved; everything that matters to a patient has.
Newspapers print the relative figure because it is bigger. One study found people eating 76g of red and processed meat daily had a 20% higher chance of bowel cancer than people eating about 21g. Roughly 1 in 15 men and 1 in 18 women in the UK are diagnosed with bowel cancer, so a fifth on top moves you nearer 1 in 12: real, but nothing like what "20% higher risk" sounds like. That arithmetic under pressure is drilled in our MMI calculation questions guide.
It is how rationing works too: NICE needs the absolute gain, because that is what converts into quality-adjusted life years. A large relative benefit on a rare event buys almost nothing, which is why some gene therapies are hard to justify.
Key Takeaway: Always ask "out of how many?" A percentage change without a baseline is not information.
How to read a forest plot in five steps
A forest plot puts one row per study, or per group of patients, on a shared axis. The square is that result, the line through it is the confidence interval, the dashed vertical line is no effect.
Dexamethasone and 28-day death rates, by how ill the patient was
Real figures from the RECOVERY trial, 28-day mortality among 6,425 patients admitted to hospital with COVID-19.
Each square is one group of patients, sized by how many people it contains. The line through it is the confidence interval. The gold diamond is the pooled result for everyone.
Dexamethasone and 28-day death rates, by how ill the patient was
Study
Estimate
95% confidence interval
On a ventilator
0.64
0.51 to 0.81
Oxygen only
0.82
0.72 to 0.94
No respiratory support
1.19
0.91 to 1.55
All patients
0.83
0.75 to 0.93
Find the line of no effect. For ratios it sits at 1, for differences at 0. Everything else is read relative to it, so find it first.
Read the axis and the arrows. Which side favours the treatment? It is written underneath, and not always the left. Getting this backwards turns a good answer into a wrong one.
Look at each square. That is the best estimate for that row, and the bigger the square, the more people behind it.
Look at the whiskers. That line is the confidence interval. If it crosses the line of no effect, that row has settled nothing on its own. An arrowhead off the edge means the study was too small.
Look at the diamond. Where there is one it pools everything above it, and its width is its confidence interval. Read it last.
Apply that to the plot above. The top two rows sit entirely left of 1, so the benefit is real and largest in the sickest patients, while the third crosses 1, so for the least ill the trial cannot tell us anything. The diamond averages all three into 0.83, and that is where the sophistication lies: reading only the diamond would have you giving this drug to everybody. When rows disagree, statisticians call it heterogeneity, and it warns you that pooling hides more than it shows, as the trial report itself says.
Key Takeaway: Line of no effect, then direction, then squares, then whiskers, then the diamond. Read the diamond last, and say what it is hiding.
Ultimate Package
Choose your Ultimate Package
Rated 5.0 from 550+ reviews. 1-1 mentoring from doctors across UCAT, personal statement and interviews.
How would you go about reading a medical paper you had never seen before?
What does a p-value of 0.05 actually mean?
What is the difference between absolute risk and relative risk?
Why is a randomised controlled trial usually stronger evidence than a cohort study?
A newspaper says a drug halves the risk of a disease. What would you want to know before believing it?
What is a confidence interval, and what does it tell you that a p-value cannot?
Less likely questions (harder, but worth knowing)
A trial finds no significant difference between two treatments. Does that mean they are equally good?
What is confounding? Give me an example from ordinary life rather than from medicine.
Why might studies that were never published change what we think we know?
When is it ethical to randomise a patient to a treatment you suspect is worse?
Model answer: "This forest plot comes from a trial of a steroid in patients admitted to hospital. Talk me through it."
The first thing I would do is find the dashed line at 1. These are ratios, so 1 means the death rate was the same in both groups, and everything to the left means fewer deaths on the steroid.
Then I would read the rows before the overall result, because they disagree. Ventilated patients are at 0.64, interval 0.51 to 0.81, entirely left of 1: a real and large reduction. Patients on oxygen alone are at 0.82, 0.72 to 0.94, also entirely below 1, so a genuine but smaller benefit.
The third row is where I would spend time. Patients needing no oxygen are at 1.19, 0.91 to 1.55, which crosses 1. That does not show the steroid is safe there, and it does not show it is harmful. It shows the trial cannot tell us: the data fit a modest benefit and a substantial harm equally well.
The diamond pools everyone at 0.83, 0.75 to 0.93. Reading only that, I would give the drug to everyone admitted, and that would be the mistake, because the group who did worst is buried inside the average. My one question would be whether those subgroups were set before the trial or found afterwards, because subgroups found afterwards are much less reliable.
Why this answer works:
It found the line of no effect first. Naming it takes five seconds and organises everything after it.
It read the intervals, not just the squares. "Entirely below 1" and "crosses 1" is the vocabulary the station wants.
It refused to over-read the uncertain row. "The trial cannot tell us" beats both "no effect" and "harmful", and few candidates say it.
It ended on a question rather than a verdict. Asking whether the subgroups were planned shows you know a finding found afterwards is worth less.
Key Takeaway: Practise on one real plot until the five steps are automatic, and you can read any plot you are handed.
The one line to take into the room
A p-value tells you whether to be surprised. A confidence interval tells you what the effect might actually be, and how confident anyone is entitled to sound. Describe what is on the page, say what the interval does and does not rule out, then say what you would want to know next.
Contact us
Want expert help with your application?
From 1-1 tutoring and personal statement editing to interview coaching and our all-in-one Ultimate Package, we support every stage. Tell us what you are working towards and we will recommend the right option.
FAQs
Frequently asked questions
What does a p-value of 0.05 mean?
It means that if the treatment did nothing at all, results at least as extreme as the one observed would still turn up by chance about 5 times in every 100 studies. It is a statement about how easily luck could fake your data. It is not the probability that the treatment works, and it says nothing about how big the effect is.
What is a confidence interval in simple terms?
It is the range of values for the true effect that your data are reasonably consistent with. A 95% interval of 0.75 to 0.93 for a death rate ratio means the real benefit is probably somewhere between 7% and 25% fewer deaths. Narrow intervals come from large studies; wide ones mean the study could not pin the answer down.
What is a forest plot?
A forest plot is a chart with one horizontal row per study or per group of patients, all drawn on a shared axis. Each square is that row’s result, the line through it is the confidence interval, and a vertical dashed line marks no effect. A diamond at the bottom, where present, is the pooled result of everything above it.
What is the difference between absolute risk and relative risk?
Absolute risk is how many people in a group are affected, such as 41 in 100. Relative risk is how much that changed, such as 29% lower. The relative figure sounds the same whether the starting risk was 41 in 100 or 1 in 100, so it is meaningless without the baseline. Always ask "out of how many?"
What is the hierarchy of evidence?
It ranks study designs by how hard it is for something other than the treatment to explain the result. From strongest to weakest: systematic reviews and meta-analyses, randomised controlled trials, cohort studies, case-control studies, case series and reports, then expert opinion. It ranks designs rather than individual papers, so a poorly run trial can be weaker evidence than a strong cohort study.
Why is randomisation important in a clinical trial?
Because it makes the two groups differ only by the treatment and by luck, including in ways nobody thought to measure. If a doctor chooses who gets the new drug, healthier patients may end up in one arm and the result becomes impossible to interpret. Randomisation is the only method that deals with confounding factors nobody has identified.
What does intention to treat mean?
Everyone is analysed in the group they were randomly allocated to, even if they stopped taking the treatment or switched. The alternative, analysing only those who completed the course, quietly removes the people the treatment did not suit and flatters the result. Intention to treat gives the more honest and usually more conservative answer.
What is confounding, with an example?
Confounding is when a third factor explains the apparent link between two others. Coffee drinkers once appeared more likely to develop lung cancer, but coffee was not the cause: heavy coffee drinkers were also more likely to smoke. Smoking was the confounder. Randomised trials handle confounding better than any other design because chance balances it out.
Does a non-significant result mean there is no difference?
No, and this is the most useful correction to make in an interview. A non-significant result usually means the study was too small or too imprecise to separate a real effect from chance. Look at the confidence interval: if it stretches from a worthwhile benefit to a worthwhile harm, the honest conclusion is that nobody knows yet.
How do I prepare for an Oxbridge medicine data or paper station?
Practise the process rather than memorising statistics. Take one real trial, find its forest plot or its main result, and talk through the design, the effect size and the confidence interval out loud until it takes ninety seconds. Then repeat with a paper you have not seen. Panels are testing how you reason under uncertainty, not what you have learned.
Ultimate Package students from our 2025/26 cycle, with their UCAT scores and offers, who trained with us for the UCAT, personal statements and interviews.
Ultimate Package
S
Sophie
Medicine, King's College London
2025 UCAT2,590 / 2,700
“Harry got my UCAT up to 2,590, working through the sections I kept dropping marks on week by week. Gemma then ran my interview practice so the MMI stations didn't catch me out, and Dr Akash mentored me the whole way through. I'm off to King's for Medicine.”
Ultimate Package
D
Daniel
Medicine, University College London
Medicine offers4 offers
“The interview prep was the part that actually moved the needle. Proper mock MMIs, not just lists of questions, and feedback that was honest about what I was getting wrong. I ended up with four offers and firmed UCL.”
Ultimate Package
A
Aisha
Dentistry, University of Birmingham
Dentistry offers4 offers
“The Ultimate Package kept me organised from UCAT through to interviews. They knew what dental schools actually ask and tightened up my personal statement. Four offers in the end, and I'm going to Birmingham.”
Ultimate Package
C
Charlotte
Veterinary Medicine, Royal Veterinary College
Vet offers4 offers
“Vet applications come down to the written SAQs as much as the interview. Dr Rebecca went through my SAQs line by line, sharpened my answers and prepped me for the panels. I came away with four offers and chose the RVC.”