Photo: National Museum of Health and Medicine (CC BY 2.0)
Photo: National Museum of Health and Medicine (CC BY 2.0)
To read a medical paper, work out what kind of study it was, how big the effect was, and how precise that estimate is: the p-value tells you how surprising the result would be by chance, and the confidence interval how large the real effect might be. Medicine interviews use data stations to test that reasoning rather than memorised statistics.
I'm Dr Akash Gandhi, an NHS GP with published research, and I've been preparing applicants for medical school interviews at TheUKCATPeople since 2012. Reading a medical paper well is mostly a matter of saying plainly what a study found and how sure anyone can be. Our MMI interview preparation guide covers these stations, and this page covers the numbers.
What do I need to understand in a medical paper?
You need to know how a paper is laid out and what a small handful of statistical terms mean. Between them, the two tables below cover most of what an interviewer is likely to hand you.
How is a medical paper structured?
Almost every paper runs in the same order: introduction, methods, results and discussion, with an abstract summarising all of it at the top. The abstract is written to be quoted, so read it as the authors' summary of their case and then check it against the numbers below.
Part
What it is for
What to check
Abstract
A summary written to be quoted
Whether its closing claim matches the numbers below it
Introduction
Why the authors think it matters
The one question the study set out to answer
Methods
How the study was done
Design, who was included, what was measured, how many people
Results
What happened
The effect size and its interval, alongside the p-value
Discussion
What the authors claim
Whether the claim grew, plus the limitations paragraph
What do the statistics actually mean?
There are only a handful of terms you need, and most are simpler than they sound. What makes them slippery is that each gets misread in the same predictable way, so I've put the common misreading alongside what the term actually means.
The term
What it means
Commonly misread as
p-value of 0.03
If the treatment did nothing, chance alone would produce a result this striking about 3 times in 100
A 3% chance that the treatment is useless
95% confidence interval, 0.75 to 0.93
The range of true effects the data are reasonably consistent with
A range the true answer is certainly inside
Absolute risk, 41 in 100
How many people in that group were actually affected
How much a treatment changed it
Relative risk, 29% lower
How much the risk changed compared with where it started
How many people that is, before you know the starting point
Number needed to treat, 8
Treat 8 people for 1 of them to benefit
That the other 7 came to harm
Key Takeaway: Read the methods before the results, because the methods are what tell you how much weight the results can carry.
Why would an interviewer show you a paper or a graph?
Because they want to see how you think when you don't already know the answer. You'll be handed an abstract, a paper or a plot, given a few minutes and an open question, and the point is to watch you reason rather than to check what you've memorised.
Key Takeaway: You're marked on how carefully you move from what the page says to what it actually means.
What is the hierarchy of evidence?
The hierarchy of evidence ranks study designs by how easy it is for something other than the treatment to have produced the result, so the designs near the top are the ones that leave fewest alternative explanations standing.
The hierarchy of evidence
Higher up the pyramid means fewer ways for something other than the treatment to have produced the result.
I use most of this pyramid in an ordinary week of general practice without thinking about it. A guideline I follow will have been built from the trials near the top, while an unusual side effect I read about might exist only as a few case reports.
Key Takeaway: Randomisation is what puts trials near the top, because it removes explanations you can't otherwise rule out.
What makes a randomised controlled trial trustworthy?
A trial earns its place near the top of that pyramid through randomisation, a control group, blinding where possible, and analysis by intention to treat.
Randomisation: a computer, rather than a doctor, decides who gets the treatment, so the two groups differ only by luck.
A control group: without one you can't know what would have happened anyway, and most colds get better whatever you do about them.
Blinding: single-blind means the patient does not know which group they are in, and double-blind means the researchers do not know either.
Intention to treat: everyone is counted in the group they were originally put into, even if they stopped taking the drug halfway through.
The RECOVERY trial shows all of this working together. It randomised 6,425 patients admitted with COVID-19 to dexamethasone, a cheap steroid, or usual care, and analysed them by intention to treat.
It was open-label, so nobody was blinded, which matters far less when the outcome is death rather than something a patient could rate subjectively. Trials report against the CONSORT checklist.
What are bias and confounding?
Confounding: a third factor explains both of the things you're looking at. Coffee drinkers once appeared more likely to get lung cancer, but they also smoked more than everyone else.
Selection bias: the people studied are not the people who will end up taking the drug. Tested only on fit 40-year-olds, it tells you very little about an 85-year-old.
Publication bias: trials that find nothing are less likely to be written up at all, so the published literature ends up looking more positive than reality.
What are the rules for research on people?
Research on people needs ethics approval first, and participants give informed consent and can leave at any point, as the NHS explains. Randomising anyone is only defensible under equipoise, meaning genuine uncertainty about which arm is better.
A p-value is the chance of getting a result at least as striking as the one you actually got, in a world where the treatment does nothing at all. That assumption, that the treatment is useless, is what the whole number is built on.
Take an invented trial of a painkiller, in which 60 patients in 100 improve on the drug and 45 in 100 on placebo, giving a p-value of 0.03. If the drug were useless and you ran that trial 100 times, luck alone would produce a gap that large about three times.
What a p-value of 0.03 actually measures
Illustrative trial, invented to show what the shaded area means. It is not a real study.
The curve is every result you could get by luck if the drug did nothing at all. The shaded tails are results at least as far from the centre as the one this trial got. Their area is the p-value.
What a p-value of 0.03 actually measures
Element
Meaning
The curve
Every result you could get by chance alone if the treatment did nothing
The centre
No difference between the groups
The shaded tails
Results at least as extreme as the one observed, 2.12 standard deviations from no effect
The shaded area
p = 0.03, the probability of a result this extreme if there were no real effect
That 3% is easy to attach to the wrong thing, and these are the readings I most often end up unpicking with students.
The probability that the drug works: the calculation starts by assuming there's no real effect, so it cannot be turned round to tell you about the drug itself.
The probability that the result is a fluke: the 3% belongs to the data, rather than to the claim built on top of the data.
A measure of how big the effect is: test enough people and a difference far too small to matter will still come out as significant.
The 0.05 cut-off is a convention rather than a law of nature, and nothing meaningful separates 0.049 from 0.051. The Cochrane Handbook now asks authors to report the effect together with its interval instead of leaning on that threshold.
Key Takeaway: A p-value answers one question only, which is whether chance alone could have produced a result like this one.
Interview coaching
Get interview-ready, 1-1
Mock MMI and panel interviews with personalised feedback.
1-1 coaching with experienced interview tutors, never a salesperson
Mock MMI and panel interviews, scored with honest feedback
A free Ultimate Interview Q&A Guide (worth £349) with every package
A 95% confidence interval is the range of true effects your data are reasonably consistent with, so it answers the question a p-value leaves open: not just whether there is an effect, but how large it might be.
In RECOVERY, dexamethasone produced a rate ratio for death of 0.83, interval 0.75 to 0.93. A ratio of 1 would mean no difference between the groups, so 0.83 represents a 17% lower death rate.
What you see written
What it is telling you
What to say out loud
0.83 (0.75 to 0.93)
The whole interval sits below 1
A real reduction, between about 7% and 25% fewer deaths
1.19 (0.91 to 1.55)
The interval crosses 1
This study cannot tell us: a benefit and a harm both fit
When an interval crosses the line of no effect, as the second row does, the honest reading is that the study has left the question open, because a real benefit and a real harm both still fit the data.
Key Takeaway: An interval crossing the line of no effect means the study can't tell us yet, which is a different finding from there being no difference.
What is the difference between absolute and relative risk?
Absolute risk is how many people are affected, such as 41 in 100, while relative risk is how much that figure changed, such as 29% lower. Both describe the same result, and the relative version always sounds larger.
Deaths within 28 days, patients on a ventilator
Real figures from the RECOVERY trial, ventilated patients, rounded to the nearest whole person in 100 (29.3% and 41.4%).
The same result stated three ways: 12 fewer deaths per 100 patients, a 29% relative reduction, and about 8 patients treated for one life saved.
Deaths within 28 days, patients on a ventilator
Group
How many out of the total
Usual care
died within 28 days
Dexamethasone
died within 28 days
Here's that same RECOVERY result written out three different ways.
Absolute risk reduction: 41 in 100 died on usual care and 29 in 100 died on dexamethasone, so 12 fewer deaths per 100 patients treated.
Relative risk reduction: those 12 are 12 out of the original 41, which works out as a reduction of about 29%.
Number needed to treat: 100 divided by 12 is about 8, so treat eight ventilated patients and one of them lives who otherwise would not.
Newspapers print the relative figure because it's the bigger number. People eating 76g of red and processed meat a day had a 20% higher chance of bowel cancer than those eating around 21g of it.
Roughly 1 in 15 men and 1 in 18 women here get bowel cancer, so adding a fifth on top moves you nearer 1 in 12, which is the conversion our MMI calculation questions guide drills.
Rationing decisions work the same way. NICE needs the absolute gain, because that's what converts into quality-adjusted life years, which is why a large relative benefit on a rare event, such as some gene therapies, is so hard to fund.
Key Takeaway: Always ask out of how many, because a percentage change without its starting number is not yet information.
How do you read a forest plot?
A forest plot puts one row per study, or per group of patients, on a shared axis, and the list below is the order I'd read one in. Working through it the same way every time stops a plot feeling overwhelming under pressure.
Dexamethasone and 28-day death rates, by how ill the patient was
Real figures from the RECOVERY trial, 28-day mortality among 6,425 patients admitted to hospital with COVID-19.
Each square is one group of patients, sized by how many people it contains. The line through it is the confidence interval. The gold diamond is the pooled result for everyone.
Dexamethasone and 28-day death rates, by how ill the patient was
Study
Estimate
95% confidence interval
On a ventilator
0.64
0.51 to 0.81
Oxygen only
0.82
0.72 to 0.94
No respiratory support
1.19
0.91 to 1.55
All patients
0.83
0.75 to 0.93
Find the line of no effect. For ratios it sits at 1 and for differences at 0, and everything else on the plot is read against it.
Read the axis and the arrows. Which side favours the treatment is written underneath, and it is not always the left.
Look at each square. That's the best estimate for that row, and the bigger the square, the more people sit behind it.
Look at the whiskers. That line is the confidence interval, and a row crossing no effect has settled nothing.
Look at the diamond last. It pools everything above it, and its width is its own confidence interval.
Read only the diamond here and you'd give this drug to everybody, when the third row shows the trial can't tell whether the least ill patients benefit at all. Statisticians call that disagreement between rows heterogeneity.
Key Takeaway: Line of no effect, then direction, then the squares, then the whiskers, and the diamond last of all.
How do I talk about a paper in my medicine interview?
Describe what's in front of you before you interpret any of it, and finish on what you'd want to know next. That order is what interviewers are listening for, and it works on any paper.
Say what the study is. The design, who was included, how many of them there were, and what was measured.
Say what it found. The effect in plain numbers, described as people rather than as percentages.
Say how confident anyone can be. The interval, and whether it crosses the line of no effect.
Say what you would want to know. A limitation, a group of patients who were left out, or who paid for the study.
You can practise on things you already meet, because a school experiment has a sample size and a control, and a health headline has a relative risk with no baseline. Here's what interviewers are testing:
Thinking aloud: narrate each step as you take it, even when you're not yet sure where you'll end up, because your reasoning is the thing being marked.
Honesty about uncertainty: "This study cannot tell us" is a correct answer, and interviewers are pleased to hear a candidate say it.
Working from the page: quote the numbers in front of you rather than a half-remembered trial, and you'll always have something solid to reason from.
Correlation and cause: two things moving together is an observation, while one causing the other is a claim that needs more behind it.
Do I need to know statistics for a medicine interview?
No. You need to know what a handful of numbers mean and how to use them in a sentence, which is a far smaller job than learning statistics. Everything here needs maths no harder than GCSE percentages.
What if I cannot do the maths in the station?
Say so, and describe what you can see instead. In the mock interviews I run, "I'd rather work that out on paper than guess" lands perfectly well, and talking your steps through out loud means a slip in the final figure does not hide sound reasoning.
Key Takeaway: Describe, then interpret, then say what you'd want to know next. You're allowed to be unsure, as long as you say what you'd do about it.
Ultimate Package
Choose your Ultimate Package
Rated 5.0 from 640+ reviews. 1-1 mentoring from doctors across UCAT, personal statement and interviews.
Example interview questions on reading medical papers
You are unlikely to be asked any of these word for word, and you do not need a prepared answer to each one. Use them to check your understanding: if you could speak for a minute on most of them, you know this topic well enough for whatever the interviewer actually asks.
Questions to get you thinking
How would you read a medical paper you had never seen before?
What does a p-value of 0.05 actually mean?
What is the difference between absolute risk and relative risk?
Why is a randomised controlled trial usually stronger evidence than a cohort study?
A newspaper says a drug halves the risk of a disease. What would you want to know?
What is a confidence interval, and what does it tell you that a p-value cannot?
Harder questions to stretch you
A trial finds no significant difference between two treatments. Does that mean they are equally good?
What is confounding? Give me an example from ordinary life rather than from medicine.
Why might studies that were never published change what we think we know?
When is it ethical to randomise a patient to a treatment you suspect is worse?
Model answer: "This forest plot comes from a trial of a steroid in patients admitted to hospital. Talk me through it."
The first thing I would do is find the dashed line at 1. These are ratios, so 1 means the death rate was the same in both groups, and anything left of that line means fewer deaths on the steroid.
Then I would read the rows before the overall result, because they do not agree. Patients on a ventilator are at 0.64, interval 0.51 to 0.81, entirely left of 1, so that is a real and large reduction. Patients on oxygen alone are at 0.82, 0.72 to 0.94, also below 1, so a smaller but genuine benefit.
The third row is the one I would slow down on. Patients needing no oxygen are at 1.19, interval 0.91 to 1.55, which crosses 1. The honest reading there is that the trial cannot tell us either way, because the data fit a modest benefit and a substantial harm equally well.
The diamond pools everyone at 0.83, 0.75 to 0.93. If I read only the diamond I would give this drug to everyone admitted, and I think that would be the wrong call, because the group who did worst is hidden inside that average.
Why this answer works:
It finds the line of no effect first: that organises everything said after it.
It reads the intervals, not just the squares: "entirely below 1" and "crosses 1" is the vocabulary the station wants.
It refuses to over-read the uncertain row: "the trial cannot tell us" beats "no effect" or "harmful".
Key Takeaway: Practise on one real plot until reading it in that order feels automatic.
The one line to take into the room
A p-value tells you whether to be surprised by a result, and a confidence interval tells you how large the real effect might be and how confident anyone is entitled to sound about it.
Key Takeaway: Describe the page, say what the interval rules out, then what you would want to know next.
Contact us
Want expert help with your application?
From 1-1 tutoring and personal statement editing to interview coaching and our all-in-one Ultimate Package, we support every stage. Tell us what you are working towards and we will recommend the right option.
FAQs
Frequently asked questions
What does a p-value of 0.05 mean?
It means that if the treatment did nothing at all, a result at least this striking would still turn up by chance about 5 times in every 100 studies. It's really a statement about how easily luck could fake the data, which is why it cannot be read as the probability that the treatment works.
What is a confidence interval in simple terms?
It's the range of values for the true effect that your data are reasonably consistent with. An interval of 0.75 to 0.93 for a death rate ratio means the real benefit is probably somewhere between about 7% and 25% fewer deaths. Narrow intervals come from large studies, and wide ones mean the study could not pin the answer down.
What is a forest plot?
A forest plot is a chart with one horizontal row per study or per group of patients, all drawn on a shared axis. Each square is that result, the line through it is the confidence interval, and a vertical dashed line marks no effect. Where a diamond appears at the bottom, it pools everything above it.
What is the difference between absolute risk and relative risk?
Absolute risk is how many people in a group are affected, such as 41 in 100, and relative risk is how much that changed, such as 29% lower. The relative figure sounds identical whether the starting risk was 41 in 100 or 1 in 100, so it means little on its own. Always ask out of how many.
What is the hierarchy of evidence?
It ranks study designs by how hard it is for something other than the treatment to explain the result. From strongest down: systematic reviews and meta-analyses, randomised controlled trials, cohort studies, case-control studies, case series and reports, then expert opinion. It ranks designs rather than papers, so a poorly run trial can be weaker than a strong cohort study.
Do I need to know statistics for a medicine interview?
No. You need to know what a handful of terms mean and how to use them in a sentence, which needs no maths beyond GCSE percentages. Understanding a p-value, a confidence interval and the difference between absolute and relative risk covers almost anything you're shown in a paper or a data station.
What if I cannot do the maths in an interview station?
Say so and describe what you can see instead. Interviewers are quite happy with "I'd want to work that out on paper rather than guess at it", so it's well worth saying. If you're given a calculation, talk through your steps out loud, so a slip in the final figure does not hide otherwise sound reasoning.
Does a non-significant result mean there is no difference?
No, and it's a useful correction to be able to make in an interview. A non-significant result usually means the study was too small or too imprecise to separate a real effect from chance. Look at the confidence interval: if it stretches from a worthwhile benefit to a worthwhile harm, nobody knows yet.
What does intention to treat mean?
Everyone's analysed in the group they were randomly allocated to, even if they stopped taking the treatment or switched to something else. The alternative, counting only those who completed the course, quietly removes the people the treatment did not suit and flatters the result. Intention to treat gives the more honest and usually more cautious answer.
How do I prepare for an Oxbridge medicine data or paper station?
Practise the process rather than memorising statistics. Take one real trial, find its main result or its forest plot, and talk through the design, the effect size and the interval out loud until it takes ninety seconds. Then repeat with a paper you've never seen. Panels are testing how you reason under uncertainty.
Ultimate Package students from our 2025/26 cycle, with their UCAT scores and offers, who trained with us for the UCAT, personal statements and interviews.
Ultimate Package
S
Sophie
Medicine, King's College London
2025 UCAT2,590 / 2,700
“Harry got my UCAT up to 2,590, working through the sections I kept dropping marks on week by week. Gemma then ran my interview practice so the MMI stations didn't catch me out, and Dr Akash mentored me the whole way through. I'm off to King's for Medicine.”
Ultimate Package
D
Daniel
Medicine, University College London
Medicine offers4 offers
“The interview prep was the part that actually moved the needle. Proper mock MMIs, not just lists of questions, and feedback that was honest about what I was getting wrong. I ended up with four offers and firmed UCL.”
Ultimate Package
A
Aisha
Dentistry, University of Birmingham
Dentistry offers4 offers
“The Ultimate Package kept me organised from UCAT through to interviews. They knew what dental schools actually ask and tightened up my personal statement. Four offers in the end, and I'm going to Birmingham.”
Ultimate Package
C
Charlotte
Veterinary Medicine, Royal Veterinary College
Vet offers4 offers
“Vet applications come down to the written SAQs as much as the interview. Dr Rebecca went through my SAQs line by line, sharpened my answers and prepped me for the panels. I came away with four offers and chose the RVC.”