A peer support community, independent and not for sale. Since February 2024.

GLP CircleA peer support community

clinical · how-to

Reading a trial without a PhD

The method our reading group uses on every paper — five questions, in order, that let anyone tell what a study actually showed from what people are claiming it showed.

11 min read1.4k wordsUpdated 3 July 2026Reviewed by Sunil

Why we bother

Sunil started the research reading group after watching the same thing happen four times in a year: a decent trial is published, a press release compresses it, a headline compresses the press release, and by the time it reaches anybody it has become a claim the researchers never made and would not endorse.

You do not need statistical training to catch most of that. You need a fixed order of questions and the willingness to stop reading when you have the answer. The group meets fortnightly, works through one paper, and nobody is expected to have read it all — most people read the abstract, the tables, and the limitations paragraph, which is genuinely most of the value.

A caution before the method. Being able to read a paper does not make you your own clinician, and we have watched people go badly wrong by deciding they have out-reasoned their prescriber from an armchair. What this skill actually buys you is better questions and a stronger immunity to nonsense. Both are worth having.

Question one: who was in it?

Always first, and it settles more arguments than everything else combined.

Find the inclusion and exclusion criteria — usually a paragraph in the methods, and always in the protocol if the paper is vague. Then compare that list with yourself, honestly.

SELECT (Lincoff, New England Journal of Medicine, 2023) required participants to be at least 45, to have a body mass index of 27 or above, to have established cardiovascular disease, and to not have diabetes. That single sentence disqualifies most of the internet arguments about SELECT, because most of them are about people who would not have been enrolled.

FLOW required type 2 diabetes and chronic kidney disease. SURPASS-2 required type 2 diabetes on metformin. STEP 1 excluded diabetes. These populations are not interchangeable and results do not automatically travel between them.

Also look at who was not there. Age ranges, the proportion of women, ethnicity, people with severe kidney or liver impairment, people with a history of an eating disorder, people already on other treatments. Every exclusion is a group the trial cannot speak for. Noor, who reads with us, calls this the shadow of the study and says it is usually more informative than the conclusion.

Question two: compared with what, and measuring what?

A drug is never good in the abstract; it is better or worse than something, at something, over some period.

The comparator. Placebo tells you whether a drug does anything. An active comparator tells you whether it beats the alternative, which is usually the question you actually care about. SURPASS-2 compared tirzepatide with semaglutide 1 mg — a real head-to-head — but note that it used the diabetes dose, so it does not tell you about the higher weight-management dose. Comparator details like that are where most of the meaning hides.

The endpoint. Was it something that happened to people — a heart attack, kidney failure, a death — or a number that moved? Both are legitimate, but they answer different questions. A surrogate endpoint is a bet that moving the number moves the outcome, and medicine has a long history of that bet failing.

Was it pre-specified? The primary endpoint is declared before the trial starts. Anything found afterwards by looking around in the data is hypothesis-generating, however striking it looks. If a result is described as exploratory, post-hoc, or a subgroup analysis, it is a suggestion, not a finding.

How long? STEP 1 ran 68 weeks. SURMOUNT-1 ran 72. SELECT followed people for a mean of a little over three years. Anything you want to know about year five is outside all of them.

Question three: how big, in absolute terms?

This is the question that defuses most headlines, and it is arithmetic anyone can do.

A relative risk reduction says how much the risk shrank in proportion. An absolute risk reduction says how many fewer people had the event. In SELECT, 6.5 per cent of the semaglutide group had a major cardiovascular event against 8.0 per cent on placebo. That is a hazard ratio of 0.80 — a 20 per cent relative reduction — and an absolute difference of about 1.5 percentage points over roughly three years.

Identical result, two framings, wildly different emotional weight. Press releases overwhelmingly choose the relative one. When a paper only gives you relative figures, dig the raw event counts out of the tables and work it out.

Then look at the confidence interval, which is the honest expression of how uncertain the estimate is. SELECT reported 0.72 to 0.90 — comfortably away from 1, so the direction is secure. An interval that scrapes past 1, or that spans anything from trivial to enormous, means the trial has not pinned the effect down, whatever the p-value says. Sunil’s line: the point estimate is the story the trial tells, the interval is how much of it it can back up.

Question four: what did the average hide?

Trial results are reported as means, and means conceal people.

STEP 1 (Wilding, NEJM, 2021) reported a mean weight change of about 14.9 per cent with semaglutide 2.4 mg against about 2.4 per cent with placebo over 68 weeks. SURMOUNT-1 (Jastreboff, NEJM, 2022) reported means ranging from around 15 to around 21 per cent across tirzepatide doses over 72 weeks. Those averages are correct and they are also the source of enormous unnecessary distress, because the distribution around them is wide in both trials. Real participants sat well above and well below.

Look for the responder analyses and the distribution charts, which most papers of this kind include. They tell you what proportion reached various thresholds, and they make plain that a substantial minority did not respond much. No reliable way to predict who will be where exists.

Which is why nothing in a trial is a target, and why we do not repeat those percentages as goals in this community. They are descriptions of what happened to a group of strangers under trial conditions. Ada puts it more bluntly in the plateau circle: a mean is a statement about a population and never a promise made to a person.

Question five: what happened to the people who left?

The most under-read part of any trial, and often the most honest.

Find the flow diagram. How many were randomised, how many completed, how many withdrew, and why? Discontinuation for gastrointestinal adverse events is reported across the weight-management trials and is a real feature of these drugs, and a trial with a research nurse on the end of the phone will always retain people better than ordinary care does.

Then ask how the analysis handled them. An intention-to-treat analysis keeps everyone in the group they were assigned to and is generally the conservative, trustworthy choice. Analyses of people who completed treatment as intended flatter the drug, because they quietly remove the people it did not suit. Papers often report both — the gap between them tells you something.

Finally, read the limitations paragraph, which authors write honestly far more often than they are given credit for, and skim the funding and conflict-of-interest statements. Nearly all of these trials were funded by the manufacturer. That is normal and it is not disqualifying; it is context, and it belongs alongside everything else rather than instead of it.

The claims that get stacked on top

Once you have the method, the failure modes become easy to spot. The recurring ones in this field:

  • Population laundering. A result found in people with established cardiovascular disease, restated as a claim about everyone.
  • Class laundering. A finding for one drug asserted for all GLP-1 medicines, or for compounded material that was never in any trial.
  • Surrogate creep. A number moved, therefore an outcome will follow. Sometimes true, historically often not.
  • The mechanism story. A confident causal explanation for a result the trial only observed. See the debate about whether SELECT’s benefit runs through weight or inflammation, where the trial simply cannot say.
  • The dropped denominator. "Twice as likely" with no baseline rate. Twice a very small number is still a very small number.
  • Mouse to person. A rodent finding presented as a human one — the ground the thyroid guide covers in detail.

If you want to practise, the group works through one paper a fortnight and newcomers are explicitly welcome to say they did not understand a table. That sentence gets said in most sessions, usually by someone who has been coming for a year.

The steps, in order

Your product leaflet and the person who prescribed for you override anything on this page.

  1. Read the abstract, then stopGet the shape: what drug, what population, what comparator, how long, what was measured. Do not form an opinion yet. Most misreadings happen when someone reaches a conclusion before checking who was enrolled.
  2. Find the inclusion and exclusion criteriaThey are in the methods, and in full in the protocol. Compare them with yourself honestly. If you would not have been eligible, everything that follows is context rather than a result about you.
  3. Identify the comparator and the primary endpointPlacebo or an active drug, at what dose. An event that happened to people, or a number that moved. Confirm the primary endpoint was pre-specified — anything found afterwards is a hypothesis.
  4. Work out the absolute difference yourselfTake the event rates from the main results table and subtract. Compare that with the relative figure being quoted at you. Do this every time; it takes thirty seconds and it recalibrates almost everything.
  5. Read the confidence interval, not the p-valueAsk how wide the interval is and whether it comes near no effect. A wide interval means the trial has not pinned the effect down, however impressive the point estimate looks.
  6. Look for the spread, not just the meanFind the responder analysis or the distribution chart. Ask what proportion had little or no response. Then refuse to treat any trial average as a target for yourself or anybody else.
  7. Follow the people who leftUse the flow diagram. How many withdrew, for what, and how were they analysed? Intention-to-treat is the conservative reading; completer analyses flatter the drug.
  8. Read the limitations and the funding statement lastAuthors are usually candid about what their trial cannot show. Note who paid, treat it as context rather than a verdict, and let it sit alongside the result rather than replacing it.
  9. Write one sentence you would defendSomething like: in people with X, over Y months, this drug reduced Z compared with placebo, by roughly this much in absolute terms. If you cannot write that sentence, you have not finished reading — and if a headline cannot survive it, it was not a finding.

Sources

  1. Lincoff AM, et al. Semaglutide and cardiovascular outcomes in obesity without diabetes. N Engl J Med. 2023;389(24):2221–2232. (SELECT — worked example for absolute versus relative effect)
  2. Wilding JPH, et al. Once-weekly semaglutide in adults with overweight or obesity. N Engl J Med. 2021;384(11):989–1002. (STEP 1)
  3. Jastreboff AM, et al. Tirzepatide once weekly for the treatment of obesity. N Engl J Med. 2022;387(3):205–216. (SURMOUNT-1)
  4. Frías JP, et al. Tirzepatide versus semaglutide once weekly in patients with type 2 diabetes. N Engl J Med. 2021;385(6):503–515. (SURPASS-2 — an active-comparator design)
  5. Perkovic V, et al. Effects of semaglutide on chronic kidney disease in patients with type 2 diabetes. N Engl J Med. 2024;391(2):109–121. (FLOW — a trial stopped early for benefit)
  6. Malhotra A, et al. Tirzepatide for the treatment of obstructive sleep apnea and obesity. N Engl J Med. 2024. (SURMOUNT-OSA — an example of a non-weight endpoint)

We name the trial, the journal and the year, because vague confidence is how people get hurt.

Read next

clinical

Cardiovascular outcomes: what SELECT showed

The trial that changed how clinicians talk about these drugs — who was in it, what it actually found, how big the effect was in plain terms, and the claims stacked on top of it.

10 min · reviewed by Sunil · updated 2 Jun 2026

clinical

Kidney numbers, and what FLOW showed

eGFR, creatinine and albumin explained without jargon, what the FLOW trial found in people with type 2 diabetes and kidney disease, and the everyday kidney risk that matters more.

9 min · reviewed by Mira · updated 21 May 2026

clinical

A1c, and what it is not

What glycated haemoglobin actually measures, the units confusion nobody warns you about, the things that make it lie, and why it is a poor scoreboard for a fortnight.

8 min · reviewed by Mira · updated 16 Apr 2026

clinical

Thyroid and GLP-1: what is known

The rodent finding behind the warning label, what human data has and has not shown since, what it means if you already take levothyroxine, and where the honest uncertainty sits.

9 min · reviewed by Mira · updated 29 Apr 2026

starting

What these medicines actually do

The mechanism in plain language, which drug is which, what the major trials found, and an honest accounting of where the evidence runs out.

11 min · reviewed by Sunil · updated 28 Jun 2026

clinical

Plateaus: what is happening, and what helps

Why change slows or stops, what the trial curves actually look like, what is worth checking, and why a plateau is a physiological event rather than a verdict on you.

10 min · reviewed by Ada · updated 25 Jun 2026