Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,601 words · 1 segments analyzed
When reading the news, you may occasionally be bemused by the flip-flopping science and health stories: one week you’re told coffee is good for you, the next week that it’s bad. Red wine extends your life, then it doesn’t. Over time, new scientific studies seemingly contradict one another, leaving us unsure what evidence to believe. As a science journalist, I’ve learnt that often the actual problem is that each study is reported in isolation, rather than weighed against everything else we know. And sometimes, the risks of harm are high. For example, in September 2025, I reported on the US Food and Drug Administration’s announcement that it would add a warning label to the painkiller acetaminophen (paracetamol), claiming that taking the drug during pregnancy increased a child’s risk of autism. It made global headlines. Yet looking closer at the science, it was clear that the FDA and its parent health department had cherrypicked a few studies, and downplayed other more robust findings. This was confirmed a couple of months later when researchers published a major review that better represented the full body of knowledge. ‘Existing evidence does not clearly link maternal paracetamol use during pregnancy with autism or ADHD in offspring,’ it concluded. By then, of course, countless women had been needlessly scared about a painkiller that is widely considered one of the safest to take during pregnancy. In 2025, academics worldwide published about 7 million scholarly articles – that’s more than 19,000 each day. On the surface, that might seem like a number to celebrate, but it also poses a problem. As the volume of research balloons, it can be hard to discern what these millions of papers actually tell us. Within the firehose, there are rigorous methodologies and important findings, but also baffling contradictions, unconfirmed results and sloppy science. And some diverse forms of knowledge, such as lived experience and Indigenous insights, are rarely captured in scholarly articles and databases at all. The causes are systemic. Scientists publish and promote one paper after another because that’s how they advance in their careers. Journalists breathlessly chase the latest, flashiest studies so their headlines get clicks online. Meanwhile, hardly any of us spend time trying to make sense of what the world already knows by carefully synthesising and taking stock of existing knowledge. Iain Chalmers, a doctor who co-founded the Cochrane Collaboration, an evidence-synthesis group based in London, once called this the ‘scandalous failure of science to cumulate evidence scientifically’. This failure was a key motivation to write my book Beyond Belief: How Evidence Shows What Really Works (2026). Researching it, I discovered better ways to make sense of the world – but these methods are not as widely known as they should be. If they were, we might pause before believing news stories based on single studies, and instead recognise the real work: finding, sorting and synthesising evidence. This unglamorous labour already shapes our lives far more than any individual study or news headline will. We might also realise that, for many of the wicked problems we face, humans already possess much of the knowledge needed to solve them. All it needs is the ability to assemble it – and then act. One of the biggest events in the history of evidence synthesis occurred in a hotel in San Francisco in 1976. There, at the annual meeting of the American Educational Research Association, the statistician Gene Glass revealed an ingenious technique that showed how to integrate findings from a large volume of individual studies. This technique would become so important to science that it has its own biography. Glass’s work had emerged from his own struggles with mental health. A decade earlier, when he had finished his PhD in psychometrics and statistics, he was suffering from anxiety and neurosis. He began weekly psychotherapy sessions – which helped him hugely – and stayed in therapy for eight years. But his positive experience ran counter to the weight of academic opinion, which maintained that psychotherapy had little benefit. This stemmed in large part from the work of the psychologist Hans Eysenck, whose influential reviews of psychotherapy research concluded that it was worthless – or had a placebo effect at best. Glass was irritated. Not only did this suggest that he’d flushed away money on ineffective therapy, but he felt Eysenck’s reviewing methods were flawed. Eysenck – like other researchers at the time – tended to do what’s called ‘vote-counting’: counting up the number of studies showing a treatment has benefits, and the number that did not find it helped. The one with the most ‘votes’ wins. Intuitive though vote-counting is, it’s also misleading as a way of synthesising studies, partly because it ignores the size of the effects. One study may find that 55 people out of 100 improved after a particular treatment, and another that 85 out of 100 did. Even though the second study showed the treatment had a much stronger effect, both studies contribute one vote. Glass could see this was problematic – and was also astonished that Eysenck arbitrarily excluded hundreds of valid studies just because they were in theses or dissertations rather than published in academic journals. The audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results Glass set out to do a better job. He and his wife, a psychologist named Mary Lee Smith, systematically hunted down every scrap of research they could find that compared the effect of psychotherapy with a control group or another therapy – ending up with more than 370 studies. Then Glass and Smith worked out a way to extract and combine the measurements of psychotherapy’s effect in each study. Even though one study measured the effects of the therapy on anxiety and another its impact on blood pressure, they devised a way to convert these results into one standard measure known as ‘effect size’ and then average them across all the studies. (This is somewhat like converting various currencies into dollars, so they can all be combined.) Glass called this new statistical method a meta-analysis – an analysis of analyses, just as metadata is data about data. After toiling on this work for two years, Glass and Smith concluded that psychotherapy had a beneficial effect, and Glass presented the method at the San Francisco hotel where the education meeting was taking place. Debate continues to this day about whether psychotherapy works, when and for whom. (Eysenck called the work ‘an exercise in mega-silliness’.) But of the meta-analysis, the audience was thunderstruck. Here was a way to extract meaning from an apparent mess of different results. The meta-analysis was not the only synthesis tool to emerge around this time. The work of Glass, Smith and other researchers also led to the systematic review, a rigorously structured method for gathering and evaluating evidence. In a systematic review, researchers scour databases worldwide of published and unpublished work for all studies that address a certain question. Then they whittle down a longlist of thousands of studies to the most relevant few, assess their reliability, extract the data, and combine the results. Many systematic reviews include a meta-analysis to pool results from the included studies and estimate the overall effect of a treatment or other intervention. These tools to make sense of a body of evidence are one of the most important developments in science over the past few decades. They have the power to identify important conclusions that would never be possible from assessing each underlying study on its own. It’s the scientific equivalent of seeing the forest, not just the trees. Without fanfare, these approaches have changed modern life – and particularly so when it comes to health decisions. In medicine, this happened partly thanks to a morning stroll by Iain Chalmers, an early champion of evidence synthesis, along the Wolvercote Mill Stream in Oxford, UK, in May 1991. Chalmers had just finished a pioneering, decade-long project to synthesise all the evidence from clinical trials on treatments in pregnancy and childbirth. This work, which involved doing hundreds of systematic reviews, had shown that many standard medical practices – such as shaving women’s pubic hair during labour, and surgical episiotomies – were based on little evidence and some were harmful. The work from Chalmers and his colleagues helped change some of these practices. Now Chalmers was thinking about expanding this work. Wouldn’t it be useful, he thought, to synthesise clinical trials in every area of medicine and healthcare? Then doctors and patients would know, based on evidence, effective ways of treating diabetes, cancer, heart disease and many other conditions. Surprisingly, many medical decisions at that time were based on conventional wisdom or the unsubstantiated opinions of senior doctors rather than on evidence from research. Some years earlier, in 1979, a doctor called Archie Cochrane working in Cardiff, Wales, had challenged the medical community to do just this. ‘It is surely a great criticism of our profession that we have not organised a critical summary, by speciality or subspecialty, adapted periodically, of all relevant randomised controlled trials,’ Cochrane wrote. Chalmers was greatly influenced by Cochrane, and he was now ready to take on the task. He and his team started searching for clinical trials in online academic databases and scouring journals by hand in the library. The task was so big that they recruited volunteers to join the hunt, including elderly people’s groups and even the unlikely source of the Headington Bowls Club in Oxford. Soon, they had tens of thousands of trials. Most people who have seen a Western doctor have unknowingly benefited from systematic reviews