The Science of Life – From Earth to the Stars

The Reproducibility Crisis: When Science Gets It Wrong

In 2011, a pharmaceutical company called Bayer reported something alarming. Its researchers had attempted to reproduce the results of 67 published studies they were considering as the basis for drug development. Only 14 of those studies, about 21%, produced results consistent with the original findings.

A year later, Glenn Begley and Lee Ellis at Amgen tried to reproduce 53 landmark cancer biology papers. Six held up. Forty-seven did not.

These weren’t fringe papers. They were published in prestigious journals, peer-reviewed, and cited hundreds of times. They had been treated as established scientific facts.

Science, it turned out, had a problem.

What Is the Reproducibility Crisis?

A researcher reviewing statistics, illustrating the reproducibility crisis in science.
A researcher reviewing graphs and statistics, illustrating the challenges of replication in scientific studies. Credit: Photo: RDNE Stock project / Pexels.

Reproducibility is the bedrock of science. If a finding is real, independent researchers using similar methods should be able to produce similar results. Reproducibility is what separates science from anecdote.

The reproducibility crisis: also called the replication crisis – refers to the widespread discovery, beginning around 2011, that a substantial fraction of published scientific results fail to replicate when other researchers attempt to verify them.

The problem has now been documented across psychology, medicine, cancer biology, neuroscience, economics, and nutrition research. The estimated rate of irreproducibility varies by field, but figures between 50% and 80% have been reported for some subfields.

This doesn’t mean science is broken. But it does mean something important about how science is done, and reported, needs to change.

The Psychology Replication Project

The clearest large-scale evidence came from the Reproducibility Project: Psychology, published in Science in 2015. A collaborative team of 270 researchers attempted to replicate 100 published psychological studies. The results were sobering:

  • Only 39% produced statistically significant results in the replication
  • Effect sizes were on average half as large as in the original studies
  • Even studies that “replicated” often showed weaker effects

This wasn’t a random sample of obscure studies. It included high-profile findings that had become staples of undergraduate textbooks: ego depletion, moral priming, power posing, and many others.

The ego depletion effect, the idea that willpower is a limited resource that gets depleted with use, failed to replicate in a large multinational study in 2016. Power posing, Amy Cuddy’s famous claim that holding a “power pose” for two minutes raises testosterone and lowers cortisol, generated major replication failures and scientific controversy. Priming effects, in which subtle cues unconsciously influence behavior, replicated only weakly and inconsistently.

Free Newsletter

Why Does Science Fail to Replicate?

The causes are multiple and intertwined.

1. Publication Bias

Journals overwhelmingly publish positive results: studies that find an effect. Studies that find nothing rarely get published. This creates a systematic distortion: the published literature overrepresents surprising, impressive-sounding findings, while negative results that would balance the picture go unreported.

If 20 labs run the same experiment and 1 finds a significant result by chance, that lab’s paper gets published. The other 19 findings never see print. The effect appears real in the literature even though it isn’t.

This is called the file drawer problem: null results are filed away, not reported.

2. P-Hacking and Researcher Degrees of Freedom

A p-value below 0.05 is the traditional threshold for a “significant” result. But there are many decisions a researcher makes in the course of a study, how many subjects to recruit, which variables to measure, which statistical tests to run, whether to exclude outliers, when to stop collecting data, and each decision creates an opportunity to nudge the p-value downward.

When researchers (consciously or not) make these decisions in ways that favor a significant result, they are engaging in p-hacking or exploiting researcher degrees of freedom. The result looks like a rigorous finding but is actually the product of selective analysis.

A simulation by Andrew Gelman and colleagues showed that with enough researcher degrees of freedom, you can find a statistically significant result for almost any hypothesis: including ones that are obviously false, like the claim that listening to “When I’m Sixty-Four” makes you younger.

3. Small Sample Sizes

Many studies, particularly in psychology and neuroscience, are conducted on small samples: sometimes fewer than 30 subjects. Small samples produce noisy estimates with wide uncertainty ranges. They are prone to detecting effects that aren’t real (false positives) and missing effects that are real (false negatives). They also produce inflated effect size estimates when they do find something, because only large (and therefore often spurious) effects cross the significance threshold in small samples.

Statistical power, the probability of detecting a real effect if one exists, is often woefully inadequate in published research. A survey of neuroscience studies found median power of around 20%. That means these studies would miss a real effect 80% of the time.

4. HARKing: Hypothesizing After Results Are Known

Science is supposed to work like this: form a hypothesis, design an experiment to test it, collect data, analyze results. But researchers sometimes reverse this order: collect data, explore patterns, find something interesting, then write the paper as if that finding was the pre-specified hypothesis.

This practice, HARKING[1] (Hypothesizing After Results are Known), inflates false positives because it takes advantage of the fact that any dataset contains random patterns. If you search enough variables, you will always find something that looks significant.

A dataset with 20 variables contains 190 possible pairwise correlations. By chance alone, about 9 of them will be statistically significant at p < 0.05.

5. The Garden of Forking Paths

Statistician Andrew Gelman coined this phrase to describe how researchers navigate countless decision points in analysis. Even without deliberate p-hacking, the accumulation of reasonable-sounding choices, each one defensible in isolation, can produce results that look significant but wouldn’t replicate.

The forking paths are invisible in the final paper, which presents one clean analysis as if it were the only possible analysis.

6. Fraud and Misconduct

A small fraction of replication failures are due to outright fraud. High-profile cases include Diederik Stapel in social psychology, who fabricated data for dozens of studies, and Hwang Woo-suk in stem cell biology, who falsified breakthrough results. Retractions have increased dramatically over the past two decades: though this may partly reflect better detection of misconduct, not an increase in its frequency.

Fraud is real but probably accounts for a small fraction of the reproducibility crisis. The larger problem is the systemic incentive structures that make honest p-hacking and publication bias so pervasive.

The Incentive Problem

Many researchers understand the statistical issues perfectly well. The problem is that the incentive system in academia pushes against solving them.

Scientists are hired, promoted, and funded based on their publication record: specifically, on publishing novel, significant findings in prestigious journals. This creates pressure to:

  • Run small, cheap studies that produce quick results
  • Report whatever came up as the central finding
  • Never publish null results
  • Never replicate other people’s work (replications don’t count for much on a CV)

The system rewards novelty over rigor. It punishes the kind of slow, careful, boring verification that makes science reliable.

How Science Is Fighting Back

The crisis has prompted genuine reform efforts across multiple fronts.

Pre-registration

Pre-registration means publicly documenting your hypothesis, methods, and analysis plan before you collect any data. Pre-registered studies show dramatically lower rates of significant results: but those results are far more likely to replicate. The Center for Open Science runs the Open Science Framework, which hosts thousands of pre-registrations[2].

Some journals now offer Registered Reports: a format where peer review happens before data collection. If the study is well-designed, the journal commits to publishing the results regardless of outcome. This breaks the link between significance and publication.

Open Data and Open Materials

Requiring researchers to share their raw data and analysis code allows others to check the analysis, attempt reproductions, and build on the work. Reproducibility failures are often caught because independent researchers can examine the data directly.

Larger, Multi-Site Studies

The replication crisis has driven a trend toward larger, collaborative studies with pre-registered protocols and samples drawn from multiple sites. The Many Labs projects in psychology have shown that some findings replicate robustly across sites while others collapse.

Laboratory glassware used in scientific research
Laboratory research; larger, pre-registered, multi-site studies are helping science replicate more reliably. Credit: jarmoluk / Pixabay.

Changing the P-Value Threshold

A 2017 proposal signed by 72 leading statisticians suggested raising the significance threshold from p < 0.05 to p < 0.005 for novel findings. This would substantially reduce the rate of false positives, though critics note it would also reduce the detection of real but subtle effects.

Others have argued for abandoning significance thresholds altogether in favor of continuous measures of evidence like Bayes factors or effect size estimates with confidence intervals.

Adversarial Collaboration

Some researchers have taken to pre-registering deliberate “adversarial collaborations”: studies designed jointly by scientists who hold opposing views, specifically to provide a fair test of their disagreement. This reduces the ability of either side to design the study to favor their hypothesis.

What the Crisis Tells Us About Science

The reproducibility crisis is alarming. But it is also, in a strange way, reassuring.

The reason we know about the crisis is because scientists tested whether their results held up. That testing capacity, the willingness to challenge even celebrated findings, is a feature of science, not a failure. Other ways of knowing (religion, ideology, tradition) don’t have built-in reproducibility checks. This is a core principle of what makes a theory scientific: the ability to be tested and potentially falsified.

The crisis reveals that published findings are not facts. They are claims with varying degrees of evidential support. A single study, even in a prestigious journal, is evidence, not proof. This was always true. The crisis has made it unmistakably obvious.

It has also forced the field to grapple with the gap between the idealized image of science and its actual practice. The idealized version is that scientists form hypotheses, test them rigorously, and report honestly whatever they find. The actual version involves career pressures, unconscious bias, flexible analysis, and journals that reward drama over rigor.

Fixing the incentive structure is harder than fixing the statistics. But the field is trying. Concepts like falsifiability, the idea that a scientific claim must be testable and open to refutation – remain essential guides for distinguishing robust science from findings that merely appear convincing.

Which Fields Are Most and Least Affected?

The crisis is not evenly distributed.

Most affected: Social psychology, nutrition research, preclinical cancer biology, and much of clinical medicine, especially studies of small effects with large numbers of confounders.

Less affected: Physics and chemistry, where experiments can often be replicated precisely, effects are large and unambiguous, and measurement errors are well-characterized. The discovery of gravitational waves, for example, was replicated independently by multiple detectors before announcement.

In between: Ecology, evolutionary biology, and economics, where replication is possible but complicated by real-world variability.

The Lesson for Reading Science News

The reproducibility crisis argues for a few simple rules when consuming science news:

  1. Wait for replication. A single study, no matter how impressive, is not a reason to change your behavior. Wait to see if it holds up.
  2. Sample size matters. A study of 30 college students is much weaker evidence than a study of 3,000 people across multiple sites.
  3. Pre-registered studies are more trustworthy. Check whether the study was pre-registered.
  4. Effect sizes matter as much as significance. A statistically significant effect can be too small to be practically meaningful.
  5. Meta-analyses are more reliable than individual studies. A systematic review of many studies is stronger evidence than any single one.

Science at its best is a slow, collective process of error correction. The reproducibility crisis has accelerated that process. The science that emerges from this reckoning will be more reliable, even if it produces fewer flashy headlines.

Sources

[1] Kerr, N.L. (1998). HARKing: Hypothesizing After the Results are Known. Personality and Social Psychology Review, 2(3), 196–217. Link

[2] Center for Open Science. Open Science Framework. Link

What is the reproducibility crisis in science?

The reproducibility crisis, also known as the replication crisis, is the widespread finding since around 2011 that many published scientific results cannot be reproduced by independent researchers, undermining the reliability of those findings.

How common is the reproducibility problem in scientific studies?

Estimates vary by field, but landmark studies found that only about 21% of preclinical drug studies and 11% of major cancer biology papers could be successfully replicated.

Which fields are affected by the reproducibility crisis?

The crisis has been documented across psychology, medicine, cancer biology, neuroscience, economics, and nutrition research.

What caused the reproducibility crisis?

Key causes include small sample sizes, selective reporting, p-hacking, publication bias favoring positive results, and insufficient replication efforts.

How does the reproducibility crisis impact scientific progress?

It wastes resources on false leads, erodes public trust in science, and can delay medical advances by basing drug development on unreliable findings.

Further reading: Replication crisis on Wikipedia