Methodological practices that increase risk of reporting errors in RCTs
A short essay
Introduction
Several commonly used interventions — understood here as methodological practices, design choices, and analytical decisions — systematically increase the risk of reporting errors in clinical trials. These can be organized by the type of error they most directly inflate.
Practices That Increase Type I Error (False Positives)
Early stopping for benefit is among the most impactful and well-documented sources of effect overestimation. Bassler et al.’s landmark JAMA systematic review of 91 truncated RCTs and 424 matching non-truncated RCTs found that trials stopped early for benefit reported treatment effects approximately 29% larger than non-truncated trials addressing the same question (pooled ratio of relative risks 0.71, 95% CI 0.65–0.77). This overestimation was independent of whether a formal statistical stopping rule was used. Critically, in 62% of the 63 questions studied, the pooled effects from non-truncated RCTs failed to demonstrate significant benefit — suggesting that many early-stopped trials may have reported benefits for interventions with minimal or no true effect.[1] The bias was greatest in smaller studies with fewer than 500 events.[1]
Multiplicity without adjustment remains pervasive. Khan et al. found that among cardiovascular RCTs published in high-impact journals, multiplicity issues were common yet frequently unaddressed.[2] Schulz and Grimes demonstrated that testing enough subgroups will produce false-positive results by chance alone, and that investigators may selectively report only significant effects, distorting the literature.[3] The CONSORT-Outcomes 2022 extension now explicitly requires authors to describe methods used to account for multiplicity.[4]
P-hacking and outcome switching are alarmingly common. Cro et al. found that only 25% of 89 RCTs published in six leading journals had no unexplained discrepancies between pre-specified and conducted analyses; 61% had one or more unexplained discrepancies, most commonly involving the analysis model (35%) or analysis population (31%).[5] In a broader sample, Kahan et al. found that only 12% of 100 RCTs had a publicly available pre-specified analysis approach, and of those, 92% had unexplained discrepancies.[6] Damen et al.’s analysis of 163,129 RCTs confirmed that indicators of questionable research practices were widespread, though improving over time with trial registration and CONSORT adherence.[7]
---
Practices That Increase Type II Error (False Negatives)
Inadequate sample size is the most fundamental driver. Domb and Sabetian emphasize that while the conventional framework tolerates a ≥20% Type II error rate (vs. ≤5% for Type I), additional factors frequently push the actual false-negative rate well beyond 20%.[8] Pocock and Stone illustrated this with the CIBIS example: the initial trial of bisoprolol (n=621) showed a non-significant hazard ratio of 0.80 (95% CI 0.56–1.15, p=0.22), but the subsequent CIBIS II trial (n=2647) demonstrated a highly significant 34% mortality reduction — a result that fell within the confidence interval of the underpowered first trial.[9]
Per-protocol analysis can dilute treatment effects in the opposite direction. Mostazir et al.’s meta-epidemiological study of 156 RCTs found that per-protocol estimates were on average 2% larger than ITT estimates (ROR 1.02, 95% CI 1.00–1.04), with divergence increasing with higher protocol non-adherence.[10] However, the more important concern is that ITT analysis in the setting of high crossover or non-adherence can mask true treatment effects, effectively increasing Type II error. The ART trial exemplifies this: 36% of patients did not receive the allocated treatment, and ITT and as-treated analyses yielded discordant results, leaving the trial question essentially unanswered.[11]
Ceiling and floor effects from overly broad inclusion criteria can mask treatment benefits by including participants unlikely to respond, diluting the signal among those who would benefit.[12]
---
Practices That Increase Both Type I and Type II Error
Inadequate randomization and allocation concealment inflate Type I error through systematic bias. Wang et al.’s meta-epidemiological synthesis of 41 studies demonstrated that inadequate allocation concealment leads to effect overestimation (ROR 0.92, 95% CI 0.88–0.97, moderate certainty), while lack of patient blinding substantially overestimates effects for patient-reported outcomes (ROR 0.36, 95% CI 0.28–0.48). Lack of outcome assessor blinding overestimates effects for subjective outcomes (ROR 0.69, 95% CI 0.51–0.93, high certainty).[13] Schulz et al.’s foundational JAMA study quantified these effects: odds ratios were exaggerated by 41% for inadequately concealed trials and by 17% for non-double-blinded trials.[14]
Use of surrogate endpoints introduces a distinct error profile. Cipriani et al. found that clinical studies using surrogate measures produced substantially exaggerated results compared with those using clinical outcomes (relative odds ratios of 1.28–1.48) and were twice as likely to find positive results.[15] This inflates Type I error when the surrogate does not predict the clinical outcome. The bevacizumab example is instructive: approved for metastatic breast cancer based on progression-free survival, subsequent trials found no overall survival benefit.[15] Conversely, surrogate endpoints can also increase Type III error — the BELLINI trial showed venetoclax was superior on surrogate measures (PFS, response rate) but produced worse overall survival, meaning the surrogate led to the wrong directional conclusion.[15]
---
Conclusion
The overarching lesson is that no single methodological safeguard is sufficient. These practices interact — a small trial stopped early for benefit on a surrogate endpoint, with inadequate blinding and no multiplicity adjustment, compounds multiple error sources simultaneously. Rigorous trial design requires addressing each vulnerability prospectively, and critical appraisal by the reader requires evaluating each one systematically.
References
Stopping Randomized Trials Early for Benefit and Estimation of Treatment Effects: Systematic Review and Meta-regression Analysis. Bassler D, Briel M, Montori VM, et al. JAMA. 2010;303(12):1180-7. doi:10.1001/jama.2010.310.
Prevalence of Multiplicity and Appropriate Adjustments Among Cardiovascular Randomized Clinical Trials Published in Major Medical Journals. Khan MS, Khan MS, Ansari ZN, et al. JAMA Network Open. 2020;3(4):e203082. doi:10.1001/jamanetworkopen.2020.3082.
Multiplicity in Randomised Trials II: Subgroup and Interim Analyses. Schulz KF, Grimes DA. Lancet (London, England). 2005 May 7-13;365(9471):1657-61. doi:10.1016/S0140-6736(05)66516-6.
Guidelines for Reporting Outcomes in Trial Reports: The CONSORT-Outcomes 2022 Extension. Butcher NJ, Monsour A, Mew EJ, et al. JAMA. 2022;328(22):2252-2264. doi:10.1001/jama.2022.21022.
Evidence of Unexplained Discrepancies Between Planned and Conducted Statistical Analyses: A Review of Randomised Trials. Cro S, Forbes G, Johnson NA, Kahan BC. BMC Medicine. 2020;18(1):137. doi:10.1186/s12916-020-01590-1.
Public Availability and Adherence to Prespecified Statistical Analysis Approaches Was Low in Published Randomized Trials. Kahan BC, Ahmad T, Forbes G, Cro S. Journal of Clinical Epidemiology. 2020;128:29-34. doi:10.1016/j.jclinepi.2020.07.015.
Indicators of Questionable Research Practices Were Identified in 163,129 Randomized Controlled Trials. Damen JA, Heus P, Lamberink HJ, et al. Journal of Clinical Epidemiology. 2023;154:23-32. doi:10.1016/j.jclinepi.2022.11.020.
The Blight of the Type II Error: When No Difference Does Not Mean No Difference. Domb BG, Sabetian PW. Arthroscopy : The Journal of Arthroscopic & Related Surgery : Official Publication of the Arthroscopy Association of North America and the International Arthroscopy Association. 2021;37(4):1353-1356. doi:10.1016/j.arthro.2021.01.057.
The Primary Outcome Fails — What Next?. Pocock SJ, Stone GW. The New England Journal of Medicine. 2016;375(9):861-70. doi:10.1056/NEJMra1510064.
Per-Protocol Analyses Produced Larger Treatment Effect Sizes Than Intention to Treat: A Meta-Epidemiological Study. Mostazir M, Taylor G, Henley WE, Watkins ER, Taylor RS. Journal of Clinical Epidemiology. 2021;138:12-21. doi:10.1016/j.jclinepi.2021.06.010.
Randomized Trials in Cardiac Surgery: JACC Review Topic of the Week. Gaudino M, Kappetein AP, Di Franco A, et al. Journal of the American College of Cardiology. 2020;75(13):1593-1604. doi:10.1016/j.jacc.2020.01.048.
Scientific rigour in psycho‐oncology trials: why and how to avoid common statistical errors. Bell ML, Olivier J, King MT. Psycho-Oncology. 2013;22(3):499-505. doi:10.1002/pon.3046.
Compelling Evidence From Meta-Epidemiological Studies Demonstrates Overestimation of Effects in Randomized Trials That Fail to Optimize Randomization and Blind Patients and Outcome Assessors. Wang Y, Parpia S, Couban R, et al. Journal of Clinical Epidemiology. 2024;165:111211. doi:10.1016/j.jclinepi.2023.11.001.
Empirical Evidence of Bias. Dimensions of Methodological Quality Associated With Estimates of Treatment Effects in Controlled Trials. Schulz KF, Chalmers I, Hayes RJ, Altman DG. JAMA. 1995;273(5):408-12. doi:10.1001/jama.273.5.408.
Generating Comparative Evidence on New Drugs and Devices After Approval. Cipriani A, Ioannidis JPA, Rothwell PM, et al. Lancet (London, England). 2020;395(10228):998-1010. doi:10.1016/S0140-6736(19)33177-0.
Empirical Evidence of Study Design Biases in Randomized Trials: Systematic Review of Meta-Epidemiological Studies. Page MJ, Higgins JP, Clayton G, et al. PloS One. 2016;11(7):e0159267. doi:10.1371/journal.pone.0159267.
Non-Adherence in Randomised Controlled Trials: Empirical Comparison of Treatment Policy and Efficacy Estimands Using Individual Participant Data. Mostazir MBA, Buckman JEJ, Wiles N, et al. BMC Medical Research Methodology. 2026;26(1):40. doi:10.1186/s12874-025-02760-6.
