The threat of serious outcome reporting bias in randomized controlled trials on acute ischemic stroke to evidence synthesis: a meta-epidemiological study
Review Article

The threat of serious outcome reporting bias in randomized controlled trials on acute ischemic stroke to evidence synthesis: a meta-epidemiological study

Na Zhang1,2,3#, Youlin Long2,3#, Xinyao Wang2,3, Xinyi Wang3, Qiong Guo3,4, Zhengchi Li3, Liang Du2,3

1General Practice Medical Center, West China Hospital, Sichuan University, Chengdu, China; 2Chinese Evidence-Based Medicine Centre, West China Hospital, Sichuan University, Chengdu, China; 3Innovation Institute for Integration of Medicine and Engineering, West China Hospital, Sichuan University, Chengdu, China; 4West China Medical Publishers, West China Hospital, Sichuan University, Chengdu, China

Contributions: (I) Conception and design: N Zhang, Y Long, L Du, Z Li; (II) Administrative support: L Du, Z Li; (III) Provision of study materials or patients: None; (IV) Collection and assembly of data: N Zhang, Y Long, Xinyao Wang, Xinyi Wang, Q Guo; (V) Data analysis and interpretation: N Zhang, Y Long, Xinyao Wang; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work as co-first authors.

Correspondence to: Liang Du, MD. Chinese Evidence-Based Medicine Centre, West China Hospital, Sichuan University, No. 37 Guoxuexiang Road, Chengdu 610041, China; Innovation Institute for Integration of Medicine and Engineering, West China Hospital, Sichuan University, No. 37 Guoxuexiang Road, Chengdu 610041, China. Email: duliang@scu.edu.cn; Zhengchi Li, MD. Innovation Institute for Integration of Medicine and Engineering, West China Hospital, Sichuan University, No. 37 Guoxuexiang Road, Chengdu 610041, China. Email: 603172078@qq.com.

Background: Stroke is the second-leading cause of death and the third-leading cause of disability, with acute ischemic stroke (AIS) being the most serious subtype. Systematic reviews of randomized controlled trials (RCTs) for AIS play a crucial role in formulating clinical guidelines and health policies. However, potential outcome reporting bias (ORB) in RCTs may skew the analytical results of systematic reviews and ultimately lead to suboptimal medical decisions. This study was conducted to investigate the prevalence and possible influencing factors of ORB in RCTs included in systematic reviews of AIS and to correct ORB at the level of systematic reviews.

Methods: A systematic literature search was conducted across three databases to retrieve subject headings and text terms related to AIS, RCTs, and systematic reviews, with the aim of identifying AIS-related systematic reviews published in 2022. ORB in trials was employed to assess the risk of ORB in RCTs, and multivariate logistic regression was used to identify the possible ORB-related factors, including registration, country, quality of journal, funding, sample size, and type of control. The correcting for ORB model was used to correct ORB evidence synthesis results.

Results: A total of 33 systematic reviews and 287 nonduplicate RCTs were included in this study. ORB was suspected in 138 (48.08%) of these RCTs. Statistically significant outcomes were more likely to be reported than were nonsignificant ones [relative risk (RR) =3.18; 95% confidence interval (CI): 2.77–3.64]. The potential factors associated with ORB were unregistered status [odds ratio (OR) =4.87; 95% CI: 1.93–12.28] and sample sizes smaller than 100 (OR =2.57; 95% CI: 1.30–5.10). The corrected results indicated that 31.58% of the therapeutic effects were overestimated due to reversal and that 16.67% of adverse reactions were underestimated due to reversal. Among outcomes without reversal, 56.52% of the effect sizes and 60.87% of the P values exceeded the clinically acceptable range.

Conclusions: The presence of ORB within the field of AIS poses a serious threat to the reliability of evidence synthesized in systematic reviews. In the future, healthcare practitioners and decision-makers should adopt a critical perspective when applying seemingly favorable results in clinical practice.

Keywords: Acute ischemic stroke (AIS); selection bias; meta-epidemiology; systematic review


Submitted Apr 24, 2025. Accepted for publication Aug 28, 2025. Published online Dec 19, 2025.

doi: 10.21037/cdt-2025-212


Highlight box

Key findings

• Selective outcome reporting bias (ORB) in randomized controlled trials (RCTs) may lead to biased estimates of evidence synthesis in the field of acute ischemic stroke (AIS).

What is known and what is new?

• ORB is recognized as a risk factor in the reliability assessment of RCTs.

• This study investigated the prevalence and possible factors of ORB in RCTs included in systematic reviews of AIS, and aimed to correct ORB at the level of systematic reviews.

What is the implication, and what should change now?

• The findings indicate that the authors of systematic reviews should prioritize the significance of ORB, improve their ability to detect RCTs at risk of ORB, and strengthen the evaluation, analysis, correction, and interpretation of such RCTs.

• We recommend that clinicians and policymakers pay close attention to the issue of ORB and cautiously apply positive results.


Introduction

According to the 2019 Global Burden of Disease Study, approximately 101 million people worldwide experience stroke, with stroke being the second-leading cause of death and the third-leading cause of disability (1). Acute ischemic stroke (AIS), accounting for approximately 87% of all strokes, arises from cerebral artery occlusion and is characterized by its sudden onset, rapid progression, high disability and mortality rates, and poor prognosis (2).

Randomized controlled trials (RCTs) are regarded as the gold standard in clinical practice for evaluating the efficacy of therapeutic interventions (3,4). However, to increase the likelihood of publication and dissemination or to align with stakeholders’ expectations, some researchers may selectively report only a subset of analyses based on the study outcomes. This specific form of bias, resulting from selectively reporting the set of study outcomes, is known as “outcome reporting bias (ORB)”, which is defined as the selection for publication of a subset of the original recorded outcome variables on the basis of the results (5,6). Such practices primarily include the following: the selective reporting of certain study outcomes (e.g., emphasizing only statistically significant or positive findings), the selective reporting of a specific outcome measured at multiple time points (e.g., presenting data from only one selected time point), and the incomplete reporting of outcomes (e.g., providing P values without corresponding effect values or confidence intervals) (7,8). Numerous studies have identified significant issues with selective outcome reporting in RCTs (9-13). Unfortunately, empirical evidence indicates that ORB in RCTs may lead to distorted estimates of intervention effects (6,7,14,15). More critically, interventions that are actually ineffective or even harmful but are obscured by ORB may pose substantial risks to patients—particularly in the case of diseases such as AIS, which severely threaten health.

Systematic reviews of RCTs play a vital role in formulating clinical guidelines and health policies, with significant implications for alleviating the public health burden of AIS (16). However, inadequate consideration of ORB in RCTs may skew the findings of systematic reviews and ultimately lead to suboptimal clinical decisions (17-20). Therefore, accurately assessing whether—and to what extent—the results of systematic reviews are influenced by ORB is essential for the sound and effective use of evidence derived from these reviews (21).

Several guidelines provide explicit instructions for authors to minimize the risk of ORB. The requirement to document discrepancies in outcome reporting and the underlying reasons for this are listed in the published “Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA)” guidelines and Cochrane Handbook (22). Specifically, the risk-of-bias tool recommended by the Cochrane Handbook evaluates issues related to ORB in RCTs, employing a three-tier classification system (high, low, and unclear) to assess the risk of such bias (23). However, the application of this tool is highly dependent on the availability of supporting materials (e.g., trial registrations and original study protocols), which may limit its utility in cases where comprehensive pre-study documentation is insufficient (22). To address this limitation, researchers from the Outcome Reporting Bias In Trials (ORBIT) project developed a classification system to assist systematic reviewers in detecting ORB. This method relies solely on information within the published trial reports (e.g., inconsistencies between sections such as the abstract, methods, and results), expert clinical judgment, and an assessment of the outcomes that researchers usually report across a set of trials in a review (6).

To further determine the extent of ORB’s impact, it is necessary to obtain missing data in RCTs. Undoubtedly, the optimal approach is to acquire unreported data directly from the RCT’s authors. However, due to factors such as the confidentiality of original data, difficulties in contacting authors, and restrictions imposed by data-sharing policies, most authors of systematic reviews often face challenges in obtaining these data (15). As a result, a statistical method for correcting ORB at the level of systematic review has been proposed. Fundamentally, ORB is a missing data issue in the context of evidence synthesis. To address this, van Aert et al. developed a correction method known as the correcting for ORB (CORB) model (20). Through an adjustment of the ORB, the pooled effect estimates in systematic reviews may shift toward null, leading to a more cautious interpretation of the statistical significance of interventions. This has important implications for enhancing the real-world applicability of medical interventions and ultimately improving patient health outcomes. At present, the scant research on this subject has focused on ORB correction methods, but no such research has been reported in the field of AIS. Therefore, we performed this study to determine the prevalence of ORB in RCTs included in systematic reviews on AIS and to identify the related influencing factors. We present this article in accordance with the Guidelines for Reporting Meta-epidemiological Methodology Research (available at https://cdt.amegroups.com/article/view/10.21037/cdt-2025-212/rc).


Methods

Study search strategy and selection criteria

To be eligible for inclusion in our analysis, studies had to be systematic reviews (including at least one binary analysis outcome, with availability of both the sample size and number of events) published in 2022 that evaluated treatments for AIS. All types of interventions and comparators were considered, and no language restrictions were applied. Network meta-analyses, umbrella meta-analyses, dose-response meta-analyses, duplicate publications, conference abstracts, protocols, and comments were excluded.

The Cochrane Database of Systematic Reviews (CDSR), MEDLINE (via Ovid), and Embase (via Ovid) were systematically searched on June 25th, 2023. The search strategy was developed based on relevant literature published in the CDSR (24-26) and subsequently refined by our expert team to align with the objectives of our study (Table S1).

Duplicate articles were removed via EndNote (Clarivate, London, UK). Four reviewers (N.Z., Y.L., Xinyao Wang, and Xinyi Wang) independently assessed the eligibility of the articles based on titles and abstracts, with those deemed potentially eligible proceeding to full-text screening. Disagreements were resolved through discussion and consensus, with input being elicited from a fifth senior researcher (L.D.) when necessary.

Data collection

Two reviewers (N.Z. and Y.L) independently extracted data, with the other two researchers (Xinyi Wang and Q.G.) conducting verification. Disagreements were resolved through discussion.

Data collection was conducted at both the systematic review and RCT levels. For the systematic review level, data collection included the following: (I) basic information such as the first author’s name and affiliated country, year of publication, published journal, intervention and comparison groups, number of RCTs included in the quantitative analysis, registration information, funding sources, and other relevant characteristics; (II) outcome information, including outcome type, number of RCTs contributing to outcomes, P value, pooled effect size with corresponding confidence intervals, and an assessment of whether the outcome of the systematic review was partially reported. Partial reporting was defined as a situation in which the number of RCTs contributing data to the meta-analysis of a specific outcome was smaller than the number of RCTs expected to report on that outcome based on the inclusion criteria of the systematic review. This definition is based on the assumption that outcomes included in systematic reviews generally reflect the most relevant endpoint questions (27). For example, if a systematic review included 10 eligible RCTs but only 8 of these trials reported data on a primary outcome, this would indicate that the remaining 2 trials might have engaged in selective outcome reporting by omitting these results. In such cases, it was also necessary to collect the sample size and the number of events for the respective outcomes. Data extraction at the RCT level mainly involved the aforementioned basic information.

Assessment of ORB in RCTs

After removal of duplicate RCTs, the likelihood of ORB was assessed via the ORBIT approach (17). According to ORBIT terminology for benefit and harm outcomes, RCTs that either did not report or only partially reported a review outcome were categorized as high risk, low risk, or no risk of ORB. This classification system includes four queries: (I) whether there is evidence that the trialists measured and analyzed (or compared) the outcomes; (II) whether the trialists measured the outcomes but did not necessarily analyze (or compare) them; (III) whether it is unclear if the outcomes were measured; and (IV) whether it is clear that outcomes were not measured at all.

Test for the likelihood of ORB being present in systematic reviews

The outcomes of all RCTs reported in the systematic reviews were categorized into five groups: (I) data provided and reported as statistically significant (P≤0.05); (II) data provided and reported as statistically nonsignificant (P>0.05); (III) data not provided but reported to be statistically significant (P≤0.05); (IV) data not provided but reported to be statistically nonsignificant (P>0.05); and (V) neither data nor statistical significance reported with these outcomes being assumed to be statistically nonsignificant (18). The distribution of these categories was compared via contingency tables, and the relative risk (RR) of reporting statistically significant versus nonsignificant results was calculated. An RR >1 with statistical significance indicated that the unreported outcomes were more likely to be negative.

ORB is essentially a missing data issue. Compared to data missing completely at random (MCAR) or missing at random (MAR), addressing data that are not MAR (NMAR) is more challenging, as any method used requires assumptions that cannot be tested with the observed data (28). Therefore, SPSS software (IBM Corp., Armonk, NY, USA) was employed to perform Little’s MCAR test to verify whether the data conformed to the NMAR assumption. If the null hypothesis was rejected (P<0.05), ORB was suspected, and the results would require corresponding adjustment.

Data analysis

Multivariate logistic regression was employed to identify potential factors of ORB, including trial registration, national economic level of the first author’s affiliated country, journal quality, funding, sample size, and type of control (29-31).

The CORB model developed by van Aert et al. was employed to correct the ORB at the systematic review level. Bland-Altman analysis was used to calculate the degree of change in the pooled effect size before and after correction. The maximum acceptable absolute change for the odds ratio (OR) was set to 0.2, and the P value was set to 0.01 (32). Reversal of OR or P value, such as the OR shifting from <1 to >1 or the P value shifting from statistically significant (P≤0.05) to nonsignificant (P>0.05), indicated that there was a statistically significant change in the OR or P value. To avoid the classification of minor or borderline changes as significant, it was required that at least one of the two P values (pre- or post-correction) fell outside the range of 0.04–0.06. In other words, changes in statistical significance within this range (e.g., from P=0.04 to P=0.06 or vice versa) were not considered meaningful changes in statistical significance (33).


Results

Description of included studies

From the 2,142 records identified through the literature search, 33 eligible systematic reviews were included in this study (Figure 1). Among these, 14 (42.42%) reviews were published in journals ranked in the first quartile (Q1) of the Journal Citation Reports (JCR). These 33 systematic reviews comprised 319 RCTs, with 287 remaining after removal of duplicates (Table 1). The majority of these RCTs were published after 2004 (276/287, 96.17%) and were funded by nonindustrial organizations (249/287, 86.76%). Over half of the RCTs were unregistered (177/287, 61.67%), conducted by authors affiliated with institutions in developing countries (186/287, 64.81%), published in a JCR non-Q1 journal (185/287, 64.46%), or had a sample size exceeding 100 participants (188/287, 65.51%). Notably, nearly half of the RCTs (138/287, 48.08%) were suspected to be at risk of ORB.

Figure 1 Study flowchart. RCT, randomized controlled trial.

Table 1

Characteristics of the included studies

Characteristic Systematic review (n=33) RCT (n=287)
Intervention
   Drug 17 (51.52) 84 (29.27)
   Medical apparatus 13 (39.39) 172 (59.93)
   Others 3 (9.09) 31 (10.80)
Control
   Drug 19 (57.58) 169 (58.89)
   Medical apparatus 5 (15.15) 70 (24.39)
   Placebo/blank 5 (15.15) 29 (10.10)
   Others 4 (12.12) 19 (6.62)
Year of publication
   Before 2004 11 (3.83)
   After 2004 276 (96.17)
National economic level of the first author
   Developed country 19 (57.58) 101 (35.19)
   Developing country 14 (42.42) 186 (64.81)
Registration
   Yes 12 (36.36) 110 (38.33)
    Prospective 12 (100.00) 94 (85.45)
    Retrospective 0 (0.00) 16 (14.55)
   No 21 (63.64) 177 (61.67)
JCR category of journal
   Q1 14 (42.42) 102 (35.54)
   Non-Q1 19 (57.58) 185 (64.46)
Sample size
   <100 0 (0.00) 99 (34.49)
   ≥100 33 (100.00) 188 (65.51)
Industry funding
   Yes 0 (0.00) 38 (13.24)
   No 33 (100.00) 249 (86.76)
Selective reporting
   Yes 138 (48.08)
   No 149 (51.92)

Data are presented as number (%). JCR, Journal Citation Reports [2022]; Q1, top 25% of journals ranked by the JCRs; non-Q1, journals categorized as Q2, Q3, or Q4 by JCRs, as well as journals not indexed by JCRs. RCT, randomized controlled trial.

Assessment of ORB in RCTs

Among the 138 RCTs suspected to be at risk of ORB, 27.54% (38/138) were assessed as high risk (Table S2). For beneficial outcomes, 38.89% (14/36) were classified as “G” (not mentioned, but clinical judgment suggests they were likely to have been measured and analyzed but not reported due to nonsignificant results), and 25.00% (9/36) were classified as “H” (not mentioned, but clinical judgment suggests they were likely never measured). For harm outcomes, 45.10% (46/102) were classified as “U” (no harms mentioned or reported), and 29.41% (30/102) were classified as “T2” (no description of specific harms).

Likelihood of ORB in systematic reviews

The 287 RCTs included a total of 699 outcomes reported in systematic reviews. Among the reported outcomes, 239 were statistically significant, while 143 were nonsignificant. Among the unreported outcomes, 2 were statistically significant and 315 were nonsignificant. Statistically significant outcomes were more likely to be reported than were nonsignificant ones [RR =3.18; 95% confidence interval (CI): 2.77–3.64], suggesting that the data were not MAR. The Little’s test results failed to find evidence that this data was not MCAR (P=0.09).

The influencing factors of ORB

Multivariate logistic regression analysis revealed a significant association between trial nonregistration and ORB (OR =4.87; 95% CI: 1.93–12.28; P=0.001), as well as between a sample size of fewer than 100 participants and ORB (OR =2.57; 95% CI: 1.30–5.10; P=0.007) (Figure 2). In contrast, country, journal quality, funding, and type of control were not significantly associated with ORB (P>0.05).

Figure 2 The potential factors that could influence ORB. CI, confidence interval; ORB, outcome reporting bias; Q1, first quartile.

Correction of ORB

The 33 systematic reviews included a total of 97 binary outcomes, of which 31 were partially reported. As shown in Figure 3, among these 31 outcomes, 19 (61.29%) were benefit outcomes, and 12 (38.71%) were harm outcomes. Correction results indicated that 31.58% (6/19) of the benefit outcomes were reversed, including mortality, modified Rankin Scale (mRS) scores, and recanalization rates, all of which changed from statistically significant in favor of the intervention group to statistically nonsignificant. Additionally, 16.67% (2/12) of the harm outcomes were reversed, involving adverse effects that shifted from statistically nonsignificant in favor of the intervention group to statistically significant.

Figure 3 The results of correction for ORB. CI, confidence interval; mRs, modified Rankin scale; ORB, outcome reporting bias; sICH, symptomatic intracranial hemorrhage.

In addition, among the 23 outcomes that were without reversed statistical significance (Table 2), the effect size estimate of 13 outcomes exceeded the clinically acceptable range (13/23, 56.52%), including 6 benefit outcomes (6/23, 26.09%) and 7 harm outcomes (7/23, 30.43%). Moreover, the P values of 14 outcomes (14/23, 60.87%) fell outside the clinically acceptable range, comprising 8 benefit outcomes (8/23, 34.78%) and 6 harm outcomes (6/23, 26.09%).

Table 2

The changes in OR and P values after correction

Statistical indicator Benefit outcome Harm outcome
Reversal Nonreversal Reversal Nonreversal
Within the clinically acceptable range Beyond the clinically acceptable range Within the clinically acceptable range Beyond the clinically acceptable range
Change in OR 6 (19.35) 7 (22.58) 6 (19.35) 2 (6.45) 3 (9.68) 7 (22.58)
Change in P value 6 (19.35) 5 (16.13) 8 (25.81) 2 (6.45) 4 (12.90) 6 (19.35)

Data are presented as number (%). OR, odds ratio.


Discussion

Our study found that nearly half of the RCTs in the field of AIS were at risk of ORB, with 27.54% assessed as being at high risk. Lack of trial registration and a sample size of fewer than 100 participants were identified as potential contributing factors to ORB. Correction results indicated that 31.58% of the treatment effects reported in systematic reviews might have been overestimated due to ORB in RCTs, while 16.67% of adverse effects might have been underestimated. Furthermore, among the outcomes that were not reversed after correction, 56.52% of the effect sizes and 60.87% of the P values exceeded the clinically acceptable range.

At the level of RCTs, our study found that unregistered trials and those with small sample sizes were more susceptible to ORB, which is aligned with findings from previous studies (29,34-36). Unregistered RCTs often lack an accessible trial protocol, making it difficult to assess, via an analysis of the original plan, whether deviations occurred due to researchers altering the study design after failing to achieve the expected results. Additionally, small-sample studies are more vulnerable to publication bias, as researchers may be more likely to selectively withhold negative outcomes due to pressure from stakeholder pressures or other publication motivations (37,38). However, Jia et al. reported that the association between sample size and outcome modification or selective reporting was not statistically significant (29). Furthermore, trial registration has been shown to improve the quality of RCTs (39) and enhance the reliability of reported treatment effects (34,35). Therefore, for small-sample studies, registration can facilitate the risk assessment of ORB and the analysis of reasons for early termination and nonpublication, thereby contributing to the evaluation of the reliability of systematic review. However, our findings revealed that unregistered RCTs accounted for 61.67% of the total, with small-sample studies constituting 25.42% of these trials. The combined risk arising from ORB and small-sample effects substantially undermines the credibility of systematic reviews based on unregistered, small-sample studies. It has been shown that meta-analyses based on small-sample studies may exhibit up to a 20% deviation in effect size compared to subsequent large-sample trials (36). Therefore, it is recommended that meta-analysis results derived from small-sample studies—especially unregistered ones—be applied with caution in the formulation of clinical decisions and policies for AIS.

We found no significant association between commercial funding and ORB, which aligns with the findings from several recent studies (29,40). This may be partly attributed to the scientific community’s growing emphasis on transparency and disclosure of funding sources in recent years, as well as the improved rate of conflicts of interest (COI) disclosures by researchers (41). Nevertheless, the completeness and adequacy of these COI disclosures remain limited, and the accuracy of disclosures and the effectiveness of managing high-risk COIs are still unclear. Consequently, relying solely on author-reported disclosure information to assess the association between COIs and ORB is unreliable. Additionally, our study also found no significant difference between journal quality and ORB. Evidence from multiple medical disciplines has shown that even high-impact journals face challenges related to selective outcome reporting (11,42-44), indicating that ORB is a widespread issue. For journal editors and reviewers, priority is often given to checking registration, with little verification of the consistency between the registered protocol and the published report, which can lead to a failure in detecting ORB. Therefore, we recommend that RCT authors minimize instances of measured but unreported outcomes and that journals enforce stricter review processes to address this issue. Authors should also clearly describe changes to prespecified outcomes, measurement methods, and analytical approaches and indicate if exploratory analyses were conducted in order to avoid misleading future researchers. Additionally, authors of systematic reviews should not rely solely on journal quality to assess ORB risk and should interpret exploratory findings cautiously.

Furthermore, given the severity of ORB in RCTs within the field of AIS, our study found that 31.58% of treatment effects were reversed after correction at the level of the systematic review. This suggests that the interventions may be ineffective or no more efficacious than existing treatments. Reversed safety outcomes would indicate the potential for unknown risks, precluding the clinical application of these treatments. We further found that 16.67% of adverse events showed reversal, highlighting the need to carefully balance the benefits and risks, including the severity and frequency of adverse events when such interventions are being evaluated.

For diseases, such as AIS, which carry high rates of mortality and disability, inaccurate estimation of treatment effects may expose patients to risks that have not been fully recognized. When the unknown risks of an intervention outweigh its potential benefits, patients could be subjected to severe harm and even death. Thus, when systematic reviews include studies at risk of ORB, researchers should adopt more rigorous and cautious approaches to the evidence design, evaluation, and correction of the evidence. However, although the systematic reviews included in this study highlight serious ORB in RCTs, none conducted in-depth analysis or correction for it. Page et al. also showed that authors of systematic reviews seldom pay attention to the impact of ORB on evidence synthesis (45).

However, even advanced methods such as Little’s test are often insufficient to detect missing data resulting from ORB in RCTs within systematic reviews. This limitation highlights the difficulty in adequately capturing the mechanisms behind the missing data caused by ORB in systematic reviews. Therefore, we recommend that authors of systematic reviews prioritize addressing ORB by enhancing their ability to identify RCTs at risk of ORB and by strengthening the processes of evaluation, analysis, correction, and interpretation. These practices can minimize the introduction of ORB during the systematic review process and mitigate its impact on conclusion reliability. Specific measures to achieve these ends include the following: (I) in the process of retrieval, researchers should judiciously consider whether to incorporate “outcome” as a search term. Inclusion may reduce sensitivity, as negative outcomes may be downplayed or even omitted in abstracts and titles. Conversely, omitting “outcome” may enhance the sensitivity at the expense of specificity, consequently increasing the volume of irrelevant records and the subsequent screening burden. Therefore, search strategies should carefully balance the tradeoff between sensitivity and specificity, while accounting for domain-specific ORB risk (45,46). (II) In the process of primary screening, whether the outcomes of interest have been reported should not be used as an exclusion criterion for full-text screening (17). (III) In the process of data analysis, measures such as contacting original authors, applying statistical methods to impute missing data, correcting for risk of bias, and conducting sensitivity analyses should be adopted to quantify the impact of missing data on the final conclusions. (IV) In the process of quality assessment, it is important to evaluate the likelihood of exploratory analyses and their consistency with prespecified outcomes. In addition, for RCTs, a high risk of ORB, it is advised that such studies be excluded from the primary analysis and that sensitivity analysis be used to test decision robustness. (V) In the process of data analysis, attention should be given not only to within-study ORB, but also to its potential harmful role in selective reporting across studies, particularly in the form of publication bias (47). ORB can lead to a “false positive” when negative results are omitted, whereas publication bias may suppress nonsignificant studies. Together, these two forms of bias can systematically lead to an overestimation of meta-analytic results. Therefore, consideration should be given to excluding RCTs with a high risk of ORB; tools such as funnel plots, the Egger test, and the Begg test should be used to assess publication bias; and corrections such as the trim-and-fill method should be applied during statistical analysis (48). Subsequently, sensitivity analysis should be conducted to ensure the robustness of the conclusions and to reduce bias-driven effect size inflation. (VI) For unstable results or those at risk of ORB, cautious and critical thinking should be applied during interpretation. Additionally, in the development of guidelines that involve weighing disadvantages and advantages, ORB should be fully considered as a criterion for downgrading both the evidence quality and recommendation strength.

Certain limitations to this study should be acknowledged. First, we focused solely on the issue of ORB during the synthesis of systematic reviews without accounting for ORB risks within the systematic reviews themselves, potentially underestimating ORB prevalence in AIS research. Second, current ORB identification relies heavily on indirect inferences from the published literature, trial registration, and protocol information and involves a degree of subjectivity. Reliable data and tools that can directly demonstrate the presence of ORB are still lacking. Due to insufficient transparency between study design and reporting, as well as inadequate oversight and management of trial registration platforms, there is a lack of robust data directly demonstrating ORB. Moving forward, collaborative efforts among research institutions, academic journals, and regulatory bodies are essential to standardizing trial registration and reporting practices, enhancing transparency, and ultimately mitigating ORB at its source. Third, in addition to the subjectivity introduce by individual items, the ORBIT tool faces the following challenges: (I) it does not classify outcomes that have been prespecified and fully reported (these outcomes should be categorized as low or no risk); (II) it does not cover certain types of outcome ORB, such as the reporting of non-prespecified outcomes and changes in data measurement or analysis plans (including “data dredging”); and (III) it requires a high degree of inferential judgment, especially clinical judgment, and it focuses more on discrepancies within study reports rather than on the consistency between study protocols and reports (22). Fourth, correction of ORB involves various tools and techniques, whose varying performance and reliability may yield divergent outcomes. Fifth, only systematic reviews published in 2022 were included in the analysis, yielding a relatively small number of RCTs, which might have led to the underestimation of ORB in the field of AIS. Sixth, only ORB among RCTs that were not incorporated into meta-analyses were assessed, which might have overlooked potential ORB in the RCTs included in meta-analyses. Nevertheless, we believe that if substantial biases are already evident in the evaluation of only the partially reported RCTs within the systematic reviews, it is reasonable to infer that the true extent of the issue may be even more severe. Finally, as this study focused exclusively on a binary outcome, it does not provide empirical evidence regarding the presence of ORB in continuous outcomes or validation of correction methods for such cases. Addressing ORB in continuous outcomes will be an important focus of our future research.


Conclusions

This study suggests that nearly half of the RCTs in the field of AIS may be at risk of ORB, and the correction results indicate that this bias significantly compromises the reliability of evidence synthesized in systematic reviews. Given the severity of ORB observed in the field of AIS, similar risks may reasonably be inferred in other medical fields as well. We recommend that clinicians and policymakers pay close attention to the issue of ORB and interpret positive findings with caution.


Acknowledgments

We appreciate the assistance provided by Ruixian Tang in data analysis and Ruizi Fu in polishing our paper.


Footnote

Reporting Checklist: The authors have completed the Guidelines for Reporting Meta-epidemiological Methodology Research. Available at https://cdt.amegroups.com/article/view/10.21037/cdt-2025-212/rc

Peer Review File: Available at https://cdt.amegroups.com/article/view/10.21037/cdt-2025-212/prf

Funding: This work was supported by the National Natural Science Foundation of China (Nos. 72074161 and 82574861).

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://cdt.amegroups.com/article/view/10.21037/cdt-2025-212/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Global, regional, and national burden of stroke and its risk factors, 1990-2019: a systematic analysis for the Global Burden of Disease Study 2019. Lancet Neurol 2021;20:795-820. [Crossref] [PubMed]
  2. Robbins BT, Howington GT, Swafford K, et al. Advancements in the management of acute ischemic stroke: A narrative review. J Am Coll Emerg Physicians Open 2023;4:e12896. [Crossref] [PubMed]
  3. Ginsburg A, Smith MS. Do randomized controlled trials meet the “gold standard”. American Enterprise Institute. 2016. Available online: https://www.carnegiefoundation.org/wp-content/uploads/2016/03/Do-randomized-controlled-trials-meet-the-gold-standard.pdf
  4. Mikulik R, Neto G, Sedani R, et al. Differences in acute ischemic stroke treatment: A cross-sectional study from international Registry of Stroke Care Quality (RES-Q). Int J Stroke 2025; Epub ahead of print. [Crossref] [PubMed]
  5. Meerpohl JJ, Schell LK, Bassler D, et al. Evidence-informed recommendations to reduce dissemination bias in clinical research: conclusions from the OPEN (Overcome failure to Publish nEgative fiNdings) project based on an international consensus meeting. BMJ Open 2015;5:e006666. [Crossref] [PubMed]
  6. Shah K, Egan G, Huan LN, et al. Outcome reporting bias in Cochrane systematic reviews: a cross-sectional analysis. BMJ Open 2020;10:e032497. [Crossref] [PubMed]
  7. Kirkham JJ, Dwan KM, Altman DG, et al. The impact of outcome reporting bias in randomised controlled trials on a cohort of systematic reviews. BMJ 2010;340:c365. [Crossref] [PubMed]
  8. Thomas ET, Heneghan C. Catalogue of bias: selective outcome reporting bias. BMJ Evid Based Med 2022;27:370-2. [Crossref] [PubMed]
  9. Wang X, Long Y, Zhang N, et al. Impact of selective reporting bias on stroke trials: potential compromise in evidence synthesis - A cross-sectional study. BMC Med Res Methodol 2024;24:255. [Crossref] [PubMed]
  10. Lin Y, Yang Y, Li Z, et al. Cohort profile: the West-China hospital alliance longitudinal epidemiology wellness (WHALE) study. Eur J Epidemiol 2025;40:1143-59. [Crossref] [PubMed]
  11. Rankin J, Ross A, Baker J, et al. Selective outcome reporting in obesity clinical trials: a cross-sectional review. Clin Obes 2017;7:245-54. [Crossref] [PubMed]
  12. Won J, Kim S, Bae I, et al. Trial registration as a safeguard against outcome reporting bias and spin? A case study of randomized controlled trials of acupuncture. PLoS One 2019;14:e0223305. [Crossref] [PubMed]
  13. Taji Heravi A, Gryaznov D, Busse JW, et al. Metaresearch on patient-reported outcomes in trial protocols and results publications suggested large outcome reporting bias. J Clin Epidemiol 2025;185:111822. [Crossref] [PubMed]
  14. Spineli LM, Pandis N. Reporting bias: Notion, many faces and implications. Am J Orthod Dentofacial Orthop 2021;159:136-8. [Crossref] [PubMed]
  15. da Rosa Oliveira L, Elagami RA, Reis TM, et al. Selective Outcome Reporting Bias in Randomized Controlled Trials on Dental Caries in Children and Adolescents: A Meta-Research Study. Caries Res 2025;59:207-18. [Crossref] [PubMed]
  16. Hollon SD, Areán PA, Craske MG, et al. Development of clinical practice guidelines. Annu Rev Clin Psychol 2014;10:213-41. [Crossref] [PubMed]
  17. Kirkham JJ, Altman DG, Chan AW, et al. Outcome reporting bias in trials: a methodological approach for assessment and adjustment in systematic reviews. BMJ 2018;362:k3802. [Crossref] [PubMed]
  18. Jackson JL, Balk EM, Hyun N, et al. Approaches to Assessing and Adjusting for Selective Outcome Reporting in Meta-analysis. J Gen Intern Med 2022;37:1247-53. [Crossref] [PubMed]
  19. Frosi G, Riley RD, Williamson PR, et al. Multivariate meta-analysis helps examine the impact of outcome reporting bias in Cochrane rheumatoid arthritis reviews. J Clin Epidemiol 2015;68:542-50. [Crossref] [PubMed]
  20. van Aert RCM, Wicherts JM. Correcting for outcome reporting bias in a meta-analysis: A meta-regression approach. Behav Res Methods 2024;56:1994-2012. [Crossref] [PubMed]
  21. Wang A, Menon R, Li T, et al. Has the degree of outcome reporting bias in surgical randomized trials changed? A meta-regression analysis. ANZ J Surg 2023;93:76-82. [Crossref] [PubMed]
  22. Littell JH, Gorman DM, Valentine JC, et al. PROTOCOL: Assessment of outcome reporting bias in studies included in Campbell systematic reviews. Campbell Syst Rev 2023;19:e1332. [Crossref] [PubMed]
  23. Sterne JAC, Savović J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ 2019;366:l4898. [Crossref] [PubMed]
  24. Minhas JS, Chithiramohan T, Wang X, et al. Oral antiplatelet therapy for acute ischaemic stroke. Cochrane Database Syst Rev 2022;1:CD000029. [PubMed]
  25. Ziganshina LE, Abakumova T, Nurkhametova D, et al. Cerebrolysin for acute ischaemic stroke. Cochrane Database Syst Rev 2023;10:CD007026. [PubMed]
  26. Guo Q, Cheng Y, Zhang C, et al. A search of only four key databases would identify most randomized controlled trials of acupuncture: A meta-epidemiological study. Res Synth Methods 2022;13:622-31. [Crossref] [PubMed]
  27. Clarke M, Williamson PR. Core outcome sets and systematic reviews. Syst Rev 2016;5:11. [Crossref] [PubMed]
  28. Wang H, Lu Z, Liu Y. Score test for missing at random or not under logistic missingness models. Biometrics 2023;79:1268-79. [Crossref] [PubMed]
  29. Jia Y, Huang D, Wen J, et al. Association between switching of primary outcomes and reported trial findings among randomized drug trials from China. J Clin Epidemiol 2021;132:10-7. [Crossref] [PubMed]
  30. Aryan R, Jagroop D, Danells CJ, et al. Publication Rate and Consistency of Registered Trials of Motor-Based Stroke Rehabilitation. Neurology 2021;96:617-26. [Crossref] [PubMed]
  31. Sendyk DI, Souza NV, César Neto JB, et al. Selective outcome reporting in root coverage randomized clinical trials. J Clin Periodontol 2021;48:867-77. [Crossref] [PubMed]
  32. Altman DG, Bland JM. Measurement in medicine: the analysis of method comparison studies. Journal of the Royal Statistical Society Series D: The Statistician 2017;32:307-17. [Crossref]
  33. Gartlehner G, Dobrescu A, Evans TS, et al. Average effect estimates remain similar as evidence evolves from single trials to high-quality bodies of evidence: a meta-epidemiologic study. J Clin Epidemiol 2016;69:16-22. [Crossref] [PubMed]
  34. Papageorgiou SN, Xavier GM, Cobourne MT, et al. Registered trials report less beneficial treatment effects than unregistered ones: a meta-epidemiological study in orthodontics. J Clin Epidemiol 2018;100:44-52. [Crossref] [PubMed]
  35. Dechartres A, Ravaud P, Atal I, et al. Association between trial registration and treatment effect estimates: a meta-epidemiological study. BMC Med 2016;14:100. [Crossref] [PubMed]
  36. Dechartres A, Trinquart L, Boutron I, et al. Influence of trial sample size on treatment effect estimates: meta-epidemiological study. BMJ 2013;346:f2304. [Crossref] [PubMed]
  37. Ayorinde AA, Williams I, Mannion R, et al. Assessment of publication bias and outcome reporting bias in systematic reviews of health services and delivery research: A meta-epidemiological study. PLoS One 2020;15:e0227580. [Crossref] [PubMed]
  38. Schwab S, Kreiliger G, Held L. Assessing treatment effects and publication bias across different specialties in medicine: a meta-epidemiological study. BMJ Open 2021;11:e045942. [Crossref] [PubMed]
  39. Ge L, Tian JH, Li YN, et al. Association between prospective registration and overall reporting and methodological quality of systematic reviews: a meta-epidemiological study. J Clin Epidemiol 2018;93:45-55. [Crossref] [PubMed]
  40. Braakhekke M, Scholten I, Mol F, et al. Selective outcome reporting and sponsorship in randomized controlled trials in IVF and ICSI. Hum Reprod 2017;32:2117-22. [Crossref] [PubMed]
  41. Sah S. The paradox of disclosure: shifting policies from revealing to resolving conflicts of interest. Behavioural Public Policy 2023; [Crossref]
  42. Gandhi R, Jan M, Smith HN, et al. Comparison of published orthopaedic trauma trials following registration in Clinicaltrials.gov. BMC Musculoskelet Disord 2011;12:278. [Crossref] [PubMed]
  43. Hannink G, Gooszen HG, Rovers MM. Comparison of registered and published primary outcomes in randomized clinical trials of surgical interventions. Ann Surg 2013;257:818-23. [Crossref] [PubMed]
  44. Milette K, Roseman M, Thombs BD. Transparency of outcome reporting and trial registration of randomized controlled trials in top psychosomatic and behavioral health journals: A systematic review. J Psychosom Res 2011;70:205-17. [Crossref] [PubMed]
  45. Page MJ, McKenzie JE, Kirkham J, et al. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Cochrane Database Syst Rev 2014;2014:MR000035. [Crossref] [PubMed]
  46. Moro-Tejedor MN, Regaira-Martínez E. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Enferm Intensiva (Engl Ed) 2021;32:45-7. [Crossref] [PubMed]
  47. Jerke J, Velicu A, Winter F, et al. Publication bias in the social sciences since 1959: Application of a regression discontinuity framework. PLoS One 2025;20:e0305666. [Crossref] [PubMed]
  48. Core GRADE 4: rating certainty of evidence-risk of bias, publication bias, and reasons for rating up certainty. BMJ 2025;390:r1468. [Crossref] [PubMed]
Cite this article as: Zhang N, Long Y, Wang X, Wang X, Guo Q, Li Z, Du L. The threat of serious outcome reporting bias in randomized controlled trials on acute ischemic stroke to evidence synthesis: a meta-epidemiological study. Cardiovasc Diagn Ther 2025;15(6):1182-1193. doi: 10.21037/cdt-2025-212

Download Citation