The Accuracy of the I2 Statistic for Detecting Baseline Imbalances and Selection Bias in Randomized Controlled Trials: A Simulation Study.
Researchers
Steffen Mickenautsch, Veerasamy Yengopal
Abstract
To assess the accuracy of the I<sup>2</sup> statistic, applied as a trial-adjusted, simulated comparator trials (SCTs) based I<sup>2</sup> test for detecting baseline imbalances and selection bias in randomized controlled trials (RCTs). Fifty non-biased and 50 biased trials set at 40% selection bias severity were simulated in MS Excel (Microsoft Corporation, Redmond, Washington, USA). All trials were tested using statistical significance testing and the I<sup>2</sup> test. Null hypotheses were tested that the sensitivities and specificities of both tests did not statistically significantly differ and that the selection bias category scores, established by the I<sup>2</sup> test, do not statistically significantly correlate with the p-values from statistical significance testing. McNemar test and Spearman's rank correlation were used. The risk of the I<sup>2</sup> test overestimating selection bias severity was also examined. I<sup>2</sup> test demonstrated higher sensitivity than significance testing for detecting both baseline imbalance and selection bias (p < 0.0001). Specificity for baseline imbalance detection was identical at 100% for both methods; 95% CI: 90%-100%. For selection bias, specificity was greater with significance testing than with the I² test (p = 0.0005). A significant negative, large (0.5 ≤ |r|) correlation emerged between I² scores and significance test p-values: Spearman's r = -0.61, p < 0.0001. All null hypotheses, with the exception of the test specificities for baseline imbalance detection, were rejected. The I<sup>2</sup> test caused 26% of 'true positive' trials to be over-classified by one severity category.  The I<sup>2</sup> test was significantly more accurate for identifying baseline imbalance and selection bias detection, and equally accurate for ruling out baseline imbalance compared with statistical significance testing. However, the I<sup>2</sup> test was significantly less accurate in detecting selection bias absence. The I<sup>2</sup> test's bias category scores were largely negatively associated with p-values and tended to overestimate bias severity. Negative I<sup>2</sup> test results appear reliable for ruling out baseline imbalances and selection bias in RCTs, while positive I<sup>2</sup> test results are highly accurate to detect baseline imbalances but should be interpreted cautiously when assessing selection bias.Source: PubMed (PMID: 42604264)View Original on PubMed