Evaluating discrimination of prediction models for recurrent clinical events.
Researchers
Thomas J Spain, Alexandra Hunt, Maria Sudell, Jane L Hutton, Hein Putter, Victoria Watson, John Blakey, Anthony Marson, Laura Jayne Bonnett
Abstract
Prediction models for clinical conditions involving recurrent events require performance metrics tailored to repeated outcomes. Although we and others have developed methods to evaluate calibration and related aspects of performance, approaches for assessing discrimination remain limited. Recent contributions, including concordance-based methods and Brier-type accuracy measures, address parts of this gap, but no widely adopted, flexible framework for discrimination exists. As discrimination is a core TRIPOD+AI-recommended metric, accessible methodology and software for evaluating how well recurrent event models separate higher- and lower-risk individuals are still needed. We propose adapted discrimination statistics based on comparing predicted and observed cumulative event counts, addressing limitations of conventional concordance approaches. Statistical uncertainty is evaluated using several resampling strategies, including delete-d jackknifing, with user-friendly R code provided. We illustrate the approach using four tie-handling methods (c statistic, Kendall's τₐ, Somers' D, Goodman-Kruskal's γ), four recurrent event models (negative binomial, zero-inflated negative binomial, Andersen-Gill, and Prentice-Williams-Peterson Total Time), and two clinical datasets: the SANAD epilepsy trial and an OPCRD asthma cohort. Discrimination varied substantially across models. In the asthma data, the Prentice-Williams-Peterson model showed the strongest discrimination (C = 0.939, 95% CI 0.935-0.944), with rank-based metrics (τₐ = 0.603; Somers' D and γ = 0.879) also indicating very strong discrimination. The Andersen-Gill model performed moderately well (C = 0.823, 95% CI 0.815-0.830), while the negative binomial (C = 0.573) and zero-inflated negative binomial models (C = 0.586) showed weak discrimination. In the epilepsy data, the Prentice-Williams-Peterson model again performed best (C = 0.943, 95% CI 0.912-0.974), with all rank-based metrics above 0.85. The negative binomial, zero-inflated negative binomial and Andersen-Gill models showed moderate discrimination. This study provides practical methods and open-source code for evaluating the discrimination of prediction models for recurrent event data. Used alongside existing calibration tools, these methods support comprehensive performance evaluation and help standardise validation and reporting for prediction models developed for recurrent clinical events as recommended by the TRIPOD + AI guidelines.Source: PubMed (PMID: 42706560)View Original on PubMed