Comparing methods for risk prediction of multicategory outcomes: dichotomized logistic regression vs. multinomial logit regression

Georgiy V. Bobashev; Abhik Das; Lei Li; Lei Li; Matthew A Rysavy; Georgiy V. Bobashev; Abhik Das

Comparing methods for risk prediction of multicategory outcomes

dichotomized logistic regression vs. multinomial logit regression

Li, L., Rysavy, M. A., Bobashev, G., & Das, A. (2024). Comparing methods for risk prediction of multicategory outcomes: dichotomized logistic regression vs. multinomial logit regression. BMC Medical Research Methodology, 24(1), 261. Article 261. https://doi.org/10.1186/s12874-024-02389-x

Copy citation

Abstract

BACKGROUND: Medical outcomes of interest to clinicians may have multiple categories. Researchers face several options for risk prediction of such outcomes, including dichotomized logistic regression and multinomial logit regression modeling. We aimed to compare these methods and provide guidance needed for practice.

METHODS: We described dichotomized logistic regression, multinomial continuation-ratio logit regression, which is an alternative to standard multinomial logit regression for ordinal outcomes, and logistic competing risks regression. We then applied these methods to develop prediction models of survival and neurodevelopmental outcomes based on the NICHD Extremely Preterm Birth Outcome Tool model. The statistical and practical advantages and flaws of these methods were examined. Both discrimination and calibration of the estimated logistic models of dichotomized outcomes and continuation-ratio logit model were assessed.

RESULTS: The dichotomized logistic models and multinomial continuation-ratio logit model had similar discrimination and calibration in predicting death and survival without neurodevelopmental impairment. But the continuation-ratio logit model had better discrimination and calibration in predicting neurodevelopmental impairment. The sum of predicted probabilities of outcome categories from the dichotomized logistic models could deviate from 100% substantially, ranging from 87.7 to 124.0%, and the dichotomized logistic model of neurodevelopmental impairment greatly overpredicted low risks and underpredicted high risks.

CONCLUSIONS: Estimating multiple logistic regression models of dichotomized outcomes may result in poorly calibrated predictions for an outcome with multiple ordinal categories. Multinomial continuation-ratio logit regression produces better calibrated predictions, constrains the sum of predicted probabilities to 100%, and has the advantages of simplicity in model interpretation, flexibility to include outcome category-specific predictors and random-effect terms for patient heterogeneity by hospital. It also accounts for mutual dependence among multiple categories and accommodates competing risks.

Publications Info

To contact an RTI author, request a report, or for additional information about publications by our experts, send us your request.

publications@rti.org

RTI shares its evidence-based research - through peer-reviewed publications and media - to ensure that it is accessible for others to build on, in line with our mission and scientific standards.

Meet the Experts

Navigate to Abhik Das

Abhik Das

Recent Publications

Article

US consumer and healthcare professional preferences for combination COVID-19 and influenza vaccines

December 31, 2025

Article

Plain language summary of mortality rates of patients with Parkinson’s disease psychosis who were treated either with pimavanserin or with different second-generation (atypical) antipsychotics

December 31, 2025

Article

A comprehensive GPS-based analysis of activity spaces in early and late pregnancy using the ActMAP framework

December 01, 2025

Article

IPECAD modeling workshop 2023 cross comparison challenge on cost-effectiveness models in Alzheimer's disease

April 01, 2025

View All Publications