Copula Based Fusion of Clinical and Genomic Machine Learning Risk Scores for Breast Cancer Risk Stratification
Abstract
Clinical and gene-expression models predict breast cancer outcomes, but simple linear fusion ignores dependence between their risk scores.
Using METABRIC, we tested whether modeling the joint distribution of clinical and gene-expression scores improved stratification of 5-year cancer-specific mortality.
We defined clinical and mRNA-expression predictor views, trained classifiers, and obtained out-of-fold probabilities through 5-fold cross-validation.
The scores were transformed into pseudo-observations on (0,1)^2 and used to fit Gaussian, Clayton, Gumbel, and Frank copulas.
The clinical model discriminated better than the gene-expression model (AUC 0.783 vs 0.721).
Frank had the smallest goodness-of-fit statistic, with Gaussian performing similarly.
Copula fusion did not improve ROC-AUC over the clinical model.
However, joint score groups showed clear survival differences, with patients scoring high on both views having the poorest outcomes.
Competing-risks analysis showed the same pattern for cancer-death incidence.
We also conducted an external evaluation in independent TCGA data using shared predictors and a harmonized 5-year overall-mortality endpoint.
Copula-fused, individual, and simple-fusion scores showed comparable discrimination with overlapping confidence intervals.
All received the same METABRIC-based recalibration.
No gene met the prespecified stability criterion under repeated cross-validated permutation importance, so gene-level findings were treated as exploratory.
Copulas provide an explicit, interpretable description of dependence between clinical and gene-expression risk scores and support descriptive joint-group analyses.
This methodological study does not establish superior prediction, validated clinical risk categories, or clinical utility.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요