Abstract
Cancer patient survival prediction remains a critical challenge in oncology, complicated by significant inter-patient and intra-tumor heterogeneity. The advent of high-throughput multi-omics technologies offers an unprecedented opportunity to characterize this heterogeneity, yet integrating and interpreting such diverse data for accurate prognosis is complex. This study proposes a novel Bayesian Dirichlet Process Mixture (DPM) model to address these challenges, specifically designed for patient survival prediction in heterogeneous multi-omics cancer datasets. Our approach leverages the non-parametric nature of the DPM to automatically discover latent patient subgroups with distinct molecular profiles and survival outcomes, without requiring a pre-specified number of clusters. By integrating genomic, transcriptomic, and clinical data, the model provides a comprehensive characterization of each subgroup, enabling more precise risk stratification. Through rigorous evaluation on publicly available cancer cohorts, we demonstrate that the DPM model significantly improves prognostic accuracy compared to traditional methods and state-of-the-art machine learning approaches, while also offering enhanced biological interpretability of the identified patient clusters. This framework represents a significant step towards personalized medicine by identifying molecularly defined patient cohorts for tailored therapeutic strategies.