Abstract
Environmental lead (Pb) exposure remains an enduring public health challenge worldwide, precipitating permanent neurodevelopmental and cognitive deficits in pediatric populations even at blood lead levels below established regulatory action thresholds. While conventional epidemiological models frequently evaluate environmental contamination and genetic vulnerability in isolation, neurodevelopmental outcomes stem from complex gene-environment interactions. In this study, we formulated, optimized, and validated a machine learning framework that integrates multi-source environmental lead exposure data—including blood lead levels (BLL), geocoded soil and municipal drinking water concentrations, and vintage housing risk metrics—with high-resolution genomic data from 1,420 pediatric subjects aged 24 to 72 months. Genomic profiling encompassed 14 candidate single nucleotide polymorphisms (SNPs) in toxicokinetic and neurodevelopmental pathways, notably within ALAD, HFE, BDNF, and COMT. Utilizing an extreme gradient boosting (XGBoost) architecture coupled with Shapley Additive Explanations (SHAP), our multimodal framework achieved an area under the receiver operating characteristic curve (AUROC) of 0.864 (95% CI: 0.832–0.896) and a root mean square error (RMSE) of 5.82 for predicting continuous cognitive performance indices, significantly surpassing unimodal exposure models (AUROC 0.718) and standard multivariable linear regression (AUROC 0.694). SHAP interaction analyses identified pronounced synergistic effects between elevated BLL and the ALAD rs1800435 and BDNF rs6265 polymorphisms, demonstrating that genetic variation substantially modulates cognitive vulnerability to low-level lead insults. This computational approach establishes a robust paradigm for precision environmental health, facilitating early, targeted screening and individualized risk mitigation for environmentally vulnerable pediatric cohorts.