Abstract
Palladium-catalyzed cross-coupling reactions represent a cornerstone of modern synthetic organic chemistry, yet optimizing reaction parameters remains a time-consuming and resource-intensive process. In this study, we evaluate the application of machine learning (ML) algorithms—specifically Random Forest, Extreme Gradient Boosting (XGBoost), and Multilayer Perceptron (MLP)—for predicting reaction yields and outcomes in Suzuki-Miyaura and Buchwald-Hartwig cross-coupling reactions. Using a curated dataset of 2,480 reaction outcomes combined with structural, electronic, and steric descriptors calculated from RDKit and Density Functional Theory (DFT), predictive models were trained and validated. Among the evaluated models, the XGBoost algorithm demonstrated superior predictive performance, achieving a coefficient of determination (R²) of 0.89 and a mean absolute error (MAE) of 6.3% on an independent test set. Feature importance analysis utilizing SHapley Additive exPlanations (SHAP) highlighted ligand steric bulk, phosphine HOMO energy levels, and solvent dielectric constants as the key parameters driving yield variance. Furthermore, the model accurately predicted high-yielding conditions for challenging, sterically hindered substrate pairs. This chemoinformatics framework provides an accessible, data-driven approach to assist synthetic chemists in rationalizing condition selection and minimizing empirical trial-and-error optimization.