Abstract
Personalized pedagogical interventions in digital learning environments are critical for improving student retention and academic performance; however, determining optimal individualized policies from observational log data is severely hindered by confounding bias. In this paper, we propose a scalable Deep Counterfactual Regressor (DCR) architecture tailored to estimate Conditional Average Treatment Effects (CATE) from high-dimensional student behavioral traces. By coupling deep representation learning with an Integral Probability Metric regularizer based on the Wasserstein distance, our model explicitly mitigates covariate shift between treatment and control cohorts while preserving predictive power for counterfactual learning outcomes. We evaluate our approach on both semi-synthetic education benchmarks and a large-scale observational dataset comprising over 280,000 online learners across multi-week modular courses. Empirical results demonstrate that the proposed framework achieves statistically significant improvements over classical causal estimators and state-of-the-art deep counterfactual baselines, reducing the Precision in Estimation of Heterogeneous Effect (PEHE) metric by 14.2%. Furthermore, policy simulation experiments indicate that deploying our causal-driven intervention recommendations yields an estimated 8.7% absolute increase in overall course completion rates compared to standard predictive heuristics, highlighting the utility of causal machine learning for evidence-based digital education.