Abstract
Personalized treatment sequencing in oncology represents one of the most critical challenges in precision medicine, where clinical decision-makers must continuously balance therapeutic efficacy against severe cumulative toxicity across dynamic disease trajectories. Standard reinforcement learning (RL) frameworks applied to observational electronic health records often struggle with treatment-selection bias, unobserved confounders, and the off-policy evaluation problem, leading to clinically unsafe policy recommendations. In this study, we propose a Causal-Aware Offline Reinforcement Learning (CA-ORL) framework that integrates structural causal models and doubly robust off-policy policy evaluation into a conservative actor-critic architecture for sequential oncology recommendations. By explicitly modeling time-varying confounding and estimating individualized counterfactual treatment effects, our framework constrains dynamic treatment regimens to clinically viable and robust strategies. We evaluate our approach using a composite cohort of longitudinal observational data from real-world non-small cell lung cancer (NSCLC) records alongside validated semi-synthetic simulation benchmarks. The experimental results demonstrate that CA-ORL achieves a 14.8% relative improvement in estimated 3-year progression-free survival while simultaneously reducing cumulative Grade III/IV hematological toxicity events by 21.3% compared to standard clinician baseline policies and non-causal offline RL baselines. These findings highlight the vital role of causal identification in guiding safe, data-driven therapeutic pathways in modern oncology.