Abstract
Electroencephalography (EEG)-based brain-computer interfaces (BCIs) offer a promising non-invasive pathway to restore motor function in individuals with upper-limb amputations. However, conventional EEG-BCI systems struggle with continuous, multidirectional control due to the non-stationarity of neural signals, low signal-to-noise ratios, and the challenge of user-system co-adaptation. In this study, we present a novel closed-loop BCI framework that utilizes deep reinforcement learning (DRL) to achieve real-time, multidirectional control of a 5-degree-of-freedom (DoF) robotic prosthetic limb. By formulating the decoding task as a continuous Markov Decision Process, we implemented a Deep Deterministic Policy Gradient (DDPG) agent that maps real-time spatial-filtered EEG features directly to prosthetic joint velocities. The system was validated with ten healthy subjects and two transradial amputees performing virtual and physical target-reaching tasks. Over multiple sessions, the closed-loop DRL agent successfully adapted to the users' shifting neural patterns, yielding an average target-reaching success rate of 89.4% and reducing average calibration times by 64% compared to traditional static decoders. These findings demonstrate that deep reinforcement learning can effectively bridge the gap between non-stationary neural signals and complex robotic control, paving the way for more intuitive, adaptive neuroprosthetic devices.