Abstract
As the commercial and scientific viability of asteroid mining transitions from theoretical speculation to operational planning, the deployment of cooperative Multi-Agent CubeSat systems in the Main Asteroid Belt presents an attractive, cost-effective paradigm. However, the highly perturbed, non-Keplerian gravitational environments surrounding irregular asteroids, combined with the strict fuel and computational constraints of CubeSats, render traditional trajectory optimization methods computationally intractable for real-time onboard execution. This paper presents an autonomous trajectory planning framework for a decentralized swarm of mining CubeSats utilizing Multi-Agent Deep Reinforcement Learning (MADRL). We implement a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm integrated with a continuous low-thrust propulsion model and a localized gravity-field simulator incorporating solar radiation pressure and mutual gravitational perturbations. Our agents are trained to cooperatively navigate from a parking orbit to target rendezvous coordinates on multiple chaotic asteroid trajectories while actively avoiding collisions and minimizing fuel consumption. Simulated results demonstrate that the proposed MADRL framework achieves a 94.6% rendezvous success rate, outperforming classical pseudominimax optimal control methods in computational efficiency by three orders of magnitude. The trained neural network policies operate within the strict computational limits of standard radiation-hardened CubeSat flight computers, enabling real-time autonomous path planning and adaptive trajectory correction under severe state-estimation uncertainties.