Session

Poster Session 1

Location

Salt Palace Convention Center, Salt Lake City, UT

Abstract

As interest in small satellite Rendezvous, Proximity Operations, and Docking (RPOD) grows, so does the need for high-level autonomy in space systems. Current systems relying on constrained hardware and ground-in-the-loop decision-making are ill-equipped for the evolving, contested space domain. Recent research on autonomous spacecraft guidance and control has moved beyond established optimal control and trajectory optimization methods to incorporate Reinforcement Learning (RL), though these efforts have predominantly focused on a single operational phase. This research advances small satellite docking autonomy by formulating the reference scenario as a Hierarchical Markov Decision Process (H-MDP), expanding the scope of the learned guidance and control algorithm to span the full mission profile rather than one phase. This structure rewards the agile deputy satellite for learning to transition between mission phases, modifying its behavior from long-range rendezvous through final approach, and for adapting to changing conditions such as a shift in the chief satellite's pointing mode. The framework uses the Basilisk astrodynamics simulator and BSK-RL library to train a Proximal Policy Optimization (PPO) agent under high-fidelity orbital dynamics, creating a lightweight, robust training pipeline. The analysis quantifies total delta-V, docking ingress characteristics, and mission success across randomized initial conditions, with the trained policy achieving a 99.2% docking success rate over a 500-run Monte Carlo analysis, demonstrating how H-MDP structures can bridge single-phase RL research and mission-ready autonomous G&C.

Document Type

Event

SSC26-P1-48 (1).pdf (7898 kB)
Paper

Share

COinS
 
Aug 23rd, 12:00 AM

Multi-Phase Guidance and Control for Small Satellite Docking via Deep Reinforcement Learning

Salt Palace Convention Center, Salt Lake City, UT

As interest in small satellite Rendezvous, Proximity Operations, and Docking (RPOD) grows, so does the need for high-level autonomy in space systems. Current systems relying on constrained hardware and ground-in-the-loop decision-making are ill-equipped for the evolving, contested space domain. Recent research on autonomous spacecraft guidance and control has moved beyond established optimal control and trajectory optimization methods to incorporate Reinforcement Learning (RL), though these efforts have predominantly focused on a single operational phase. This research advances small satellite docking autonomy by formulating the reference scenario as a Hierarchical Markov Decision Process (H-MDP), expanding the scope of the learned guidance and control algorithm to span the full mission profile rather than one phase. This structure rewards the agile deputy satellite for learning to transition between mission phases, modifying its behavior from long-range rendezvous through final approach, and for adapting to changing conditions such as a shift in the chief satellite's pointing mode. The framework uses the Basilisk astrodynamics simulator and BSK-RL library to train a Proximal Policy Optimization (PPO) agent under high-fidelity orbital dynamics, creating a lightweight, robust training pipeline. The analysis quantifies total delta-V, docking ingress characteristics, and mission success across randomized initial conditions, with the trained policy achieving a 99.2% docking success rate over a 500-run Monte Carlo analysis, demonstrating how H-MDP structures can bridge single-phase RL research and mission-ready autonomous G&C.