Spacecraft Active Defense Tactics in Uncertain Skies

Beijing Institute of Technology Press Co., Ltd

As high-value orbital regions become increasingly congested, pursuit–evasion confrontations among spacecraft have emerged as a critical issue in the field of space security. Active defense strategies, in which a target spacecraft releases defensive vehicles to implement counter-interception against the pursuer, constitute a three-body engagement scenario involving the target, the pursuer, and the defender. However, in actual engagements, the presence of observation noise, incomplete information, and multi-strategy adversarial conditions arising from uncertain pursuit policies render existing guidance methods—whether based on unilateral optimal control, differential games, or conventional reinforcement learning—incapable of simultaneously addressing the multiple challenges posed by unknown pursuit strategies, information deficiencies, and high-maneuverability confrontations. Consequently, how to achieve cooperative adaptive guidance between the target and the defender under conditions of incomplete observations and variable pursuit policies has become a key technological bottleneck for enhancing spacecraft survivability.

In a recent study published in Space: Science & Technology, the research team from the School of Aeronautics and Astronautics, Sun Yat-sen University, proposed an active defense guidance method for spacecraft based on an adaptive dueling double deep Q-network (D3QN). The study models the three-body engagement scenario as a multi-agent partially observable Markov decision process (Multi-POMDP) and designs a fusion network architecture integrating convolutional neural networks (CNNs) and gated recurrent units (GRUs). By stacking incomplete observations over the time dimension and extracting situational features via CNNs, while capturing temporal dependencies through GRUs, the approach effectively addresses the partial observability problem under multi-strategy adversarial conditions. Simultaneously, a continuous reward function in the form of potential-field differences is constructed based on the zero-effort miss distance, ensuring training stability and policy optimality. Simulation results demonstrate that the proposed method achieves a decision frequency of 85 Hz in a simulated single-chip microprocessor environment, satisfying engineering real-time requirements. In 1,000 Monte Carlo simulations, the evasion success rate reaches 99.7%, significantly outperforming the optimal switching cooperative guidance law (51.8%) and various reinforcement learning baselines. Even under extreme conditions with observation noise up to 40 times the typical value, the method maintains an evasion success rate of 73.5%, exhibiting exceptional robustness. This research provides a highly adaptive and robust solution for active defense guidance of spacecraft under incomplete information and multi-strategy adversarial conditions, offering significant engineering application value for enhancing the survivability of high-value spacecraft in complex space confrontation scenarios.

First, this paper focuses on the active defense guidance problem for spacecraft within a three-body engagement scenario involving a target, a pursuer, and a defender, where the target spacecraft must coordinate with the defender to evade a pursuer possessing maneuver advantages. In actual confrontations, the presence of observation noise, incomplete information, and the possibility that the pursuer may adopt multiple interception strategies—such as optimal control guidance, proportional navigation, or differential game guidance—endows the problem with dual complexities, namely multi-strategy adversarial conditions and partial observability. As illustrated in Fig. 1, the three-body engagement scenario comprises two pairs of pursuit–evasion relationships, namely target–pursuer and defender–pursuer, and the target and defender are required to maneuver cooperatively to achieve effective defense. The study models the problem as a multi-agent partially observable Markov decision process (Multi-POMDP), in which the agents cannot access the complete state information and must rely on noisy local observations (range and line-of-sight angles) for decision-making. To address this challenge, a reinforcement learning guidance framework based on an adaptive dueling double deep Q-network (D3QN) is designed. The core concept lies in: extracting situational features and temporal dependencies from sequences of incomplete observations through a fusion network integrating convolutional neural networks and gated recurrent units, while enhancing estimation stability via value function decomposition in the dueling double deep Q-network, ultimately generating cooperative maneuver commands for both the target and the defender.

Second, this paper elaborates on the working principles and key design of the AD³QN algorithm. As shown in Fig. 2, the algorithm stacks current and historical incomplete observations along the time dimension into a 2D tensor; a convolutional neural network (CNN) is then applied to extract spatial features to identify critical situational information in the multi-strategy adversarial environment. Subsequently, a gated recurrent unit (GRU) models the temporal features, producing a history tensor that encapsulates trajectory historical characteristics through the update gate and reset gate mechanisms, thereby providing a dynamic basis for policy decision-making. To address the issue of incomplete information, the algorithm restructures the experience replay buffer to store the stacked observation tensor, action, history tensor, and reward, enabling the network to leverage temporal correlation information during training. In the dueling double deep Q-network component, the state value function and the action advantage function are estimated separately, effectively reducing the estimation variance in multi-strategy environments. Additionally, a continuous reward function in the form of potential-field differences is constructed based on the zero-effort miss distance—prior to the defender–pursuer encounter, the reward guides the defender to reduce its zero-effort miss distance with respect to the pursuer to achieve interception; thereafter, it guides the target to increase its zero-effort miss distance from the pursuer to accomplish evasion. It is theoretically proven that this reward formulation does not alter the optimal policy, thereby ensuring training stability and policy optimality.

Finally, this paper validates the effectiveness of the proposed method through a series of numerical experiments. Fig. 3 presents a comparison of training stability; mainstream reinforcement learning algorithms—including Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), Proximal Policy Optimization (PPO), and Deep Recurrent Q-Learning (DRQN)—all struggle to achieve stable policy optimization under multi-strategy adversarial conditions and incomplete information. In contrast, the AD³QN algorithm demonstrates excellent convergence and stability, owing to its fusion network architecture and incomplete-information processing mechanism. Computational efficiency analysis indicates that the AD³QN achieves a decision frequency of 85 Hz in a simulated single-chip microprocessor environment, satisfying the requirements of spacecraft actuation systems and representing an improvement of approximately 30% over the DRQN method. In sample simulations, the defender successfully deviates the pursuer from the target by approximately 20 m. Monte Carlo statistical results from 1,000 simulations, as shown in Table 10, reveal that the proposed method attains an evasion success rate of 99.7%, significantly outperforming the optimal switching cooperative guidance law (51.8%) and various reinforcement learning baselines (below 0.2%). Robustness analysis in Fig. 4 demonstrates that even under extreme conditions with observation noise up to 40 times the typical value—i.e., range noise of 400 m and line-of-sight angle noise of 40 mrad—the method maintains an evasion success rate of 73.5%. Parameter sensitivity analysis reveals that an increase in the pursuer's maneuver capability leads to a reduction in the target's miss distance by approximately 30%, whereas enhanced maneuverability of the defender can effectively improve the evasion performance. This study provides a highly adaptive and robust solution for active defense guidance of spacecraft under incomplete information and multi-strategy adversarial conditions.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.