J. Guo
Please Note
48 records found
1
Reproducing Orbital Approach Trajectories using Wheeled Robotic Arms
A Method for Hardware-in-the-Loop Validation of Satellite Systems during Close-Proximity Operations
Defining the Engineering Envelope of VMA Based Fogponics for Lunar Greenhouses
Framework Development and Boundary Evidence from Decontextualized Lab Venue Experiments
Fault Tolerance in CubeSat Attitude Determination
Applying Machine Learning to Sensor Fault Detection in Federated Kalman Filters
...
Design of Attitude Controller for Alticube+
Design of LQR Controller for Alticube+ and Evaluation of the Pointing Performance Based on Reaction Wheel Jitter and Flexible Structure Interactions
This research aims to design a centralised LQR attitude controller that enables effective utilisation of reactions wheel to satisfy scientific motivated pointing control and knowledge requirements. The second objective is to research how reaction wheel jitter interacts with the large aggregated structure to degrade the pointing control, affecting the measurement accuracy. To achieve this, a comprehensive simulation framework was developed in MATLAB/Simulink, capable of modelling both rigid-body and flexible spacecraft dynamics within a unified environment. A centralised Linear Quadratic Regulator (LQR) combined with a Kalman filter was designed to address multi-axis attitude regulation and pointing knowledge requirements under realistic actuator and sensor constraints.
The flexible spacecraft model was formulated using Kane’s equations and a lumped-parameter representation of the dominant structural modes. Reaction wheel jitter was modelled via static imbalance effects, enabling realistic excitation of flexible modes. The simulation framework was verified through analytical comparisons for the rigid-body case and validated for the flexible model by comparison with a finite element model, including modal frequency alignment and structural response under controlled torque excitation.
Simulation results demonstrate that the centralised LQR controller achieves stable attitude convergence within the allocated \(450\,\mathrm{s}\) operation window for the majority of initial conditions. Monte Carlo analysis shows that approximately \(80\%\) of the simulated cases satisfy the absolute pointing error requirement of \(0.2\,\mathrm{deg}\), with performance primarily limited by reaction wheel saturation. Flexible dynamics introduce oscillations, particularly in the pitch axis, where reaction wheel jitter and inertia uncertainty result in pointing errors up to \(0.3\,\mathrm{deg}\) under worst-case conditions. Despite this, internal antenna misalignment and pointing knowledge errors remain within mission requirements for nominal operating conditions.
The results indicate that centralised LQR-based control remains a viable and effective solution for flexible assembled CubeSat systems such as Alticube+, provided that actuator saturation, structural flexibility, and uncertainty effects are explicitly accounted for during design. This work provides insight into the achievable pointing performance envelope of Alticube+'s architecture.
...
This research aims to design a centralised LQR attitude controller that enables effective utilisation of reactions wheel to satisfy scientific motivated pointing control and knowledge requirements. The second objective is to research how reaction wheel jitter interacts with the large aggregated structure to degrade the pointing control, affecting the measurement accuracy. To achieve this, a comprehensive simulation framework was developed in MATLAB/Simulink, capable of modelling both rigid-body and flexible spacecraft dynamics within a unified environment. A centralised Linear Quadratic Regulator (LQR) combined with a Kalman filter was designed to address multi-axis attitude regulation and pointing knowledge requirements under realistic actuator and sensor constraints.
The flexible spacecraft model was formulated using Kane’s equations and a lumped-parameter representation of the dominant structural modes. Reaction wheel jitter was modelled via static imbalance effects, enabling realistic excitation of flexible modes. The simulation framework was verified through analytical comparisons for the rigid-body case and validated for the flexible model by comparison with a finite element model, including modal frequency alignment and structural response under controlled torque excitation.
Simulation results demonstrate that the centralised LQR controller achieves stable attitude convergence within the allocated \(450\,\mathrm{s}\) operation window for the majority of initial conditions. Monte Carlo analysis shows that approximately \(80\%\) of the simulated cases satisfy the absolute pointing error requirement of \(0.2\,\mathrm{deg}\), with performance primarily limited by reaction wheel saturation. Flexible dynamics introduce oscillations, particularly in the pitch axis, where reaction wheel jitter and inertia uncertainty result in pointing errors up to \(0.3\,\mathrm{deg}\) under worst-case conditions. Despite this, internal antenna misalignment and pointing knowledge errors remain within mission requirements for nominal operating conditions.
The results indicate that centralised LQR-based control remains a viable and effective solution for flexible assembled CubeSat systems such as Alticube+, provided that actuator saturation, structural flexibility, and uncertainty effects are explicitly accounted for during design. This work provides insight into the achievable pointing performance envelope of Alticube+'s architecture.
Five imperfection-sensitive strategies were assessed using NASA Shell Buckling Knockdown Factor Project cylinders at 0, 2, and 4 bar, including SPLA, MPLA, EIA, MGI, and a distributed-force perturbation approach (DFPA). Reliability was quantified using the coefficient of variation, Kendall’s W, and the intraclass correlation coefficient.
Results show that pressurization reduces imperfection sensitivity by increasing geometric stiffness. Distributed and multiple-perturbation methods were the most stable across pressures, whereas localized approaches exhibited strong pressure sensitivity. Accurate pressurized buckling prediction, therefore, requires pressure-aware imperfection mechanisms.
...
Five imperfection-sensitive strategies were assessed using NASA Shell Buckling Knockdown Factor Project cylinders at 0, 2, and 4 bar, including SPLA, MPLA, EIA, MGI, and a distributed-force perturbation approach (DFPA). Reliability was quantified using the coefficient of variation, Kendall’s W, and the intraclass correlation coefficient.
Results show that pressurization reduces imperfection sensitivity by increasing geometric stiffness. Distributed and multiple-perturbation methods were the most stable across pressures, whereas localized approaches exhibited strong pressure sensitivity. Accurate pressurized buckling prediction, therefore, requires pressure-aware imperfection mechanisms.
This research addresses these challenges by developing visualization-based methods to interpret and analyze the learning dynamics of actor–critic RL algorithms applied to spacecraft attitude control. Specifically, it introduces a critic match loss landscape visualization method for online actor–critic algorithms, allowing the evolution of the critic network during training to be examined systematically. Network parameters are recorded at the end of each episode and projected onto a low-dimensional subspace using Principal Component Analysis (PCA). A fixed-target critic match loss is then defined using reference state samples and corresponding temporal-difference (TD) targets from a selected policy. Evaluating the loss over the principal component plane generates both three-dimensional landscapes and two-dimensional contour plots with overlaid training trajectories, thereby illustrating how the critic optimizes its parameters over time. Quantitative indices and random-direction projections are used to systematically compare learning behavior across different training runs and reduce reliance on a single PCA projection.
The method is demonstrated on the Action-Dependent Heuristic Dynamic Programming (ADHDP) algorithm, applied to cart-pole and spacecraft attitude control tasks. Comparative analysis shows how the loss landscape geometry corresponds to training stability and control performance. Moreover, the visualization framework is extended to off-policy RL by adapting it to the Soft Actor–Critic (SAC) algorithm, which is capable of convergent control performance in spacecraft tasks. Adjustments account for SAC’s twin-critic structure, target computations, and replay-based training while maintaining interpretive consistency with online learning results.
Finally, the framework is expanded to capture both actor and critic dynamics, integrating four complementary components: a three-dimensional critic match loss landscape, an actor loss landscape, trajectories combining time, TD error, and actor weight evolution, and state–TD plots highlighting areas of large TD fluctuations. Applied across multiple ADHDP variants with deep learning, this multi-perspective framework provides interpretable insights into the interactions between value estimation, policy optimization, and TD signals during training. The framework allows for diagnosis of instability mechanisms, improves understanding of actor–critic learning behavior, and supports the design and analysis of RL algorithms for dynamic and uncertain control systems.
Altogether, this research establishes a comprehensive visualization and interpretation framework for RL in spacecraft attitude control and other dynamic environments, enhancing transparency, interpretability, and reliability of actor–critic algorithms under uncertainty, and providing a foundation for future work in dynamic visualization, theoretical stability analysis, and experimental validation on integrated physical platforms. ...
This research addresses these challenges by developing visualization-based methods to interpret and analyze the learning dynamics of actor–critic RL algorithms applied to spacecraft attitude control. Specifically, it introduces a critic match loss landscape visualization method for online actor–critic algorithms, allowing the evolution of the critic network during training to be examined systematically. Network parameters are recorded at the end of each episode and projected onto a low-dimensional subspace using Principal Component Analysis (PCA). A fixed-target critic match loss is then defined using reference state samples and corresponding temporal-difference (TD) targets from a selected policy. Evaluating the loss over the principal component plane generates both three-dimensional landscapes and two-dimensional contour plots with overlaid training trajectories, thereby illustrating how the critic optimizes its parameters over time. Quantitative indices and random-direction projections are used to systematically compare learning behavior across different training runs and reduce reliance on a single PCA projection.
The method is demonstrated on the Action-Dependent Heuristic Dynamic Programming (ADHDP) algorithm, applied to cart-pole and spacecraft attitude control tasks. Comparative analysis shows how the loss landscape geometry corresponds to training stability and control performance. Moreover, the visualization framework is extended to off-policy RL by adapting it to the Soft Actor–Critic (SAC) algorithm, which is capable of convergent control performance in spacecraft tasks. Adjustments account for SAC’s twin-critic structure, target computations, and replay-based training while maintaining interpretive consistency with online learning results.
Finally, the framework is expanded to capture both actor and critic dynamics, integrating four complementary components: a three-dimensional critic match loss landscape, an actor loss landscape, trajectories combining time, TD error, and actor weight evolution, and state–TD plots highlighting areas of large TD fluctuations. Applied across multiple ADHDP variants with deep learning, this multi-perspective framework provides interpretable insights into the interactions between value estimation, policy optimization, and TD signals during training. The framework allows for diagnosis of instability mechanisms, improves understanding of actor–critic learning behavior, and supports the design and analysis of RL algorithms for dynamic and uncertain control systems.
Altogether, this research establishes a comprehensive visualization and interpretation framework for RL in spacecraft attitude control and other dynamic environments, enhancing transparency, interpretability, and reliability of actor–critic algorithms under uncertainty, and providing a foundation for future work in dynamic visualization, theoretical stability analysis, and experimental validation on integrated physical platforms.
This thesis provides a new, improved coverage analysis method and applies it for research on constellation design optimization. More specifically, the new coverage analysis method is used to evaluate the revisit time and discontinuous coverage of any constellation configuration. Using a semi-analytical ground-to-spacecraft calculation, the coverage analysis method proves to be both accurate and fast.
In optimization, the coverage analysis method can be used to evaluate the candidate constellation configurations. By applying the method in a satellite constellation optimization framework, various optimization algorithms could be researched and compared. It was found that the choice of optimization algorithm is dependent on the desired result of the user, such as providing one singular optimal constellation or providing a wider range of viable configurations. The genetic algorithm and the covariance matrix adaptation evolution strategy proved to perform the best in single objective optimization, while the non-dominated sorting genetic algorithm II proved a good alternative for multi objective optimization. ...
This thesis provides a new, improved coverage analysis method and applies it for research on constellation design optimization. More specifically, the new coverage analysis method is used to evaluate the revisit time and discontinuous coverage of any constellation configuration. Using a semi-analytical ground-to-spacecraft calculation, the coverage analysis method proves to be both accurate and fast.
In optimization, the coverage analysis method can be used to evaluate the candidate constellation configurations. By applying the method in a satellite constellation optimization framework, various optimization algorithms could be researched and compared. It was found that the choice of optimization algorithm is dependent on the desired result of the user, such as providing one singular optimal constellation or providing a wider range of viable configurations. The genetic algorithm and the covariance matrix adaptation evolution strategy proved to perform the best in single objective optimization, while the non-dominated sorting genetic algorithm II proved a good alternative for multi objective optimization.
Design of an Utility-Frontier Based Exploration Strategy Coupled with 3D Object Localization for Modular Open-Source Autonomous Rovers
A State Machine Framework Approach Implemented and Tested on the Lunar Rover Mini
Exploration is decomposed into reusable low-level state machines within a hierarchical architecture that handles frontier detection, clustering, and filtering, as well as position and orientation monitoring, implemented in RAFCON, DLR's open-source software tool to manage autonomous tasks. A utility function balances information gain, computed by performing 3D ray casting, and travel cost, determined by the estimated travel time to a frontier centroid, to decide the next frontier centroid.
Moreover, a frontier coverage algorithm is employed to determine the most efficient set of orientations that maximize information gain, based on the normalized cumulative entropy within the camera’s iFOV across the full azimuth range around each frontier centroid.
A parallel perception pipeline runs, in real time, a quantized, custom-trained YOLOv7 model on the LRM’s Intel NUC to detect objects of interest, and compute their 3D coordinates in the global map frame using stereo depth data and the camera’s intrinsic and extrinsic parameters.
The design was tested in DLR's Planetary Exploration Laboratory across a set of benchmarks and mission scenarios. Results show the proposed Utility With Edges strategy performs better than classical Closest-frontier and Entropy-only methods, with the Utility With Edges strategy improving exploration efficiency by 27% over the Closest-frontier baseline, while achieving accurate real time CPU-only object detection and localization.
Key limitations of this implementation include the computational cost of ray casting, the use of a full OctoMap that stores the full range of occupancy probabilities instead of a binary map, and drift in the visual odometry estimates.
Future work recommendation entail studying how more complex utility functions affect selecting the next frontier centroid, building a fully autonomous mission pipeline with autonomous object grasping, and testing the implementation of the open-source ready-to-use exploration strategy on other robotic platforms with user-defined parameters. ...
Exploration is decomposed into reusable low-level state machines within a hierarchical architecture that handles frontier detection, clustering, and filtering, as well as position and orientation monitoring, implemented in RAFCON, DLR's open-source software tool to manage autonomous tasks. A utility function balances information gain, computed by performing 3D ray casting, and travel cost, determined by the estimated travel time to a frontier centroid, to decide the next frontier centroid.
Moreover, a frontier coverage algorithm is employed to determine the most efficient set of orientations that maximize information gain, based on the normalized cumulative entropy within the camera’s iFOV across the full azimuth range around each frontier centroid.
A parallel perception pipeline runs, in real time, a quantized, custom-trained YOLOv7 model on the LRM’s Intel NUC to detect objects of interest, and compute their 3D coordinates in the global map frame using stereo depth data and the camera’s intrinsic and extrinsic parameters.
The design was tested in DLR's Planetary Exploration Laboratory across a set of benchmarks and mission scenarios. Results show the proposed Utility With Edges strategy performs better than classical Closest-frontier and Entropy-only methods, with the Utility With Edges strategy improving exploration efficiency by 27% over the Closest-frontier baseline, while achieving accurate real time CPU-only object detection and localization.
Key limitations of this implementation include the computational cost of ray casting, the use of a full OctoMap that stores the full range of occupancy probabilities instead of a binary map, and drift in the visual odometry estimates.
Future work recommendation entail studying how more complex utility functions affect selecting the next frontier centroid, building a fully autonomous mission pipeline with autonomous object grasping, and testing the implementation of the open-source ready-to-use exploration strategy on other robotic platforms with user-defined parameters.
Initially, the figures of merit most commonly used in mega-constellation design were identified. The most important figure of merit used in all types of missions, except for Earth observation, is visibility, i.e., the number of satellites visible from a point on the ground. For Earth observation missions, the revisit time is more relevant than visibility. Thus, the thesis focused on mission objectives where visibility is the main figure of merit, such as Satcom, Satnav, IoT, etc. There is also more literature and data available on such constellations, including Starlink, OneWeb, Project Kuiper, and others. Consequently, two closely related figures of merit were selected for this study: minimum visibility and mean visibility. The former ensures an N-fold uninterrupted coverage, while the latter provides a general metric of how many satellites are visible over time. Also, the main focus was on Walker constellations as this is the most common geometry used for mega-constellations.
The visibility computation begins with constellation propagation. Due to the large number of satellites in a mega-constellation, existing tools used at ESOC, such as Godot, are insufficient. Therefore, a self-written tool was developed, named mcdo (Mega-Constellation Design Optimisation), which utilises the basic functionality of Godot and takes advantage of NumPy by applying vectorisation to notoriously slow Python. This approach resulted in a decrease in computational time by a factor of 100x with respect to godot.cosmos.BallisticPropagator.
Moreover, the visibility computation was accelerated with the utilisation of graphic processing units (GPU) and simplifications such as North-South symmetry, longitude averaging, or estimating visibility for a single time instance. Depending on the methods and simulation setups used, a reduction of 400x-120,000x in computational time was achieved for visibility computation.
Additionally, the parametric analysis of large Walker constellations yielded some valuable discoveries. For instance, the number of planes P and phasing parameter F do not influence the mean visibility but may have a significant effect on minimum visibility. This means that P and F can be omitted when designing for the mean visibility. It was also revealed that the mean visibility curve scales linearly with the number of satellites N, allowing to avoid propagation of very large constellations, and to scale up the mean visibility curve from smaller constellations instead. Parametric analysis of other parameters provided a general insight into their effect on visibility.
The improved computational efficiency of visibility computation for large Walker constellations enabled the application of multi-shell constellation design. A case study was set up with a requirement of uninterrupted coverage of 50 satellites over European latitudes (35-70 deg), assuming that all shells were at the same altitude of 700 km with a goal to minimise N. The analysis revealed that for two- and three-shell layouts, the minimum N was attained when higher-inclination shells had more satellites than lower-inclination ones. Moreover, methods for design acceleration were discussed.
The "building blocks" method was proposed for designing mega-constellations with more than three shells. This method starts from placing shells at high inclinations and then gradually lowers the inclination of the subsequent shells. It takes advantage of visibility properties of Walker constellations: the higher-inclination shells can cover both high and low latitudes, while the lower-inclination shells can only cover low latitudes. This method reduces the computational time drastically: by a factor of 150x for three-shell layouts and even more for larger numbers of shells.
The presented research offered valuable insights into initial phases of design of the orbital layout of mega-constellations. The mcdo tool and the obtained results were already applied for internal projects at ESA, demonstrating their relevance and usefulness. ...
Initially, the figures of merit most commonly used in mega-constellation design were identified. The most important figure of merit used in all types of missions, except for Earth observation, is visibility, i.e., the number of satellites visible from a point on the ground. For Earth observation missions, the revisit time is more relevant than visibility. Thus, the thesis focused on mission objectives where visibility is the main figure of merit, such as Satcom, Satnav, IoT, etc. There is also more literature and data available on such constellations, including Starlink, OneWeb, Project Kuiper, and others. Consequently, two closely related figures of merit were selected for this study: minimum visibility and mean visibility. The former ensures an N-fold uninterrupted coverage, while the latter provides a general metric of how many satellites are visible over time. Also, the main focus was on Walker constellations as this is the most common geometry used for mega-constellations.
The visibility computation begins with constellation propagation. Due to the large number of satellites in a mega-constellation, existing tools used at ESOC, such as Godot, are insufficient. Therefore, a self-written tool was developed, named mcdo (Mega-Constellation Design Optimisation), which utilises the basic functionality of Godot and takes advantage of NumPy by applying vectorisation to notoriously slow Python. This approach resulted in a decrease in computational time by a factor of 100x with respect to godot.cosmos.BallisticPropagator.
Moreover, the visibility computation was accelerated with the utilisation of graphic processing units (GPU) and simplifications such as North-South symmetry, longitude averaging, or estimating visibility for a single time instance. Depending on the methods and simulation setups used, a reduction of 400x-120,000x in computational time was achieved for visibility computation.
Additionally, the parametric analysis of large Walker constellations yielded some valuable discoveries. For instance, the number of planes P and phasing parameter F do not influence the mean visibility but may have a significant effect on minimum visibility. This means that P and F can be omitted when designing for the mean visibility. It was also revealed that the mean visibility curve scales linearly with the number of satellites N, allowing to avoid propagation of very large constellations, and to scale up the mean visibility curve from smaller constellations instead. Parametric analysis of other parameters provided a general insight into their effect on visibility.
The improved computational efficiency of visibility computation for large Walker constellations enabled the application of multi-shell constellation design. A case study was set up with a requirement of uninterrupted coverage of 50 satellites over European latitudes (35-70 deg), assuming that all shells were at the same altitude of 700 km with a goal to minimise N. The analysis revealed that for two- and three-shell layouts, the minimum N was attained when higher-inclination shells had more satellites than lower-inclination ones. Moreover, methods for design acceleration were discussed.
The "building blocks" method was proposed for designing mega-constellations with more than three shells. This method starts from placing shells at high inclinations and then gradually lowers the inclination of the subsequent shells. It takes advantage of visibility properties of Walker constellations: the higher-inclination shells can cover both high and low latitudes, while the lower-inclination shells can only cover low latitudes. This method reduces the computational time drastically: by a factor of 150x for three-shell layouts and even more for larger numbers of shells.
The presented research offered valuable insights into initial phases of design of the orbital layout of mega-constellations. The mcdo tool and the obtained results were already applied for internal projects at ESA, demonstrating their relevance and usefulness.
This research utilizes the fusion of data from a scanning LiDAR with a long-wavelength infrared camera to estimate the relative pose of an unknown uncooperative target. Two separate bespoke pose estimation algorithms, color-ICP and Feature Matching, were developed and tested with laboratory experiments mimicking the close-approach phase with a target under various lighting conditions and relative motion rates. The color-ICP algorithm uses a thermal infrared-infused color-assisted Generalized Iterative Closest Points method, while the Feature Matching algorithm uses computer vision on LiDAR point-infused thermal images to track BRISK feature points in each frame to estimate pose.
In general, the color-ICP algorithm delivered more accurate results throughout the range of experiments, though the fusion was slightly detrimental while the target is being heated or cooled. The Feature Matching algorithm contains a large amount of tunable parameters, making the estimation highly sensitive yet
versatile, demonstrating that harsh lighting conditions can be mitigated with accurate features tracked after the implementation of image processing techniques. Overall, the end product shows promise as a light-agnostic remote sensing and pose estimation solution.
This research contributes to the advancement of active debris removal theory and explores two promising avenues for LiDAR-infrared sensor fusion for pose estimation, laying the groundwork for further iterations exploring this sensor pairing. The resulting use case is a conceivable scenario in which these sensors work together to supplement individual strengths and mitigate disadvantages throughout the approach phase of a debris removal mission. ...
This research utilizes the fusion of data from a scanning LiDAR with a long-wavelength infrared camera to estimate the relative pose of an unknown uncooperative target. Two separate bespoke pose estimation algorithms, color-ICP and Feature Matching, were developed and tested with laboratory experiments mimicking the close-approach phase with a target under various lighting conditions and relative motion rates. The color-ICP algorithm uses a thermal infrared-infused color-assisted Generalized Iterative Closest Points method, while the Feature Matching algorithm uses computer vision on LiDAR point-infused thermal images to track BRISK feature points in each frame to estimate pose.
In general, the color-ICP algorithm delivered more accurate results throughout the range of experiments, though the fusion was slightly detrimental while the target is being heated or cooled. The Feature Matching algorithm contains a large amount of tunable parameters, making the estimation highly sensitive yet
versatile, demonstrating that harsh lighting conditions can be mitigated with accurate features tracked after the implementation of image processing techniques. Overall, the end product shows promise as a light-agnostic remote sensing and pose estimation solution.
This research contributes to the advancement of active debris removal theory and explores two promising avenues for LiDAR-infrared sensor fusion for pose estimation, laying the groundwork for further iterations exploring this sensor pairing. The resulting use case is a conceivable scenario in which these sensors work together to supplement individual strengths and mitigate disadvantages throughout the approach phase of a debris removal mission.
Adaptive Control for Spacecraft Rendezvous
A reinforcement meta-learning approach
Thus, throughout this project a recurrent neural network was trained via reinforcement meta-learning to generate a control policy that can perform the final approach maneuver of a chaser spacecraft towards a rotating target. A feedforward network was also trained for comparison. The learning algorithm used to train the policy is the Proximal Policy Optimization algorithm, which is a modern actor-critic method that has shown good performance in several continuous control settings. A virtual environment was developed in Python to simulate the rendezvous scenario and collect data to train the policy.
Before beginning with the training process, the hyperparameters of the model were tuned to ensure a smooth and efficient learning process. The three components that required tuning were the learning algorithm, the architecture of the neural networks, and the reward function. Each of these components was tuned in turn, primarily through trial and error. This process required the learning algorithm to be executed multiple times, using a different combination of hyperparameters on each iteration. By repeating this process over a large search space, suitable hyperparameters were found for the learning algorithm and the neural networks. The hyperparameters were chosen to maximize the amount of reward achieved by the policy while maintaining a reasonable training runtime. The reward function was split into several components to guide the policy towards its objective, thereby improving the speed at which the policy learns. Each of the components of the reward function represented some partial goal that the controller had to accomplish. Tuning the relative weight of these components was a challenging process since it often leads to trade-offs between different policy behaviors. Once the tuning process was completed, a sensitivity study was performed to ensure that the model can be used for different kinds of rendezvous trajectories. The sensitivity analysis was performed by training the model on different scenarios, including different orbit altitudes, different distances from the target, different target sizes, and different target rotation speeds. The results of this study showed that the model could be applied to most of these scenarios without the need for any major changes.
After completing the tuning and the sensitivity analysis, the recurrent and the feedforward policies were each trained on a partially observable environment, and their performance was evaluated using a Monte Carlo simulation for a total of one thousand trajectories. The results showed that the recurrent policy was able to learn how to infer hidden information from the environment, which led it to have a much better performance than the feedforward policy. However, the recurrent policy was not without its limitations, since it could not always generate collision-free trajectories, especially when the target rotated at a faster rate. Overall, this thesis showed that reinforcement meta-learning can be a valuable tool for executing complex rendezvous maneuvers, which may be useful now that active debris removal missions are becoming a reality. Furthermore, this thesis also presented a description of how the model was designed and tuned, so that other machine learning practitioners may apply the technique to different scenarios. ...
Thus, throughout this project a recurrent neural network was trained via reinforcement meta-learning to generate a control policy that can perform the final approach maneuver of a chaser spacecraft towards a rotating target. A feedforward network was also trained for comparison. The learning algorithm used to train the policy is the Proximal Policy Optimization algorithm, which is a modern actor-critic method that has shown good performance in several continuous control settings. A virtual environment was developed in Python to simulate the rendezvous scenario and collect data to train the policy.
Before beginning with the training process, the hyperparameters of the model were tuned to ensure a smooth and efficient learning process. The three components that required tuning were the learning algorithm, the architecture of the neural networks, and the reward function. Each of these components was tuned in turn, primarily through trial and error. This process required the learning algorithm to be executed multiple times, using a different combination of hyperparameters on each iteration. By repeating this process over a large search space, suitable hyperparameters were found for the learning algorithm and the neural networks. The hyperparameters were chosen to maximize the amount of reward achieved by the policy while maintaining a reasonable training runtime. The reward function was split into several components to guide the policy towards its objective, thereby improving the speed at which the policy learns. Each of the components of the reward function represented some partial goal that the controller had to accomplish. Tuning the relative weight of these components was a challenging process since it often leads to trade-offs between different policy behaviors. Once the tuning process was completed, a sensitivity study was performed to ensure that the model can be used for different kinds of rendezvous trajectories. The sensitivity analysis was performed by training the model on different scenarios, including different orbit altitudes, different distances from the target, different target sizes, and different target rotation speeds. The results of this study showed that the model could be applied to most of these scenarios without the need for any major changes.
After completing the tuning and the sensitivity analysis, the recurrent and the feedforward policies were each trained on a partially observable environment, and their performance was evaluated using a Monte Carlo simulation for a total of one thousand trajectories. The results showed that the recurrent policy was able to learn how to infer hidden information from the environment, which led it to have a much better performance than the feedforward policy. However, the recurrent policy was not without its limitations, since it could not always generate collision-free trajectories, especially when the target rotated at a faster rate. Overall, this thesis showed that reinforcement meta-learning can be a valuable tool for executing complex rendezvous maneuvers, which may be useful now that active debris removal missions are becoming a reality. Furthermore, this thesis also presented a description of how the model was designed and tuned, so that other machine learning practitioners may apply the technique to different scenarios.
A key requirement of any ASDR missions is that during capture, no new space debris is to be generated during the process. However, when simulating net capturing in the literature, the potential to break of vulnerable structures, like antennas or solar panels is often neglected. Such elements may show an enhanced risk of failure, especially if these are already damaged, potentially contributing to even more space debris.
A discrete Multi-Spring-Damper net model was used to simulate the 20 m/s-frontal impact of a 30 m x 30 m net onto an ESA Envisat mock-up. The Envisat was modelled as a two rigid-body system with a Single-Degree-of-Freedom hinge connection. A sequential modelling strategy was implemented, which de-coupled all the necessary dynamic and structural models. More than two large sub-structures (the Ka-band antenna dish and solar array) were found to have a high likelihood of breaking, leading to the recommendation of several design mitigation strategies using two types of sensitivity analysis. With secondary space debris being generated, net capturing is found to be riskier than originally assumed throughout the literature. ...
A key requirement of any ASDR missions is that during capture, no new space debris is to be generated during the process. However, when simulating net capturing in the literature, the potential to break of vulnerable structures, like antennas or solar panels is often neglected. Such elements may show an enhanced risk of failure, especially if these are already damaged, potentially contributing to even more space debris.
A discrete Multi-Spring-Damper net model was used to simulate the 20 m/s-frontal impact of a 30 m x 30 m net onto an ESA Envisat mock-up. The Envisat was modelled as a two rigid-body system with a Single-Degree-of-Freedom hinge connection. A sequential modelling strategy was implemented, which de-coupled all the necessary dynamic and structural models. More than two large sub-structures (the Ka-band antenna dish and solar array) were found to have a high likelihood of breaking, leading to the recommendation of several design mitigation strategies using two types of sensitivity analysis. With secondary space debris being generated, net capturing is found to be riskier than originally assumed throughout the literature.