ADAPTIVE SPACECRAFT RELATIVE ON-OFF CONTROL VIA R EINFORCEMENT LEARNING WITH STABILITY GUARANTEES
Heading:
| KHOROSHYLOV, SV, SOROCHINSK, II, VV |
| Spase Sci.&Tehnol. 2026, 32 ;(3):03-16 |
| https://doi.org/10.15407/knit2026.03.003 |
| Publication Language: English |
Abstract: Spacecraft relative motion control under on-off TH constraints presents challenges for achieving required performance in terms of
accuracy, effi ciency, and robustness. Th is study introduces a stability-guided deep reinforcement learning framework that directly generates TH-fi ring policies without relying on thrust modulators. Th e proposed agent integrates a sliding-mode control structure with an explicit discrete-time Lyapunov-decrease penalty into the proximal policy optimization algorithm, ensuring stability while enabling effi cient learning. Numerical simulations demonstrate that the method satisfi es established performance requirements, including settling time, steady-state accuracy in both position and rate, and propellant effi ciency. Th e sliding-mode structure enhances the interpretability of the learned policies and provides a principled basis for selecting initial agent parameters, thereby improving sample effi ciency during training. Furthermore, the agent exhibits strong robustness to parameter uncertainties and maintains reliable performance across varying operating conditions. Its adaptability is further underscored by the capacity for progressive improvement through online learning. Th ese fi ndings highlight the potential of stability-guided deep reinforcement learning as a viable and eff ective approach for spacecraft relative control under discrete actuation constraints, off ering both theoretical rigor and practical applicability for future space missions. |
| Keywords: neural network, on-off control, reinforcement learning, sliding-mode control, spacecraft relative control, TH fi ring |
References:
1. Alpatov A. P., Cichocki F., Fokov A. A., Khoroshylov S. V., Merino M., Zakrzhevskii A. E. (2015). Algorithm for determination of force transmitted by plume of ion TH to orbital object using photo camera. 66th Int. Astronautical Congress (Jerusalem, Israel), 2239-2247.
2. Anthony T., Wie B., Carroll S. (1989). Pulse modulated control synthesis for a fl exible spacecraft . Guidance, Navigation and Control Conference.
https://doi.org/10.2514/6.1989-3433
3. Bernelli-Zazzera F., Mantegazza P., Nurzia V. (1998). Multi-pulse-width modulated control of linear systems. J. Guidance, Control, and Dynamics, 21(1), 64-70.
https://doi.org/10.2514/2.4198
4. Finn C., Abbeel P., Levine S. (2017). Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. Proc. 34th Int.Conf. on Machine Learning (Sydney, Australia), PMLR 70.
https://doi.org/10.48550/arXiv.1703.03400
5. Gaudet B., Linares R., Furfaro R. (2019). Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. arXiv:1907.06098, 1-19.
https://doi.org/10.48550/arXiv.1907.06098
https://doi.org/10.1016/j.actaastro.2020.02.036
6. Gaudet B., Linares R., Furfaro R. (2020). Adaptive guidance and integrated navigation with reinforcement meta-learning. Acta Astronautica, 169, 180-190.
https://doi.org/10.1016/j.actaastro.2020.01.007
7. Hovell K., Ulrich S. (2020). On deep reinforcement learning for spacecraft guidance. AIAA Scitech 2020 Forum.
https://doi.org/10.2514/6.2020-1600
8. Ieko T., Ochi Y., Kanai K., Ieko T., Ochi Y., Kanai K. (1997). A new digital redesign method for pulse-width modulation control systems. Guidance, Navigation, and Control Conference.
https://doi.org/10.2514/6.1997-3770
9. Izzo D., Martens M., Pan B. (2019). A survey on Artifi cial Intelligence Trends in spacecraft guidance dynamics and Control. Astrodynamics, 3(4), 287-299.
https://doi.org/10.1007/s42064-018-0053-6
10. Khoroshylov S. V. (2018). Relative Motion Control System of spacecraft for contactless space debris removal. Science and
Innovation, 14(4), 5-16.
https://doi.org/10.15407/scine14.04.005
11. Khoroshylov S. V., Redka M. O. (2020). Relative control of an underactuated spacecraft using reinforcement learning. Technical
Mechanics, 2020(4), 43-54.
https://doi.org/10.15407/itm2020.04.043
12. Khoroshylov S. V., Redka M. O. (2021). Deep learning for spacecraft guidance, navigation, and Control. Space Science and
Technology, 27(6), 38-52.
https://doi.org/10.15407/knit2021.06.038
13. Khoroshylov S., Redka M. (2025). Deep Learning for Space Applications. Communications in Computer and Information
Science, 39-46.
https://doi.org/10.1007/978-3-032-04731-1_5
14. Khoroshylov S. V., Wang C. (2024). Spacecraft relative on-off control via reinforcement learning. Space Science and Technology,
30(2), 3-14.
https://doi.org/10.15407/knit2024.02.003
15. Khosravi A., Sarhadi P. (2016). Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to
spacecraft attitude controller design. Automatika, 57(1), 212-220.
https://doi.org/10.7305/automatika.2016.07.618
16. Lapkhanov E., Khoroshylov S. (2019). Development of the aeromagnetic space debris deorbiting system. Eastern-European J.
Enterprise Technologies, 5 (101), 30-37.
https://doi.org/10.15587/1729-4061.2019.179382
17. Li S., Liu K., Liu M. (2024). Adaptive dynamic programming-based spacecraft attitude control under a tube-based framework.
Electronics, 13(22), 4575.
https://doi.org/10.3390/electronics13224575
18. Li W.-J., Cheng D.-Y., Liu X.-G., Wang Y.-B., Shi W.-H., Tang Z.-X., Gao F., Zeng F.-M., Chai H.-Y., Luo W.-B., Cong Q., Gao
Z.-L. (2019). On-orbit service (OOS) of spacecraft : A review of Engineering Developments. Progress in Aerospace Sciences,
108, 32-120.
https://doi.org/10.1016/j.paerosci.2019.01.004
19. Miller B. M., Rubinovich E. Ya. (2013). Discontinuous solutions in the optimal control problems and their representation
by singular space-time transformations. Automation and Remote Control, 74(12), 1969-2006.
https://doi.org/10.1134/s0005117913120047
20. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. (2016). Asynchronous Methods for Deep Reinforcement
Learning. arXiv preprint, ArXiv:1602.01783.
https://doi.org/10.48550/arXiv.1602.01783
21. Oestreich C. E., Linares R., Gondhalekar R. (2021). Autonomous Six-degree-of-freedom spacecraft docking with rotating
targets via reinforcement learning. J. Aerospace Information Systems, 18(7), 417-428.
https://doi.org/10.2514/1.I010914
22. Pan S. J., Yang Q. (2010). A survey on Transfer Learning. IEEE Transactions on Knowledge and Data Engineering, 22(10),
1345-1359.
https://doi.org/10.1109/TKDE.2009.191
23. Redka M. O., Khoroshylov S. V. (2022). Determination of the force impact of an ion thruster plume on an orbital object via
Deep Learning. Space Science and Technology, 28(5), 15-26.
https://doi.org/10.15407/knit2022.05.015
24. Redka M., Khoroshylov S. (2023). Convolutional neural networks for determining the ion beam impact on a space debris
object. Science and Innovation, 19(6), 19-30.
https://doi.org/10.15407/scine19.06.019
25. Robinett R. D., Parker G. G., Schaub H., Junkins J. L. (1997). Lyapunov optimal saturated control for Nonlinear Systems.
https://doi.org/10.2514/6.1997-112
26. J. Guidance, Control, and Dynamics, 20(6), 1083-1088.
https://doi.org/10.2514/2.4189
27. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. (2017). Proximal policy optimization algorithms. arXiv preprint,
arXiv:1707.06347.
https://doi.org/10.48550/arXiv.1707.06347
28. Serpelloni E., Maggiore M., Damaren C. J. (2017). Quaternion-based bang-bang attitude stabilizer for rotating rigid bodies.
https://doi.org/10.2514/6.2016-0368
29. J. Guidance, Control, and Dynamics, 40(6), 1523-1534.
https://doi.org/10.2514/1.g002271
https://doi.org/10.2514/1.G002271
30. Sharifi E., Damaren C. J. (2024). Hybrid nonlinear passivity-based control approach to magnetic-impulsive spacecraft attitude
regulation. J. Guidance, Control, and Dynamics, 47(7), 1377-1393.
https://doi.org/10.2514/1.g005915
https://doi.org/10.2514/1.G005915
31. Silver D., Schrittwieser J., Simonyan K., Antonoglou I., Huang A., Guez A., Hubert T., Baker L., Lai M., Bolton A., Chen Y.,
Lillicrap T., Hui F., Sifre L., van den Driessche G., Graepel T., Hassabis D. (2017). Mastering the game of Go without human
knowledge. Nature, 550(7676), 354-359.
https://doi.org/10.1038/nature24270
32. Song G., Buck N. V., Agrawal B. N. (1999). Spacecraft vibration reduction using pulse-width pulse-frequency modulated input
shaper. J. Guidance, Control, and Dynamics, 22(3), 433-440.
https://doi.org/10.2514/2.4415
33. Sorochinskii V. V., Khoroshylov S. I., Levchuk I. L., Dubovyk T. M., Huz H. M., Romanchuk O. O. (2025). On-off spacecraft
relative control in sliding mode via reinforcement learning. Technical Mechanics, 2025(4), 77-92.
https://doi.org/10.15407/itm2025.04.077
34. Steinberger M., Horn M., Fridman L. (2020). Variable-structure systems and sliding-mode control: From theory to practice. Springer.
https://doi.org/10.1007/978-3-030-36621-6
35. Taborda P., Matias H., Silvestre D., Lourenco P. (2024). Convex MPC and thrust allocation with Deadband for spacecraft
rendezvous. IEEE Control Systems Letters, 8, 1132-1137.
https://doi.org/10.1109/LCSYS.2024.3407611
36. Vallado A. (2007) Fundamentals of Astrodynamics and Applications (3rd ed.). Hawthorne, CA: Microcosm Press. 1055 p.
37. Yamanaka K., Ankersen F. (2002). New state transition matrix for relative motion on an arbitrary elliptical orbit. J. Guidance,
Control, and Dynamics, 25(1), 60-66.
https://doi.org/10.2514/2.4875
2. Anthony T., Wie B., Carroll S. (1989). Pulse modulated control synthesis for a fl exible spacecraft . Guidance, Navigation and Control Conference.
https://doi.org/10.2514/6.1989-3433
3. Bernelli-Zazzera F., Mantegazza P., Nurzia V. (1998). Multi-pulse-width modulated control of linear systems. J. Guidance, Control, and Dynamics, 21(1), 64-70.
https://doi.org/10.2514/2.4198
4. Finn C., Abbeel P., Levine S. (2017). Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. Proc. 34th Int.Conf. on Machine Learning (Sydney, Australia), PMLR 70.
https://doi.org/10.48550/arXiv.1703.03400
5. Gaudet B., Linares R., Furfaro R. (2019). Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. arXiv:1907.06098, 1-19.
https://doi.org/10.48550/arXiv.1907.06098
https://doi.org/10.1016/j.actaastro.2020.02.036
6. Gaudet B., Linares R., Furfaro R. (2020). Adaptive guidance and integrated navigation with reinforcement meta-learning. Acta Astronautica, 169, 180-190.
https://doi.org/10.1016/j.actaastro.2020.01.007
7. Hovell K., Ulrich S. (2020). On deep reinforcement learning for spacecraft guidance. AIAA Scitech 2020 Forum.
https://doi.org/10.2514/6.2020-1600
8. Ieko T., Ochi Y., Kanai K., Ieko T., Ochi Y., Kanai K. (1997). A new digital redesign method for pulse-width modulation control systems. Guidance, Navigation, and Control Conference.
https://doi.org/10.2514/6.1997-3770
9. Izzo D., Martens M., Pan B. (2019). A survey on Artifi cial Intelligence Trends in spacecraft guidance dynamics and Control. Astrodynamics, 3(4), 287-299.
https://doi.org/10.1007/s42064-018-0053-6
10. Khoroshylov S. V. (2018). Relative Motion Control System of spacecraft for contactless space debris removal. Science and
Innovation, 14(4), 5-16.
https://doi.org/10.15407/scine14.04.005
11. Khoroshylov S. V., Redka M. O. (2020). Relative control of an underactuated spacecraft using reinforcement learning. Technical
Mechanics, 2020(4), 43-54.
https://doi.org/10.15407/itm2020.04.043
12. Khoroshylov S. V., Redka M. O. (2021). Deep learning for spacecraft guidance, navigation, and Control. Space Science and
Technology, 27(6), 38-52.
https://doi.org/10.15407/knit2021.06.038
13. Khoroshylov S., Redka M. (2025). Deep Learning for Space Applications. Communications in Computer and Information
Science, 39-46.
https://doi.org/10.1007/978-3-032-04731-1_5
14. Khoroshylov S. V., Wang C. (2024). Spacecraft relative on-off control via reinforcement learning. Space Science and Technology,
30(2), 3-14.
https://doi.org/10.15407/knit2024.02.003
15. Khosravi A., Sarhadi P. (2016). Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to
spacecraft attitude controller design. Automatika, 57(1), 212-220.
https://doi.org/10.7305/automatika.2016.07.618
16. Lapkhanov E., Khoroshylov S. (2019). Development of the aeromagnetic space debris deorbiting system. Eastern-European J.
Enterprise Technologies, 5 (101), 30-37.
https://doi.org/10.15587/1729-4061.2019.179382
17. Li S., Liu K., Liu M. (2024). Adaptive dynamic programming-based spacecraft attitude control under a tube-based framework.
Electronics, 13(22), 4575.
https://doi.org/10.3390/electronics13224575
18. Li W.-J., Cheng D.-Y., Liu X.-G., Wang Y.-B., Shi W.-H., Tang Z.-X., Gao F., Zeng F.-M., Chai H.-Y., Luo W.-B., Cong Q., Gao
Z.-L. (2019). On-orbit service (OOS) of spacecraft : A review of Engineering Developments. Progress in Aerospace Sciences,
108, 32-120.
https://doi.org/10.1016/j.paerosci.2019.01.004
19. Miller B. M., Rubinovich E. Ya. (2013). Discontinuous solutions in the optimal control problems and their representation
by singular space-time transformations. Automation and Remote Control, 74(12), 1969-2006.
https://doi.org/10.1134/s0005117913120047
20. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. (2016). Asynchronous Methods for Deep Reinforcement
Learning. arXiv preprint, ArXiv:1602.01783.
https://doi.org/10.48550/arXiv.1602.01783
21. Oestreich C. E., Linares R., Gondhalekar R. (2021). Autonomous Six-degree-of-freedom spacecraft docking with rotating
targets via reinforcement learning. J. Aerospace Information Systems, 18(7), 417-428.
https://doi.org/10.2514/1.I010914
22. Pan S. J., Yang Q. (2010). A survey on Transfer Learning. IEEE Transactions on Knowledge and Data Engineering, 22(10),
1345-1359.
https://doi.org/10.1109/TKDE.2009.191
23. Redka M. O., Khoroshylov S. V. (2022). Determination of the force impact of an ion thruster plume on an orbital object via
Deep Learning. Space Science and Technology, 28(5), 15-26.
https://doi.org/10.15407/knit2022.05.015
24. Redka M., Khoroshylov S. (2023). Convolutional neural networks for determining the ion beam impact on a space debris
object. Science and Innovation, 19(6), 19-30.
https://doi.org/10.15407/scine19.06.019
25. Robinett R. D., Parker G. G., Schaub H., Junkins J. L. (1997). Lyapunov optimal saturated control for Nonlinear Systems.
https://doi.org/10.2514/6.1997-112
26. J. Guidance, Control, and Dynamics, 20(6), 1083-1088.
https://doi.org/10.2514/2.4189
27. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. (2017). Proximal policy optimization algorithms. arXiv preprint,
arXiv:1707.06347.
https://doi.org/10.48550/arXiv.1707.06347
28. Serpelloni E., Maggiore M., Damaren C. J. (2017). Quaternion-based bang-bang attitude stabilizer for rotating rigid bodies.
https://doi.org/10.2514/6.2016-0368
29. J. Guidance, Control, and Dynamics, 40(6), 1523-1534.
https://doi.org/10.2514/1.g002271
https://doi.org/10.2514/1.G002271
30. Sharifi E., Damaren C. J. (2024). Hybrid nonlinear passivity-based control approach to magnetic-impulsive spacecraft attitude
regulation. J. Guidance, Control, and Dynamics, 47(7), 1377-1393.
https://doi.org/10.2514/1.g005915
https://doi.org/10.2514/1.G005915
31. Silver D., Schrittwieser J., Simonyan K., Antonoglou I., Huang A., Guez A., Hubert T., Baker L., Lai M., Bolton A., Chen Y.,
Lillicrap T., Hui F., Sifre L., van den Driessche G., Graepel T., Hassabis D. (2017). Mastering the game of Go without human
knowledge. Nature, 550(7676), 354-359.
https://doi.org/10.1038/nature24270
32. Song G., Buck N. V., Agrawal B. N. (1999). Spacecraft vibration reduction using pulse-width pulse-frequency modulated input
shaper. J. Guidance, Control, and Dynamics, 22(3), 433-440.
https://doi.org/10.2514/2.4415
33. Sorochinskii V. V., Khoroshylov S. I., Levchuk I. L., Dubovyk T. M., Huz H. M., Romanchuk O. O. (2025). On-off spacecraft
relative control in sliding mode via reinforcement learning. Technical Mechanics, 2025(4), 77-92.
https://doi.org/10.15407/itm2025.04.077
34. Steinberger M., Horn M., Fridman L. (2020). Variable-structure systems and sliding-mode control: From theory to practice. Springer.
https://doi.org/10.1007/978-3-030-36621-6
35. Taborda P., Matias H., Silvestre D., Lourenco P. (2024). Convex MPC and thrust allocation with Deadband for spacecraft
rendezvous. IEEE Control Systems Letters, 8, 1132-1137.
https://doi.org/10.1109/LCSYS.2024.3407611
36. Vallado A. (2007) Fundamentals of Astrodynamics and Applications (3rd ed.). Hawthorne, CA: Microcosm Press. 1055 p.
37. Yamanaka K., Ankersen F. (2002). New state transition matrix for relative motion on an arbitrary elliptical orbit. J. Guidance,
Control, and Dynamics, 25(1), 60-66.
https://doi.org/10.2514/2.4875
