Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for safe reinforcement learning in optimal control.

Bibliographic Details
Title: Self-Organizing Dual-Buffer Adaptive Clustering Experience Replay (SODACER) for safe reinforcement learning in optimal control.
Authors: Khalili-Amirabadi R; Department of Applied Mathematics, Ferdowsi University of Mashhad, Mashhad, Iran. roya.khalili.a@gmail.com., Jalaeian-Farimani M; Department of Electronics, Information and Bioengineering (DEIB), Politecnico di Milano, Milan, Italy., Solaymani-Fard O; Department of Applied Mathematics, Ferdowsi University of Mashhad, Mashhad, Iran.
Source: Scientific reports [Sci Rep] 2026 Mar 25; Vol. 16 (1). Date of Electronic Publication: 2026 Mar 25.
Publication Type: Journal Article
Language: English
Journal Info: Publisher: Nature Publishing Group Country of Publication: England NLM ID: 101563288 Publication Model: Electronic Cited Medium: Internet ISSN: 2045-2322 (Electronic) Linking ISSN: 20452322 NLM ISO Abbreviation: Sci Rep Subsets: MEDLINE
Imprint Name(s): Original Publication: London : Nature Publishing Group, copyright 2011-
MeSH Terms: Clustering Algorithms* , Reinforcement Machine Learning*, Humans ; Adaptive Algorithms ; Cluster Analysis ; Nonlinear Dynamics ; Soft Computing
Abstract: This paper proposes a novel reinforcement learning framework, named Self-Organizing Dual-buffer Adaptive Clustering Experience Replay (SODACER), designed to achieve safe and scalable optimal control of nonlinear systems. The proposed SODACER mechanism consists of a Fast-Buffer for rapid adaptation to recent experiences and a Slow-Buffer equipped with a self-organizing adaptive clustering mechanism to maintain diverse and non-redundant historical experiences. The adaptive clustering mechanism dynamically prunes redundant samples, optimizing memory efficiency while retaining critical environmental patterns. The approach integrates SODACER with Control Barrier Functions (CBFs) to guarantee safety by enforcing state and input constraints throughout the learning process. To enhance convergence and stability, the framework is combined with the Sophia optimizer, enabling adaptive second-order gradient updates. The proposed SODACER-Sophia's architecture ensures reliable, effective, and robust learning in dynamic, safety-critical environments, offering a generalizable solution for applications in robotics, healthcare, and large-scale system optimization. The proposed approach is validated on a nonlinear Human Papillomavirus (HPV) transmission model with multiple control inputs and safety constraints. Comparative evaluations against random and clustering-based experience replay methods demonstrate that SODACER achieves faster convergence, improved sample efficiency, and a superior bias-variance trade-off, while maintaining safe system trajectories, validated via the Friedman test.
(© 2026. The Author(s).)
Competing Interests: Declarations. Competing interests: The authors declare no competing interests.
References: Bian, T. & Jiang, Z. P. Reinforcement learning and adaptive optimal control for continuous-time nonlinear systems: A value iteration approach. IEEE Trans. Neural Netw. Learn. Syst. 33(7), 2781–2790 (2021). (PMID: 10.1109/TNNLS.2020.3045087)
Li, S. E. Deep reinforcement learning. In Reinforcement Learning for Sequential Decision and Optimal Control 365–402 (Springer Nature Singapore, Singapore, 2023). (PMID: 10.1007/978-981-19-7784-8_10)
Marvi, Z. & Kiumarsi, B. Safe reinforcement learning: A control barrier function optimization approach. Int. J. Robust Nonlinear Control 31(6), 1923–1940 (2021). (PMID: 10.1002/rnc.5132)
Ames, A. D., Xu, X., Grizzle, J. W. & Tabuada, P. Control barrier function based quadratic programs for safety critical systems. IEEE Trans. Autom. Control 62(8), 3861–3876 (2016). (PMID: 10.1109/TAC.2016.2638961)
Khalili-Amirabadi, R. & Solaymani-Fard, O. Combining hybrid metaheuristic algorithms and reinforcement learning to improve the optimal control of nonlinear continuous-time systems with input constraints. Comput. Electr. Eng. 116, 109179 (2024). (PMID: 10.1016/j.compeleceng.2024.109179)
Berkenkamp, F., Turchetta, M., Schoellig, A., & Krause, A. Safe model-based reinforcement learning with stability guarantees. Advances in neural information processing systems, 30 (2017).
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., & Ghavamzadeh, M. Lyapunov-based safe policy optimization for continuous control. arXiv preprint arXiv:1901.10031 (2019).
Adam, S., Busoniu, L. & Babuska, R. Experience replay for real-time reinforcement learning control. IEEE Trans. Syst. Man Cybern. Part C Appl. Rev. 42(2), 201–212 (2011). (PMID: 10.1109/TSMCC.2011.2106494)
Yang, D., Qin, X., Xu, X., Li, C. & Wei, G. Sample efficient reinforcement learning method via high efficient episodic memory. IEEE Access 8, 129274–129284 (2020). (PMID: 10.1109/ACCESS.2020.3009329)
Mnih, V. et al. Human-level control through deep reinforcement learning. Nature 518(7540), 529–533 (2015). (PMID: 10.1038/nature1423625719670)
Schaul, T. Prioritized Experience Replay. arXiv preprint arXiv:1511.05952 (2015).
Isele, D., & Cosgun, A. Selective experience replay for lifelong learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32, No. 1 (2018).
Zhang, S., & Sutton, R. S. A deeper look at experience replay. arXiv preprint arXiv:1712.01275 (2017).
Nagabandi, A., Finn, C., & Levine, S. Deep online learning via meta-learning: Continual adaptation for model-based rl. arXiv preprint arXiv:1812.07671 (2018).
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., & Abbeel, P. Continuous adaptation via meta-learning in nonstationary and competitive environments. arXiv preprint arXiv:1710.03641 (2017).
Li, M., Huang, T. & Zhu, W. Clustering experience replay for the effective exploitation in reinforcement learning. Pattern Recognit. 131, 108875 (2022). (PMID: 10.1016/j.patcog.2022.108875)
Sinha, S., Song, J., Garg, A., & Ermon, S., Experience replay with likelihood-free importance weights. In Learning for Dynamics and Control Conference, 110-123 (PMLR, 2022).
Chen, Z., Li, H. & Yan, B. Efficient training framework for multi-USV system based on off-policy deep reinforcement learning. Inf. Technol. Control 54(4), 1159–1177 (2025). (PMID: 10.5755/j01.itc.54.4.40496)
Chen, Z., Li, H. & Wang, Z. Directly Attention loss adjusted prioritized experience replay. Complex Intell. Syst. 11(6), 1–11 (2025). (PMID: 10.1007/s40747-025-01852-6)
Saldaña, F., Korobeinikov, A. & Barradas, I. Optimal control against the human papillomavirus: protection versus eradication of the infection. Abstr. Appl. Anal. 201(1), 4567825 (2019).
Malik, T., Imran, M. & Jayaraman, R. Optimal control with multiple human papillomavirus vaccines. J. Theor. Biol. 393, 179–193 (2016). (PMID: 10.1016/j.jtbi.2016.01.00426796222)
Malik, T., Reimer, J., Gumel, A., Elbasha, E. H. & Mahmud, S. The impact of an imperfect vaccine and pap cytology screening on the transmission of human papillomavirus and occurrence of associated cervical dysplasia and cancer. Math. Biosci. Eng. 10(4), 1173–1205 (2013). (PMID: 10.3934/mbe.2013.10.117323906207)
Brown, V. L. & Jane White, K. A. The role of optimal control in assessing the most cost-effective implementation of a vaccination programme: HPV as a case study. Math. Biosci. 231(2), 126–134 (2011). (PMID: 10.1016/j.mbs.2011.02.00921377481)
Jalaeian-F, M., Fateh, M. M. & Rahimiyan, M. Bi-level adaptive computed-current impedance controller for electrically driven robots. Robotica 39(2), 200–216 (2021). (PMID: 10.1017/S0263574720000314)
Liu, H., Li, Z., Hall, D., Liang, P., & Ma, T. Sophia: A scalable stochastic second-order optimizer for language model pre-training. arXiv preprint arXiv:2305.14342 (2023).
Jalaeian Farimani, M., Khalili Amirabadi, R., Esmaeili Ranjbar, M. & Samadzadeh, S. Event-triggered dynamic seed invasive weed optimization (ET-DSIWO): a nature-inspired approach for non-stationary optimization. Nonlinear Dyn. 113(20), 27611–27636 (2025). (PMID: 10.1007/s11071-025-11513-5)
Dawson, C., Gao, S. & Fan, C. Safe control with learned certificates: A survey of neural Lyapunov, barrier, and contraction methods for robotics and control. IEEE Trans. Robot. 39(3), 1749–1767 (2023). (PMID: 10.1109/TRO.2022.3232542)
Benatia, M. A., Hafsi, M. & Ayed, S. B. A continual learning approach for failure prediction under non-stationary conditions: Application to condition monitoring data streams. Comput. Ind. Eng. 204, 111049 (2025). (PMID: 10.1016/j.cie.2025.111049)
Khalili-Amirabadi, R., Solaymai-Fard, O. & Jalaeian-Farimani, M. Towards optimal control of HPV model using safe reinforcement learning with actor-critic neural networks. Expert Syst. Appl. 264, 125783 (2025). (PMID: 10.1016/j.eswa.2024.125783)
Khalili-Amirabadi, R., Jalaeian-Farimani, M. & Solaymani-Fard, O. LSTM-empowered reinforcement learning in bi-level optimal control for nonlinear systems with uncertain dynamics. ISA Transactions 168, 465–478 (2026). (PMID: 10.1016/j.isatra.2025.11.02741309415)
López-Vázquez, C. & Hochsztain, E. Extended and updated tables for the Friedman rank test. Commun. Stat.-Theory Methods 48(2), 268–281 (2019). (PMID: 10.1080/03610926.2017.1408829)
Contributed Indexing: Keywords: Adaptive clustering; Dual-buffer experience replay; HPV model; Nonlinear optimal control; Safe reinforcement learning
Entry Date(s): Date Created: 20260326 Date Completed: 20260713 Latest Revision: 20260714
Update Code: 20260715
PubMed Central ID: PMC13168364
DOI: 10.1038/s41598-026-44517-1
PMID: 41882163
Database: MEDLINE
Description
ISSN:2045-2322
DOI:10.1038/s41598-026-44517-1