Conference
Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems
| Τίτλος: | Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems |
|---|---|
| Συγγραφείς: | Ahmed, Hani Hazza A., Fabri, Simon G., Bugeja, Marvin K., Camilleri, Kenneth P., ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics |
| Στοιχεία εκδότη: | SCITEVENTS |
| Έτος έκδοσης: | 2025 |
| Συλλογή: | University of Malta: OAR@UM / L-Università ta' Malta |
| Θεματικοί όροι: | Reinforcement learning, Machine learning, Algorithms -- Mathematical models, Nonlinear systems, Python (Computer program language) |
| Περιγραφή: | This paper introduces QARSA, a novel reinforcement learning algorithm that combines the strengths of off-policy and on-policy methods, specifically Q-learning and SARSA, for the dynamic control of nonlinear systems. Designed to leverage the sample efficiency of off-policy learning while preserving the stability and lower variance of on-policy approaches, QARSA aims to offer a balanced and robust learning framework. The algorithm is evaluated on the CartPole-v1 simulation environment using the OpenAI Gym framework, with performance compared against standalone Q-learning and SARSA implementations. The comparison is based on three critical metrics: average reward, stability, and sample efficiency. Experimental results demonstrate that QARSA outperforms both Q-learning and SARSA, achieving higher average rewards, stability, sample efficiency, and improved consistency in learned policies. These results demonstrate QARSA’s effectiveness in environments were maximizing long-term performance while maintaining learning stability is crucial. The study provides valuable insights for the design of hybrid reinforcement learning algorithms for continuous control tasks. ; peer-reviewed |
| Τύπος εγγράφου: | conference object |
| Γλώσσα: | English |
| Relation: | https://www.um.edu.mt/library/oar/handle/123456789/138993 |
| DOI: | 10.5220/0013836700003982 |
| Διαθεσιμότητα: | https://www.um.edu.mt/library/oar/handle/123456789/138993 https://doi.org/10.5220/0013836700003982 |
| Rights: | info:eu-repo/semantics/openAccess ; The copyright of this work belongs to the author(s)/publisher. The rights of this work are as defined by the appropriate Copyright Legislation or as modified by any successive legislation. Users may access this work and can make use of the information contained in accordance with the Copyright Legislation provided that the author must be properly acknowledged. Further distribution or reproduction in any format is prohibited without the prior permission of the copyright holder |
| Αριθμός Καταχώρησης: | edsbas.FCADC913 |
| Βάση Δεδομένων: | BASE |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://www.um.edu.mt/library/oar/handle/123456789/138993# Name: EDS - BASE (ns324271) Category: fullText Text: View record from BASE |
|---|---|
| Header | DbId: edsbas DbLabel: BASE An: edsbas.FCADC913 RelevancyScore: 993 AccessLevel: 3 PubType: Conference PubTypeId: conference PreciseRelevancyScore: 992.8095703125 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Ahmed%2C+Hani+Hazza+A%2E%22">Ahmed, Hani Hazza A.</searchLink><br /><searchLink fieldCode="AR" term="%22Fabri%2C+Simon+G%2E%22">Fabri, Simon G.</searchLink><br /><searchLink fieldCode="AR" term="%22Bugeja%2C+Marvin+K%2E%22">Bugeja, Marvin K.</searchLink><br /><searchLink fieldCode="AR" term="%22Camilleri%2C+Kenneth+P%2E%22">Camilleri, Kenneth P.</searchLink><br /><searchLink fieldCode="AR" term="%22ICINCO+2025+-+22nd+International+Conference+on+Informatics+in+Control%2C+Automation+and+Robotics%22">ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics</searchLink> – Name: Publisher Label: Publisher Information Group: PubInfo Data: SCITEVENTS – Name: DatePubCY Label: Publication Year Group: Date Data: 2025 – Name: Subset Label: Collection Group: HoldingsInfo Data: University of Malta: OAR@UM / L-Università ta' Malta – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms+--+Mathematical+models%22">Algorithms -- Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Nonlinear+systems%22">Nonlinear systems</searchLink><br /><searchLink fieldCode="DE" term="%22Python+%28Computer+program+language%29%22">Python (Computer program language)</searchLink> – Name: Abstract Label: Description Group: Ab Data: This paper introduces QARSA, a novel reinforcement learning algorithm that combines the strengths of off-policy and on-policy methods, specifically Q-learning and SARSA, for the dynamic control of nonlinear systems. Designed to leverage the sample efficiency of off-policy learning while preserving the stability and lower variance of on-policy approaches, QARSA aims to offer a balanced and robust learning framework. The algorithm is evaluated on the CartPole-v1 simulation environment using the OpenAI Gym framework, with performance compared against standalone Q-learning and SARSA implementations. The comparison is based on three critical metrics: average reward, stability, and sample efficiency. Experimental results demonstrate that QARSA outperforms both Q-learning and SARSA, achieving higher average rewards, stability, sample efficiency, and improved consistency in learned policies. These results demonstrate QARSA’s effectiveness in environments were maximizing long-term performance while maintaining learning stability is crucial. The study provides valuable insights for the design of hybrid reinforcement learning algorithms for continuous control tasks. ; peer-reviewed – Name: TypeDocument Label: Document Type Group: TypDoc Data: conference object – Name: Language Label: Language Group: Lang Data: English – Name: NoteTitleSource Label: Relation Group: SrcInfo Data: https://www.um.edu.mt/library/oar/handle/123456789/138993 – Name: DOI Label: DOI Group: ID Data: 10.5220/0013836700003982 – Name: URL Label: Availability Group: URL Data: https://www.um.edu.mt/library/oar/handle/123456789/138993<br />https://doi.org/10.5220/0013836700003982 – Name: Copyright Label: Rights Group: Cpyrght Data: info:eu-repo/semantics/openAccess ; The copyright of this work belongs to the author(s)/publisher. The rights of this work are as defined by the appropriate Copyright Legislation or as modified by any successive legislation. Users may access this work and can make use of the information contained in accordance with the Copyright Legislation provided that the author must be properly acknowledged. Further distribution or reproduction in any format is prohibited without the prior permission of the copyright holder – Name: AN Label: Accession Number Group: ID Data: edsbas.FCADC913 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.FCADC913 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.5220/0013836700003982 Languages: – Text: English Subjects: – SubjectFull: Reinforcement learning Type: general – SubjectFull: Machine learning Type: general – SubjectFull: Algorithms -- Mathematical models Type: general – SubjectFull: Nonlinear systems Type: general – SubjectFull: Python (Computer program language) Type: general Titles: – TitleFull: Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Ahmed, Hani Hazza A. – PersonEntity: Name: NameFull: Fabri, Simon G. – PersonEntity: Name: NameFull: Bugeja, Marvin K. – PersonEntity: Name: NameFull: Camilleri, Kenneth P. – PersonEntity: Name: NameFull: ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2025 Identifiers: – Type: issn-locals Value: edsbas – Type: issn-locals Value: edsbas.oa |
| ResultId | 1 |