Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems

Λεπτομέρειες βιβλιογραφικής εγγραφής
Τίτλος: Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems
Συγγραφείς: Ahmed, Hani Hazza A., Fabri, Simon G., Bugeja, Marvin K., Camilleri, Kenneth P., ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics
Στοιχεία εκδότη: SCITEVENTS
Έτος έκδοσης: 2025
Συλλογή: University of Malta: OAR@UM / L-Università ta' Malta
Θεματικοί όροι: Reinforcement learning, Machine learning, Algorithms -- Mathematical models, Nonlinear systems, Python (Computer program language)
Περιγραφή: This paper introduces QARSA, a novel reinforcement learning algorithm that combines the strengths of off-policy and on-policy methods, specifically Q-learning and SARSA, for the dynamic control of nonlinear systems. Designed to leverage the sample efficiency of off-policy learning while preserving the stability and lower variance of on-policy approaches, QARSA aims to offer a balanced and robust learning framework. The algorithm is evaluated on the CartPole-v1 simulation environment using the OpenAI Gym framework, with performance compared against standalone Q-learning and SARSA implementations. The comparison is based on three critical metrics: average reward, stability, and sample efficiency. Experimental results demonstrate that QARSA outperforms both Q-learning and SARSA, achieving higher average rewards, stability, sample efficiency, and improved consistency in learned policies. These results demonstrate QARSA’s effectiveness in environments were maximizing long-term performance while maintaining learning stability is crucial. The study provides valuable insights for the design of hybrid reinforcement learning algorithms for continuous control tasks. ; peer-reviewed
Τύπος εγγράφου: conference object
Γλώσσα: English
Relation: https://www.um.edu.mt/library/oar/handle/123456789/138993
DOI: 10.5220/0013836700003982
Διαθεσιμότητα: https://www.um.edu.mt/library/oar/handle/123456789/138993
https://doi.org/10.5220/0013836700003982
Rights: info:eu-repo/semantics/openAccess ; The copyright of this work belongs to the author(s)/publisher. The rights of this work are as defined by the appropriate Copyright Legislation or as modified by any successive legislation. Users may access this work and can make use of the information contained in accordance with the Copyright Legislation provided that the author must be properly acknowledged. Further distribution or reproduction in any format is prohibited without the prior permission of the copyright holder
Αριθμός Καταχώρησης: edsbas.FCADC913
Βάση Δεδομένων: BASE
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://www.um.edu.mt/library/oar/handle/123456789/138993#
    Name: EDS - BASE (ns324271)
    Category: fullText
    Text: View record from BASE
Header DbId: edsbas
DbLabel: BASE
An: edsbas.FCADC913
RelevancyScore: 993
AccessLevel: 3
PubType: Conference
PubTypeId: conference
PreciseRelevancyScore: 992.8095703125
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Ahmed%2C+Hani+Hazza+A%2E%22">Ahmed, Hani Hazza A.</searchLink><br /><searchLink fieldCode="AR" term="%22Fabri%2C+Simon+G%2E%22">Fabri, Simon G.</searchLink><br /><searchLink fieldCode="AR" term="%22Bugeja%2C+Marvin+K%2E%22">Bugeja, Marvin K.</searchLink><br /><searchLink fieldCode="AR" term="%22Camilleri%2C+Kenneth+P%2E%22">Camilleri, Kenneth P.</searchLink><br /><searchLink fieldCode="AR" term="%22ICINCO+2025+-+22nd+International+Conference+on+Informatics+in+Control%2C+Automation+and+Robotics%22">ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics</searchLink>
– Name: Publisher
  Label: Publisher Information
  Group: PubInfo
  Data: SCITEVENTS
– Name: DatePubCY
  Label: Publication Year
  Group: Date
  Data: 2025
– Name: Subset
  Label: Collection
  Group: HoldingsInfo
  Data: University of Malta: OAR@UM / L-Università ta' Malta
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms+--+Mathematical+models%22">Algorithms -- Mathematical models</searchLink><br /><searchLink fieldCode="DE" term="%22Nonlinear+systems%22">Nonlinear systems</searchLink><br /><searchLink fieldCode="DE" term="%22Python+%28Computer+program+language%29%22">Python (Computer program language)</searchLink>
– Name: Abstract
  Label: Description
  Group: Ab
  Data: This paper introduces QARSA, a novel reinforcement learning algorithm that combines the strengths of off-policy and on-policy methods, specifically Q-learning and SARSA, for the dynamic control of nonlinear systems. Designed to leverage the sample efficiency of off-policy learning while preserving the stability and lower variance of on-policy approaches, QARSA aims to offer a balanced and robust learning framework. The algorithm is evaluated on the CartPole-v1 simulation environment using the OpenAI Gym framework, with performance compared against standalone Q-learning and SARSA implementations. The comparison is based on three critical metrics: average reward, stability, and sample efficiency. Experimental results demonstrate that QARSA outperforms both Q-learning and SARSA, achieving higher average rewards, stability, sample efficiency, and improved consistency in learned policies. These results demonstrate QARSA’s effectiveness in environments were maximizing long-term performance while maintaining learning stability is crucial. The study provides valuable insights for the design of hybrid reinforcement learning algorithms for continuous control tasks. ; peer-reviewed
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: conference object
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: NoteTitleSource
  Label: Relation
  Group: SrcInfo
  Data: https://www.um.edu.mt/library/oar/handle/123456789/138993
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.5220/0013836700003982
– Name: URL
  Label: Availability
  Group: URL
  Data: https://www.um.edu.mt/library/oar/handle/123456789/138993<br />https://doi.org/10.5220/0013836700003982
– Name: Copyright
  Label: Rights
  Group: Cpyrght
  Data: info:eu-repo/semantics/openAccess ; The copyright of this work belongs to the author(s)/publisher. The rights of this work are as defined by the appropriate Copyright Legislation or as modified by any successive legislation. Users may access this work and can make use of the information contained in accordance with the Copyright Legislation provided that the author must be properly acknowledged. Further distribution or reproduction in any format is prohibited without the prior permission of the copyright holder
– Name: AN
  Label: Accession Number
  Group: ID
  Data: edsbas.FCADC913
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.FCADC913
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.5220/0013836700003982
    Languages:
      – Text: English
    Subjects:
      – SubjectFull: Reinforcement learning
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Algorithms -- Mathematical models
        Type: general
      – SubjectFull: Nonlinear systems
        Type: general
      – SubjectFull: Python (Computer program language)
        Type: general
    Titles:
      – TitleFull: Combining off-policy and on-policy reinforcement learning for dynamic control of nonlinear systems
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Ahmed, Hani Hazza A.
      – PersonEntity:
          Name:
            NameFull: Fabri, Simon G.
      – PersonEntity:
          Name:
            NameFull: Bugeja, Marvin K.
      – PersonEntity:
          Name:
            NameFull: Camilleri, Kenneth P.
      – PersonEntity:
          Name:
            NameFull: ICINCO 2025 - 22nd International Conference on Informatics in Control, Automation and Robotics
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-locals
              Value: edsbas
            – Type: issn-locals
              Value: edsbas.oa
ResultId 1