Academic Journal

Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction.

Bibliographic Details
Title: Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction.
Authors: YING, DONGHAO1 donghaoy@berkeley.edu, GUO, MENGZI AMY1 mengzi_guo@berkeley.edu, LEE, HYUNIN1 hyunin@berkeley.edu, DING, YUHAO2 yuhao.ding3@gmail.com, LAVAEI, JAVAD1 lavaei@berkeley.edu, SHEN, ZUO-JUN MAX3 maxshen@hku.hk
Source: Journal of Artificial Intelligence Research. 2025, Vol. 83, p1-53. 53p.
Subject Terms: Markov processes, Algorithms
Abstract: We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an Õ(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments. [ABSTRACT FROM AUTHOR]
Database: Supplemental Index
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edo&genre=article&issn=10769757&ISBN=&volume=83&issue=&date=20250501&spage=1&pages=1-53&title=Journal of Artificial Intelligence Research&atitle=Policy-based%20Primal-Dual%20Methods%20for%20Concave%20CMDP%20with%20Variance%20Reduction.&aulast=YING%2C%20DONGHAO&id=DOI:10.1613/jair.1.18129
    Name: Full Text Finder (for New FTF UI) (ns324271)
    Category: fullText
    Text: Full Text Finder
    MouseOverText: Full Text Finder
Header DbId: edo
DbLabel: Supplemental Index
An: 187740087
RelevancyScore: 994
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 993.548583984375
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22YING%2C+DONGHAO%22">YING, DONGHAO</searchLink><relatesTo>1</relatesTo><i> donghaoy@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22GUO%2C+MENGZI+AMY%22">GUO, MENGZI AMY</searchLink><relatesTo>1</relatesTo><i> mengzi_guo@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22LEE%2C+HYUNIN%22">LEE, HYUNIN</searchLink><relatesTo>1</relatesTo><i> hyunin@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22DING%2C+YUHAO%22">DING, YUHAO</searchLink><relatesTo>2</relatesTo><i> yuhao.ding3@gmail.com</i><br /><searchLink fieldCode="AR" term="%22LAVAEI%2C+JAVAD%22">LAVAEI, JAVAD</searchLink><relatesTo>1</relatesTo><i> lavaei@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22SHEN%2C+ZUO-JUN+MAX%22">SHEN, ZUO-JUN MAX</searchLink><relatesTo>3</relatesTo><i> maxshen@hku.hk</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Artificial+Intelligence+Research%22">Journal of Artificial Intelligence Research</searchLink>. 2025, Vol. 83, p1-53. 53p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Markov+processes%22">Markov processes</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an Õ(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments. [ABSTRACT FROM AUTHOR]
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edo&AN=187740087
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1613/jair.1.18129
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 53
        StartPage: 1
    Subjects:
      – SubjectFull: Markov processes
        Type: general
      – SubjectFull: Algorithms
        Type: general
    Titles:
      – TitleFull: Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: YING, DONGHAO
      – PersonEntity:
          Name:
            NameFull: GUO, MENGZI AMY
      – PersonEntity:
          Name:
            NameFull: LEE, HYUNIN
      – PersonEntity:
          Name:
            NameFull: DING, YUHAO
      – PersonEntity:
          Name:
            NameFull: LAVAEI, JAVAD
      – PersonEntity:
          Name:
            NameFull: SHEN, ZUO-JUN MAX
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Text: 2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 10769757
          Numbering:
            – Type: volume
              Value: 83
          Titles:
            – TitleFull: Journal of Artificial Intelligence Research
              Type: main
ResultId 1