Academic Journal
Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction.
| Title: | Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction. |
|---|---|
| Authors: | YING, DONGHAO1 donghaoy@berkeley.edu, GUO, MENGZI AMY1 mengzi_guo@berkeley.edu, LEE, HYUNIN1 hyunin@berkeley.edu, DING, YUHAO2 yuhao.ding3@gmail.com, LAVAEI, JAVAD1 lavaei@berkeley.edu, SHEN, ZUO-JUN MAX3 maxshen@hku.hk |
| Source: | Journal of Artificial Intelligence Research. 2025, Vol. 83, p1-53. 53p. |
| Subject Terms: | Markov processes, Algorithms |
| Abstract: | We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an Õ(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments. [ABSTRACT FROM AUTHOR] |
| Database: | Supplemental Index |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://resolver.ebsco.com/c/fiv2js/result?sid=EBSCO:edo&genre=article&issn=10769757&ISBN=&volume=83&issue=&date=20250501&spage=1&pages=1-53&title=Journal of Artificial Intelligence Research&atitle=Policy-based%20Primal-Dual%20Methods%20for%20Concave%20CMDP%20with%20Variance%20Reduction.&aulast=YING%2C%20DONGHAO&id=DOI:10.1613/jair.1.18129 Name: Full Text Finder (for New FTF UI) (ns324271) Category: fullText Text: Full Text Finder MouseOverText: Full Text Finder |
|---|---|
| Header | DbId: edo DbLabel: Supplemental Index An: 187740087 RelevancyScore: 994 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 993.548583984375 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22YING%2C+DONGHAO%22">YING, DONGHAO</searchLink><relatesTo>1</relatesTo><i> donghaoy@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22GUO%2C+MENGZI+AMY%22">GUO, MENGZI AMY</searchLink><relatesTo>1</relatesTo><i> mengzi_guo@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22LEE%2C+HYUNIN%22">LEE, HYUNIN</searchLink><relatesTo>1</relatesTo><i> hyunin@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22DING%2C+YUHAO%22">DING, YUHAO</searchLink><relatesTo>2</relatesTo><i> yuhao.ding3@gmail.com</i><br /><searchLink fieldCode="AR" term="%22LAVAEI%2C+JAVAD%22">LAVAEI, JAVAD</searchLink><relatesTo>1</relatesTo><i> lavaei@berkeley.edu</i><br /><searchLink fieldCode="AR" term="%22SHEN%2C+ZUO-JUN+MAX%22">SHEN, ZUO-JUN MAX</searchLink><relatesTo>3</relatesTo><i> maxshen@hku.hk</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Artificial+Intelligence+Research%22">Journal of Artificial Intelligence Research</searchLink>. 2025, Vol. 83, p1-53. 53p. – Name: Subject Label: Subject Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Markov+processes%22">Markov processes</searchLink><br /><searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: We study Concave Constrained Markov Decision Processes (Concave CMDPs) where both the objective and constraints are defined as concave functions of the state-action occupancy measure. We propose the Variance-Reduced Primal-Dual Policy Gradient Algorithm (VR-PDPG), which updates the primal variable via policy gradient ascent and the dual variable via projected sub-gradient descent. Despite the challenges posed by the loss of additivity structure and the nonconcave nature of the problem, we establish the global convergence of VR-PDPG by exploiting a form of hidden concavity. In the exact setting, we prove an O(T-1/3) convergence rate for both the average optimality gap and constraint violation, which further improves to O(T-1/2) under strong concavity of the objective in the occupancy measure. In the sample-based setting, we demonstrate that VR-PDPG achieves an Õ(ε-4) sample complexity for ε-global optimality. Moreover, by incorporating a diminishing pessimistic term into the constraint, we show that VR-PDPG can attain a zero constraint violation without compromising the convergence rate of the optimality gap. Finally, we validate our methods through numerical experiments. [ABSTRACT FROM AUTHOR] |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edo&AN=187740087 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1613/jair.1.18129 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 53 StartPage: 1 Subjects: – SubjectFull: Markov processes Type: general – SubjectFull: Algorithms Type: general Titles: – TitleFull: Policy-based Primal-Dual Methods for Concave CMDP with Variance Reduction. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: YING, DONGHAO – PersonEntity: Name: NameFull: GUO, MENGZI AMY – PersonEntity: Name: NameFull: LEE, HYUNIN – PersonEntity: Name: NameFull: DING, YUHAO – PersonEntity: Name: NameFull: LAVAEI, JAVAD – PersonEntity: Name: NameFull: SHEN, ZUO-JUN MAX IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 05 Text: 2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 10769757 Numbering: – Type: volume Value: 83 Titles: – TitleFull: Journal of Artificial Intelligence Research Type: main |
| ResultId | 1 |