Robust Peak-cost Constrained Reinforcement Learning
Mirrored from arXiv — Machine Learning for archival readability. Support the source by reading on the original site.
Computer Science > Machine Learning
Title:Robust Peak-cost Constrained Reinforcement Learning
Abstract:We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most epsilon. Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15457 [cs.LG] |
| (or arXiv:2607.15457v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15457
arXiv-issued DOI via DataCite (pending registration)
|
Submission history
From: Shilpa Mukhopadhyay [view email][v1] Thu, 16 Jul 2026 21:00:50 UTC (6,061 KB)
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — Machine Learning
-
Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory
Aug 12
-
Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification
Aug 12
-
CurveFP: Rational-Radix Logarithmic Datatypes with Closed Products for Language Models
Aug 12
-
Sheaf-Based Federated Representation Learning
Aug 12
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.