BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//UM//UM*Events//EN
CALSCALE:GREGORIAN
BEGIN:VTIMEZONE
TZID:America/Detroit
TZURL:http://tzurl.org/zoneinfo/America/Detroit
X-LIC-LOCATION:America/Detroit
BEGIN:DAYLIGHT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
DTSTART:20070311T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
DTSTART:20071104T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20260903T140455
DTSTART;TZID=America/Detroit:20261028T160000
DTEND;TZID=America/Detroit:20261028T170000
SUMMARY:Workshop / Seminar:Policy Gradient Flow for Stochastic Control
DESCRIPTION:Reinforcement learning has achieved remarkable success in landmark applications such as AlphaGo and AlphaFold\, and is increasingly expected to play a central role in emerging domains such as autonomous driving. A substantial body of work has established theoretical foundations for classical reinforcement learning algorithms\, including temporal-difference learning\, policy iteration\, and Q-learning. By contrast\, policy gradient descent\, despite its widespread use\, remains less well understood from a theoretical perspective\, largely because of its non-convex structure and the complexity of the underlying functional space. Existing convergence analyses typically require uniform regularity assumptions on the policy\, viewed as a feedback control function\, along the gradient flow\; moreover\, the resulting convergence rates depend on these regularity bounds. In this work\, we significantly strengthen the existing convergence theory. Our key insight is to relate policy gradient descent to a mirror flow on the space of probability measures over controlled trajectories\, where the Bregman divergence is induced by the convex control cost. We rigorously prove a JKO-type convergence result showing that the discrete-time mirror flow converges to its continuous-time counterpart\, and we identify the limiting dynamics as a preconditioned policy gradient descent flow on the space of control processes. Leveraging this mirror-flow perspective\, we establish an exponential convergence rate that is independent of policy regularity. The convergence holds both in Bregman divergence and in the value of the associated control problem.
UID:148617-21904532@events.umich.edu
URL:https://events.umich.edu/event/148617
CLASS:PUBLIC
STATUS:CONFIRMED
CATEGORIES:Mathematics
LOCATION:East Hall - 1360
CONTACT:
END:VEVENT
END:VCALENDAR