Presented By: Financial/Actuarial Mathematics Seminar - Department of Mathematics
Exploratory Optimal Reinsurance under the Mean-Variance Criterion
Austin Riis-Due, Waterloo
This paper proposes a Reinforcement Learning (RL) approach to the optimal reinsurance problem
when the insurer faces uncertainty about the insurance claim dynamics. To this end, we first formulate
an exploratory version of the problem as a relaxed stochastic control problem. Within a broad class of
parametric retention functions and general risk loading functions, we derive the closed-form equilibrium
policy under the continuous-time mean-variance criterion. This is achieved through a formal verification
theorem and solving classical solutions of a system of exploratory extended Hamilton-Jacobi-Bellman
(EEHJB) equations. We then establish a policy iteration theorem, showing that starting from any timeand state-homogeneous policy, policy iteration converges to the derived equilibrium policy. Next, we
develop a martingale orthogonality theorem, which serves as the foundation of our RL algorithm. The
algorithm is evaluated through simulation studies and real data from the U.S. National Flood Insurance
Program. Results demonstrate that the RL approach effectively learns unknown claim distributions,
tracks unobserved changes in claim dynamics and produces higher insurer surplus trajectories than the
maximum likelihood estimation (MLE) approach while simultaneously exhibiting greater robustness to
the choice of training window.
when the insurer faces uncertainty about the insurance claim dynamics. To this end, we first formulate
an exploratory version of the problem as a relaxed stochastic control problem. Within a broad class of
parametric retention functions and general risk loading functions, we derive the closed-form equilibrium
policy under the continuous-time mean-variance criterion. This is achieved through a formal verification
theorem and solving classical solutions of a system of exploratory extended Hamilton-Jacobi-Bellman
(EEHJB) equations. We then establish a policy iteration theorem, showing that starting from any timeand state-homogeneous policy, policy iteration converges to the derived equilibrium policy. Next, we
develop a martingale orthogonality theorem, which serves as the foundation of our RL algorithm. The
algorithm is evaluated through simulation studies and real data from the U.S. National Flood Insurance
Program. Results demonstrate that the RL approach effectively learns unknown claim distributions,
tracks unobserved changes in claim dynamics and produces higher insurer surplus trajectories than the
maximum likelihood estimation (MLE) approach while simultaneously exhibiting greater robustness to
the choice of training window.