Keep the gradient flowing

Policy Gradients Part 2: Baselines

The Power of Subtracting Zero

In Part 1, we saw that the asymptotic variance of the vanilla REINFORCE estimator scales cubically ($O(T^3)$) with trajectory length $T$.

In this second post, we explore how baselines systematically tame this variance. Rather than introducing baselines as an ad-hoc trick, we formulate them through the classical Monte …