Policy Gradients Part 2: Baselines
The Power of Subtracting Zero
In Part 1, we saw that the asymptotic variance of the vanilla REINFORCE estimator scales cubically ($O(T^3)$) with trajectory length $T$.
In this second post, we explore how baselines systematically tame this variance. We formulate them through the classical Monte Carlo technique of control variates, showing how centering …