Keep the gradient flowing

Policy Gradients Part 2: Baselines

The Power of Subtracting Zero

In Part 1, we saw that the asymptotic variance of the vanilla REINFORCE estimator scales cubically ($O(T^3)$) with trajectory length $T$.

In this second post, we explore how baselines systematically tame this variance. We formulate them through the classical Monte Carlo technique of control variates, showing how centering …

Policy Gradients Part 1: The REINFORCE Estimator

I have a dirty secret. Well, I actually have many. But one of them is that I never understood the basic algorithms behind reinforcement learning. So I plan to remedy this with a series of blog posts, where I will cover foundational RL algorithms, from REINFORCE to the frontier.

Here's …