RL Bite: Policy Gradient and Reinforce
Abstract
Till now we have considered only learning the Value or Q function and estimating the policy from those. In the next few posts, we are going to look into directly learning the policy. Why directly learn the policy? First, Q learning has a lot of issues involving the Deadly Triad; second, if we have continuous actions we cannot really use it; and lastly, Q learning always learns a deterministic policy, and in cases of partially observed stochastic environments (which is nearly always what we have), having a stochastic policy is proven to be better.
Policy Gradient
I teased a bit, yes we are going to use (stochastic) Gradient Descent to optimize (this is also known as policy search) the following loss:
- this measures the time spent in non-terminal states
- it is NOT a probability measure, since it is not normalised but can be normalized by exploiting if this yields
- it is NOT a probability measure, since it is not normalised but can be normalized by exploiting if this yields
- this is the marginal probability of being in state at time
- is the probability of going from to in steps
We abuse the notation a bit by treating as a probability measure we get:
And our final loss is:
Theorem
Now we differentiate the loss to get:
- this is also known as the score function and it is totally unrelated to the score function in Denoising Diffusion, which is the gradient with respect to a log probability:
Since it is a gradient, we can follow it to regions with higher reward! Now you can also see that the equation contains the Q function! Even in policy-based methods, we frequently use the Value, Q, or Advantage function (difference between the Value and Q function) since they stabilize the loss.
Reinforce
So naive policy gradient is said to have high variance! This is because we do a Monte Carlo rollout of the policy, which has low bias but, as said, high variance. To reduce the variance we need to introduce a baseline function, let’s look at the equations:
- this is the reward-to-go and is estimated using Markov Chain rollout of the policy, hence the high variance part
As mentioned before, we introduce a baseline function :
Baseline Functions
It can be anything, it has just one requirement:
A common choice is either the Value Function or Advantage Function
Estimator
Thus our update for the parameters becomes:
The update has an intuitive explanation:
We compute the sum of discounted future reward induced by a trajectory, compared to a baseline and if it is positive we increase to make this trajectory more likely otherwise we decrease , or in simpler terms we reinforce good behavior and penalize bad.
Algorithm

Final Remarks
This is the basis of Policy-Based Reinforcement Learning. In the next posts, we will look into an alternative approach called Actor-Critic methods, where instead of MC policy rollout we use Temporal Difference to estimate . However, as it turns out, neither Reinforce nor Actor-Critic methods guarantee monotonic improvement in the learned policy, and we turn to Policy Improvement methods which contain the currently popular Proximal Policy Optimization (PPO) method known from Large Language Models!