Aosong Feng 冯傲松

the world, rebuilt small enough to run

← ../writing/

Two policy gradients, one blind spot

This distinction appears across reinforcement learning. PPO and GRPO use the score-function route [1], [2], [3]. SAC and differentiable simulation use pathwise derivatives, although they obtain those derivatives from different places [4], [5]. The names will make more sense after we solve one small robot problem. First, here is the whole setup in one picture: the input stays fixed, but a stochastic policy can sample different actions.

A left-to-right flow diagram. The same robot observation enters a box labelled policy pi theta, which branches into three gripper closures at different horizontal positions around a fixed cup. Black triangular markers show the three sampled action positions relative to the cup-centre guide.
figure 1. The camera view and task stay the same, but the stochastic policy πθ does not return one guaranteed closing position. Each run samples a new action. The dashed guide marks the cup centre; the three black triangles mark three possible samples.

Start with one noisy robot action

Imagine that the robot repeats the same reach many times. The camera image, instruction, cup, and all but one joint stay fixed. The only number we record is where the gripper closes along the horizontal axis. Call this position aa, measured in centimetres from the cup’s centre.

We model the action as

a=μθ+σϵ,ϵ∼N(0,1).a = \mu_\theta + \sigma \epsilon, \qquad \epsilon \sim \mathcal{N}(0,1).

Here μθ\mu_\theta is the mean position chosen by policy parameters θ\theta, σ\sigma is the amount of exploration noise, and ϵ\epsilon is one draw from a standard normal distribution. In plain language: the policy aims at μθ\mu_\theta, but each attempt lands a little to the left or right.

The robot receives one point if the gripper closes inside the success interval G=[ℓ,r]G = [\ell,r], and zero otherwise:

R(a)=1{a∈G}.R(a) = \mathbf{1}\{a \in G\}.

The symbol 1{⋅}\mathbf{1}\{\cdot\} is an indicator: it equals 11 when the condition is true and 00 when it is false. Throughout the post, σ=3\sigma = 3 cm and G=[−2,2]G = [-2,2] cm. This reward has no gentle slope: crossing either edge changes the reward instantly from 00 to 11, or from 11 to 00.

Two panels compare nearby grasp positions. On the left, the action marker a lies inside the hatched success interval G, the cup stays upright, and the reward is one. On the right, a lies just outside G, the cup tips, and the reward is zero.
figure 2. The lower strips turn the robot scene into the one-dimensional model: a is the closing position and the hatched band is the success interval G. A small shift moves a across the boundary, so the outcome jumps from R = 1 to R = 0. There is no partial credit at the edge.

The second figure deliberately tells the same story twice. The robot drawings show the physical consequence; the strips below them show the mathematical abstraction. The next diagram adds the policy’s full action distribution on top of that line.

A Gaussian policy density centred at μ = −3 cm, with the area under the curve inside the hatched success interval G = [−2, 2] cm shaded as the success probability J(μ).
figure 3. The policy is centred at μ = −3 cm. The hatched band is the fixed success interval G = [−2, 2] cm, and the shaded area under the density is the success probability J(μ). Moving μ changes that area.

A policy update moves probability

Before deriving either gradient estimator, we need a precise description of what learning changes. Let λ\lambda run continuously from 00 to 11 as the policy moves from its old parameters to its new ones: θ(0)=θold\theta(0)=\theta_{\text{old}} and θ(1)=θnew\theta(1)=\theta_{\text{new}}. This is not the robot’s physical time tt. The index tt moves the robot through one rollout; λ\lambda moves the policy through one update.In the one-action example, we can simply set λ=μ\lambda=\mu. Increasing λ\lambda then means sliding the Gaussian policy to the right.

A rollout τ=(s0,a0,…,sT)\tau=(s_0,a_0,\ldots,s_T) is one complete sequence of states and actions. Let ρλ(τ)\rho_\lambda(\tau) be the probability density that the current policy and environment assign to that rollout, and let R(τ)R(\tau) be its total reward. The objective is the expected reward:

J(λ)=∫R(τ) ρλ(τ) dτ(1)J(\lambda) = \int R(\tau)\, \rho_\lambda(\tau) \, d\tau \tag{1}

Equation (1) says: consider every possible rollout, multiply its reward by its probability, and add the results. The integration variable dτd\tau means that the integral ranges over all possible rollouts. For a fixed task, the scoring rule RR stays the same. The policy update changes ρλ\rho_\lambda, so it changes which rollouts are likely.

We can describe this change as a flow of probability. Let vλ(τ)v_\lambda(\tau) be a velocity field: it tells us the direction and speed at which probability near rollout τ\tau moves as λ\lambda increases. Conservation of probability gives the continuity equation:

∂λρλ(τ)+∇τ⋅(ρλ(τ)vλ(τ))=0(2)\partial_\lambda \rho_\lambda(\tau) + \nabla_\tau \cdot \big(\rho_\lambda(\tau) v_\lambda(\tau)\big) = 0 \tag{2}

The first term, ∂λρλ\partial_\lambda\rho_\lambda, is the rate at which the density changes at a fixed location. The second term measures net probability flowing out of that location: ρλvλ\rho_\lambda v_\lambda is the probability flux, ∇τ ⁣⋅\nabla_\tau\!\cdot is its divergence, and the minus sign says that net outflow lowers the local density. In short, density changes because probability flows in or out; probability is not created or destroyed.

The one-dimensional robot example makes (2) easy to check. Set λ=μ\lambda=\mu and shift the Gaussian mean to the right. Every sampled action moves right at unit speed, so v=1v=1. The equation becomes ∂μρμ(a)=−∂aρμ(a)\partial_\mu\rho_\mu(a)=-\partial_a\rho_\mu(a): changing the mean has the same effect as translating the whole density in the opposite coordinate direction.

Probability in a region falls not because probability was destroyed, but because it went somewhere else.

Two Gaussian curves before and after a rightward policy shift cross a fixed hatched interval G, with inflow at the left boundary ℓ and outflow at the right boundary r.
figure 4. The success interval stays fixed while the policy density moves right. Probability enters through ℓ and leaves through r, so the change in success probability is ΔJ = inflow − outflow.

View one: keep the outcomes fixed

The first view keeps each possible rollout τ\tau fixed and asks how its probability changes. Differentiate (1) with respect to λ\lambda. Because the task reward RR does not depend on the policy update, the derivative acts only on ρλ\rho_\lambda:

dJdλ=∫R ∂λρλ dτ=Eτ∼ρλ[R(τ) ∂λlog⁡ρλ(τ)](3)\frac{dJ}{d\lambda} = \int R\, \partial_\lambda \rho_\lambda \, d\tau = \mathbb{E}_{\tau \sim \rho_\lambda}\big[R(\tau)\, \partial_\lambda \log \rho_\lambda(\tau)\big] \tag{3}

The last expression is the score-function, or likelihood-ratio, gradient. The reward R(τ)R(\tau) says how good the sampled rollout was. The score ∂λlog⁡ρλ(τ)\partial_\lambda\log\rho_\lambda(\tau) says how quickly the update changes that rollout’s relative probability. Their product is the rollout’s contribution to the gradient, and the expectation averages that contribution over sampled rollouts.

Why can this estimator work with a black-box environment? The probability of a rollout factors into the initial-state distribution p0p_0, the policy πθ\pi_\theta, and the environment dynamics PP:

ρθ(τ)=p0(s0)∏tπθ(at∣st)P(st+1∣st,at),∇θlog⁡ρθ(τ)=∑t∇θlog⁡πθ(at∣st).\rho_\theta(\tau)=p_0(s_0)\prod_t \pi_\theta(a_t\mid s_t)P(s_{t+1}\mid s_t,a_t), \qquad \nabla_\theta\log\rho_\theta(\tau)=\sum_t\nabla_\theta\log\pi_\theta(a_t\mid s_t).

The product on the left assigns a probability to the whole rollout. After taking the logarithm, that product becomes a sum. If the environment does not itself depend on θ\theta, the derivatives of p0p_0 and PP are zero, so only derivatives of the policy remain on the right. We need sampled states, actions, and rewards, plus ∇θlog⁡πθ\nabla_\theta\log\pi_\theta from the policy network. We do not need derivatives of the environment or reward. This is the key idea behind REINFORCE [7] and the policy gradient theorem [8].

This route is tolerant of black boxes, but it is not assumption-free. Moving the derivative through the integral must be valid, and the set of possible rollouts—the support of ρθ\rho_\theta—must behave regularly. If the support itself changes shape with θ\theta, the usual formula needs extra care [9].

Two aligned rows of six fixed bins. The bins remain in the same positions while their probability weights shift right into the success interval G.
figure 5. Score-function view: the possible outcomes—the bins—stay fixed. A policy update changes only how much probability weight each bin receives; it does not match an old sample to a new one.

View two: keep the randomness fixed

The second view matches an outcome under the old policy with an outcome under the new one. Write ϵ∼q\epsilon\sim q for a random draw from a fixed noise distribution, and write τ=Tλ(ϵ)\tau=T_\lambda(\epsilon) for the rollout produced by passing that draw through the policy and environment. In the robot example, this is simply aλ=μθ(λ)+σϵa_\lambda=\mu_{\theta(\lambda)}+\sigma\epsilon.

The important step is to give the old and new policies the same ϵ\epsilon. Think of ϵ\epsilon as a numbered lottery ticket. If both policies receive the same ticket, we can say how that particular sampled outcome moved. If they receive independent tickets, we can compare two distributions, but we cannot match one old sample to one new sample.

With shared noise, ∂Tλ(ϵ)/∂λ\partial T_\lambda(\epsilon)/\partial\lambda is the velocity of the sampled rollout. If the reward is differentiable along this path, the chain rule gives the pathwise gradient:

dJdλ=Eϵ∼q[∇τR(Tλ(ϵ)) ⁣⊤∂Tλ(ϵ)∂λ](4)\frac{dJ}{d\lambda} = \mathbb{E}_{\epsilon \sim q}\left[\nabla_\tau R\big(T_\lambda(\epsilon)\big)^{\!\top} \frac{\partial T_\lambda(\epsilon)}{\partial \lambda}\right] \tag{4}

The vector ∇τR\nabla_\tau R is the local reward slope: it points toward nearby rollouts with higher reward. The vector ∂Tλ/∂λ\partial T_\lambda/\partial\lambda says how the sampled rollout moves as the policy changes. Their dot product measures how quickly the reward changes along that motion, and the expectation averages over noise draws.

For a multi-step rollout, we can write at=gθ(st,ϵt)a_t=g_\theta(s_t,\epsilon_t) for the policy and st+1=f(st,at,ξt)s_{t+1}=f(s_t,a_t,\xi_t) for the environment. Here gθg_\theta maps a state and policy noise to an action; ff maps the current state, action, and environment noise ξt\xi_t to the next state. Once every ϵt\epsilon_t and ξt\xi_t is held fixed, the rollout becomes an ordinary computation graph, θ→a0→s1→⋯→R\theta\rightarrow a_0\rightarrow s_1\rightarrow\cdots\rightarrow R, and backpropagation can differentiate it. Stochastic computation graphs formalize how score-function and pathwise estimators can be mixed at different random nodes [10].

Eight action samples before and after an update are connected one to one by equal rightward arrows. One bold arrow carries a sample across the left boundary into G.
figure 6. Pathwise view: keeping the same noise ε pairs every old sample with its new location. Each action moves by the same amount; the bold arrow shows a sample crossing into G. The reward changes at that crossing, not along the flat regions on either side.

Integration by parts connects the two views

The two formulas look unrelated: (3) differentiates probability, while (4) differentiates reward. The continuity equation shows that they describe the same flow. Substitute ∂λρλ=−∇τ⋅(ρλvλ)\partial_\lambda\rho_\lambda=-\nabla_\tau\cdot(\rho_\lambda v_\lambda) into the score-function derivation, then use integration by parts to move the derivative from the probability flux onto the reward:

E[R ∂λlog⁡ρ]⏟watch the weights change  =  −∫R ∇⋅(ρv)⏟weights change because probability flows  =  E[v ⁣⊤∇R]⏟follow probability toward higher reward\underbrace{\mathbb{E}\big[R\, \partial_\lambda \log \rho\big]}_{\text{watch the weights change}} \;=\; \underbrace{-\int R\, \nabla \cdot (\rho v)}_{\text{weights change because probability flows}} \;=\; \underbrace{\mathbb{E}\big[v^{\!\top} \nabla R\big]}_{\text{follow probability toward higher reward}}

The left expression is the score-function view: keep locations fixed and watch their probability weights change. The middle expression replaces that weight change with probability flow. The right expression is the pathwise view: follow the moving probability and measure the reward slope along its path. Integration by parts is the operation that moves the derivative from ρ\rho to RR.

claim 1 (two views of one gradient).

Suppose (2) holds, the reward is differentiable along the flow, and the boundary term at infinity is zero. Then the score-function estimator (3) and the pathwise estimator (4) have the same expectation. They are two representations of the same change in the rollout distribution.

This transport interpretation of pathwise derivatives is developed by Jankowiak and Obermeyer [6]. Parmas and Sugiyama give a unified treatment of likelihood-ratio and reparameterization gradients as estimators of the same probability movement [11].

Three caveats are important. First, equal expectations do not imply equal finite-sample variance. Second, the velocity field is not unique, so different pathwise estimators can have the same mean. Third, claim 1 assumes that the ordinary reward gradient ∇τR\nabla_\tau R exists. Our hard 0/10/1 reward breaks that last assumption.

Where common algorithms fit

We can now sort familiar algorithms by one practical question: what derivative can the algorithm obtain outside the policy itself? The “view” column follows from that requirement.

familyviewderivative needed outside the policyhard R=1{a∈G}R=\mathbf{1}\{a\in G\}
REINFORCE, PPO, GRPO [1], [2], [7]fixed outcomesnone; only sampled returnsunbiased, but can be noisy
DDPG, TD3, SAC [4], [12], [13]fixed noise∇aQ(s,a)\nabla_a Q(s,a) from a learned criticdepends on the critic’s learned smoothness
SVG(1), SVG(∞\infty) [14]fixed noisederivatives through a learned dynamics modeldepends on the learned model and reward
differentiable simulation [5]fixed noisedynamics and reward derivatives from the simulatornaive autodiff returns zero in this example
Gumbel-Softmax, straight-through [15]fixed noisea differentiable relaxation of a discrete choiceproduces a biased surrogate gradient
table 1. The derivative each algorithm family needs outside the policy, and what a hard zero-or-one reward means for it. The last column refers to the robot example in this post.

The actor-critic row is easy to misunderstand. DDPG and SAC do not differentiate the real environment. They differentiate a learned critic Q(s,a)Q(s,a) with respect to the action. The smoothness requirement therefore falls on a neural network the algorithm controls, not directly on the world’s reward or dynamics. By contrast, methods that backpropagate through a learned dynamics model or a differentiable simulator need a longer differentiable path. Evolution strategies and random search sit outside both columns: they estimate changes in performance without forming ∇θlog⁡πθ\nabla_\theta\log\pi_\theta at all [16], [17].

One boundary is enough to fool autodiff

Return to the robot’s actual 0/10/1 reward and set the update coordinate λ\lambda equal to the policy mean μ\mu. The objective J(μ)J(\mu) is now simply the probability that a Gaussian action lands between ℓ\ell and rr. Let Φ\Phi denote the standard normal cumulative distribution function and ϕ\phi its probability density. Then both the objective and its exact derivative have closed forms:

J(μ)=Φ ⁣(r−μσ)−Φ ⁣(ℓ−μσ),∂J∂μ=1σ[ϕ ⁣(ℓ−μσ)−ϕ ⁣(r−μσ)].J(\mu) = \Phi\!\left(\frac{r - \mu}{\sigma}\right) - \Phi\!\left(\frac{\ell - \mu}{\sigma}\right), \qquad \frac{\partial J}{\partial \mu} = \frac{1}{\sigma}\left[\phi\!\left(\frac{\ell - \mu}{\sigma}\right) - \phi\!\left(\frac{r - \mu}{\sigma}\right)\right].

The formula for JJ subtracts the Gaussian probability to the left of ℓ\ell from the probability to the left of rr, leaving exactly the probability inside the success interval. The derivative has an equally concrete meaning: ϕ((ℓ−μ)/σ)/σ\phi((\ell-\mu)/\sigma)/\sigma is probability density entering through the left edge as the policy moves right, while ϕ((r−μ)/σ)/σ\phi((r-\mu)/\sigma)/\sigma is density leaving through the right edge. Their difference is inflow minus outflow.

The same conclusion follows directly from the continuity equation. Integrating (2) over G=[ℓ,r]G=[\ell,r] gives

dJdλ=ρλ(ℓ)vλ(ℓ)−ρλ(r)vλ(r).\frac{dJ}{d\lambda}=\rho_\lambda(\ell)v_\lambda(\ell)-\rho_\lambda(r)v_\lambda(r).

The two terms are the probability flux through the interval’s left and right edges. In our example v=1v=1, so the gradient is ρμ(ℓ)−ρμ(r)\rho_\mu(\ell)-\rho_\mu(r), matching the formula above and figure 4.This boundary is not the boundary term at infinity in claim 1. It appears because we integrate over the finite success interval GG.

The score-function view has no problem with this boundary because it never differentiates the reward. For one sampled action, its gradient estimate is

g^score=R(a)a−μσ2.\widehat g_{\mathrm{score}}=R(a)\frac{a-\mu}{\sigma^2}.

The reward R(a)R(a) says whether the sample succeeded. The factor (a−μ)/σ2(a-\mu)/\sigma^2 is the derivative of the Gaussian log-probability with respect to its mean. Averaging these products gives an unbiased estimate: over repeated sample sets, its mean equals the exact gradient.

Now apply ordinary autodiff through the same sampled action. The chain rule produces

g^path=∂R∂a∂a∂μ=∂R∂a⋅1=0almost surely.\widehat g_{\mathrm{path}}=\frac{\partial R}{\partial a}\frac{\partial a}{\partial\mu}=\frac{\partial R}{\partial a}\cdot 1=0 \quad \text{almost surely}.

The action moves one-for-one with μ\mu, which explains ∂a/∂μ=1\partial a/\partial\mu=1. But the indicator reward is flat everywhere except at its two edges, so ordinary autodiff sees ∂R/∂a=0\partial R/\partial a=0 for every sample that does not land exactly on an edge. A continuous distribution hits either exact edge with probability zero. Therefore the estimator is not merely noisy: it returns exactly zero for every practical sample set, at every μ\mu.

Why is the true gradient nonzero? Because success improves when probability crosses the boundary, not because reward has a local slope inside either flat region. In figure 6, all particles shift by the same amount. The change in success count comes entirely from particles crossing into or out of GG. Pointwise autodiff examines the neighborhood around each sampled particle, so it misses a contribution concentrated exactly at the edges.

In generalized-function notation, the missing derivative is

dda1{a∈[ℓ,r]}=δ(a−ℓ)−δ(a−r),\frac{d}{da}\mathbf{1}\{a\in[\ell,r]\}=\delta(a-\ell)-\delta(a-r),

where the Dirac delta δ\delta represents a unit contribution concentrated at one point. Substituting this generalized derivative into (4) with v=1v=1 recovers ρμ(ℓ)−ρμ(r)\rho_\mu(\ell)-\rho_\mu(r). So the pathwise identity is not mathematically wrong. The problem is that ordinary autodiff differentiates the program’s pointwise operations and does not automatically recover this boundary contribution.

μ (cm)J(μ)exact ∂J/∂μscore, MC mean± s.e.naive pathwisesmoothed, β = 2
−9.00.00970.008580.008360.0002700.01046
−6.00.08740.050870.050880.0005200.05169
−3.00.32170.092640.092080.0004600.08558
0.00.49500.000000.000310.0002500.00092
3.00.3217−0.09264−0.091780.000460−0.08516
9.00.0097−0.00858−0.008740.000260−0.01016
table 2. Four estimates of the same gradient with σ = 3 cm and G = [−2, 2] cm. Monte Carlo values use 256 samples per estimate and 400 repeated estimates. The score-function mean follows the exact derivative, while naive pathwise autodiff stays at zero.

For example, when μ=−3\mu=-3 cm, the robot succeeds about 32%32\% of the time and the exact gradient is 0.092640.09264. The score-function Monte Carlo mean is 0.092080.09208, close to the exact answer. Naive pathwise autodiff still reports 00, while the smoothed reward reports 0.085580.08558—a useful direction, but a biased value for the original objective.

A common workaround is to replace each hard step with a smooth sigmoid:

Rβ(a)=ς ⁣(β(a−ℓ))−ς ⁣(β(a−r)).R_\beta(a)=\varsigma\!\left(\beta(a-\ell)\right)-\varsigma\!\left(\beta(a-r)\right).

Here ς(x)=1/(1+e−x)\varsigma(x)=1/(1+e^{-x}) is the logistic sigmoid, and β\beta controls how sharp the two softened edges are. This surrogate reward has nonzero slopes near ℓ\ell and rr, so autodiff can see a pathwise signal. But it is a different objective. At β=2\beta=2, the resulting gradient is about 8%8\% too small at μ=−3\mu=-3 and about 19%19\% too large at μ=−9\mu=-9; the bias even changes sign. Increasing β\beta makes the surrogate closer to the hard reward, but it also concentrates the useful gradient into narrower regions near the edges, which tends to increase variance.

Gradient against policy mean: an exact positive lobe on the left and negative lobe on the right, score Monte Carlo markers on that curve, a nearby smoothed curve, and a flat zero line for naive autodiff.
figure 7. Across policy means, the score-function Monte Carlo markers follow the exact gradient. The smoothed reward gives a nearby but biased curve. Naive pathwise autodiff remains zero everywhere and therefore misses both lobes.

The exact objective, exact derivative, score-function estimate, and variance calculation fit in a few lines of dependency-free Python. The code below makes the numerical claims reproducible.1

import math

SIGMA, L, R = 3.0, -2.0, 2.0

def phi(x): return math.exp(-0.5 * x * x) / math.sqrt(2.0 * math.pi)
def Phi(x): return 0.5 * (1.0 + math.erf(x / math.sqrt(2.0)))

def J(mu):
    """Success probability in closed form."""
    return Phi((R - mu) / SIGMA) - Phi((L - mu) / SIGMA)

def dJ(mu):
    """Its exact derivative: density at the left edge minus density at the right."""
    return (phi((L - mu) / SIGMA) - phi((R - mu) / SIGMA)) / SIGMA

def score(mu, n, rng):
    """Unbiased at every mu; needs only sampled successes and failures."""
    xs = [mu + SIGMA * rng.gauss(0.0, 1.0) for _ in range(n)]
    return sum((1.0 if L <= a <= R else 0.0) * (a - mu) / SIGMA**2 for a in xs) / n

def score_sd(mu):
    """Exact standard deviation of one score sample, via int u^2 phi(u) du."""
    ul, ur = (L - mu) / SIGMA, (R - mu) / SIGMA
    m2 = (ul * phi(ul) - ur * phi(ur) + Phi(ur) - Phi(ul)) / SIGMA**2
    return math.sqrt(m2 - dJ(mu) ** 2)

# The naive pathwise estimator is a function that sums zeros, so it is not written here:
# dR/da is 0 wherever it exists, and the mean of n zeros is 0 for every n.

What each view costs

The score-function estimator can see a hard boundary, but rare successes make it noisy. If one sample has standard deviation ss and the true gradient has magnitude ∣g∣|g|, then the number of independent samples needed for a standard error equal to 10%10\% of ∣g∣|g| is approximately

n≈(s0.1∣g∣)2.n \approx \left(\frac{s}{0.1|g|}\right)^2.

This equation comes from the usual standard error s/ns/\sqrt{n}. As the policy mean moves away from the cup, success becomes rare, ∣g∣|g| becomes small relative to the estimator’s noise, and the required sample count grows rapidly:

μ (cm)J(μ)exact ∂J/∂μs.d. of one sampleratio to gradientsamples for 10% error
−3.03.2e−019.26e−020.1511.6265
−6.08.7e−025.09e−020.1683.31084
−9.09.7e−038.58e−030.08710.210331
−12.04.3e−045.12e−040.02548.5234806
−15.07.3e−061.11e−050.004369.613657178
table 3. The cost of the score-function estimator as the policy mean moves away from the success interval. The gradient and one-sample standard deviation are exact. An unbiased estimator can still be too noisy to use cheaply.

The two views therefore fail in opposite ways on the same task:

  • The score-function view sees the correct boundary flux without differentiating the reward, but it needs many samples when successes are rare.
  • The pathwise view often provides a low-variance local signal when the computation is smooth, but ordinary autodiff sees no signal at all from this hard boundary.

This does not contradict claim 1. The theorem equates the ideal expectations under its regularity assumptions; it does not promise that every finite-sample implementation is useful, or that ordinary autodiff will construct the generalized derivative of a discontinuity.

Read table 1 with that tradeoff in mind. PPO and GRPO can learn from a 0/10/1 reward, but may pay heavily in samples. A differentiable simulator given the same raw indicator reward can return zero because the gradient it computes contains no boundary term.One practical compromise is to use an accurate simulator for the forward pass and a smoothed surrogate for the backward pass [19].

The main lesson is simple: a policy gradient measures how probability moves between outcomes. You can estimate that movement by watching probability weights change or by following samples through the computation. Hard reward boundaries make the choice visible: the true gradient lives in probability crossing the boundary, and an estimator is useful only if it can see that crossing.

Footnotes

  1. The closed-form standard deviation in score_sd uses ∫u2ϕ(u) du=−uϕ(u)+Φ(u)\int u^2 \phi(u)\,du = -u\phi(u) + \Phi(u); it agrees with a 400,000-sample Monte Carlo estimate to three digits at μ=−3\mu = -3, which is the check I would want before believing the last column of table 3. ↩

references

[1]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
[2]
Z. Shao et al., “DeepSeekMath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024.
[3]
D. Guo et al., “DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning,” Nature, vol. 645, no. 8081, pp. 633–638, 2025, doi: 10.1038/s41586-025-09422-z.
[4]
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th International Conference on Machine Learning, in PMLR, vol. 80. 2018, pp. 1861–1870.
[5]
J. Xu et al., “Accelerated policy learning with parallel differentiable simulation,” in International Conference on Learning Representations, 2022.
[6]
M. Jankowiak and F. Obermeyer, “Pathwise derivatives beyond the reparameterization trick,” in Proceedings of the 35th International Conference on Machine Learning, in PMLR, vol. 80. 2018, pp. 2235–2244.
[7]
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, no. 3–4, pp. 229–256, 1992.
[8]
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems, 2000, pp. 1057–1063.
[9]
S. Mohamed, M. Rosca, M. Figurnov, and A. Mnih, “Monte Carlo gradient estimation in machine learning,” Journal of Machine Learning Research, vol. 21, no. 132, pp. 1–62, 2020.
[10]
J. Schulman, N. Heess, T. Weber, and P. Abbeel, “Gradient estimation using stochastic computation graphs,” in Advances in Neural Information Processing Systems, 2015.
[11]
P. Parmas and M. Sugiyama, “A unified view of likelihood ratio and reparameterization gradients,” in Proceedings of the 24th International Conference on Artificial Intelligence and Statistics, in PMLR, vol. 130. 2021, pp. 4078–4086.
[12]
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning, in PMLR, vol. 32. 2014, pp. 387–395.
[13]
T. P. Lillicrap et al., “Continuous control with deep reinforcement learning,” in International Conference on Learning Representations, 2016.
[14]
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa, “Learning continuous control policies by stochastic value gradients,” in Advances in Neural Information Processing Systems, 2015.
[15]
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with Gumbel-Softmax,” in International Conference on Learning Representations, 2017.
[16]
T. Salimans, J. Ho, X. Chen, S. Sidor, and I. Sutskever, “Evolution strategies as a scalable alternative to reinforcement learning,” arXiv preprint arXiv:1703.03864, 2017.
[17]
H. Mania, A. Guy, and B. Recht, “Simple random search of static linear policies is competitive for reinforcement learning,” in Advances in Neural Information Processing Systems, 2018.
[18]
H. J. T. Suh, M. Simchowitz, K. Zhang, and R. Tedrake, “Do differentiable simulators give better policy gradients?,” in Proceedings of the 39th International Conference on Machine Learning, in PMLR, vol. 162. 2022, pp. 20668–20696.
[19]
Y. Song, S. Kim, and D. Scaramuzza, “Learning quadruped locomotion using differentiable simulation,” in Proceedings of the 8th Conference on Robot Learning, in PMLR, vol. 270. 2024, pp. 258–271.

cite

Found this useful or inspiring? Consider citing it — and follow along on X.

@article{feng2026robottwo,
  title   = {Two policy gradients, one blind spot},
  author  = {Feng, Aosong},
  journal = {asfeng.dev},
  year    = {2026},
  month   = {Aug},
  url     = {https://asfeng.dev/blog/robot-two-gradients-one-flow/}
}

© 2026 Aosong Feng · text licensed CC BY 4.0 — quote and reuse with attribution.