<Callout kind="note">
**The short version.** A stochastic policy does not choose one future; it assigns probabilities to many possible futures. Training changes those probabilities.

There are two common ways to measure that change. Score-function methods such as REINFORCE, PPO, and GRPO watch each future's probability change. Pathwise methods follow the same random sample through the update and differentiate its reward. When the required derivatives are well behaved, both compute the same gradient; [integration by parts connects the two formulas](#integration-by-parts-connects-the-two-views).

The difference becomes important when the reward has a hard success boundary. In the example below, [ordinary pathwise autodiff returns exactly zero](#one-boundary-is-enough-to-fool-autodiff), even though moving the policy clearly changes the chance of success. The score-function estimator still sees the change, but it may need many samples.

You only need basic probability, derivatives, and the idea that a policy can sample an action. Every other symbol is defined where it appears.
</Callout>

This distinction appears across reinforcement learning. PPO and GRPO use the score-function route [@schulman2017ppo; @shao2024deepseekmath; @guo2025deepseekr1]. SAC and differentiable simulation use pathwise derivatives, although they obtain those derivatives from different places [@haarnoja2018sac; @xu2022shac]. The names will make more sense after we solve one small robot problem. First, here is the whole setup in one picture: the input stays fixed, but a stochastic policy can sample different actions.

<Figure src="/assets/posts/robot-two-gradients-one-flow/the-task.png" alt="A left-to-right flow diagram. The same robot observation enters a box labelled policy pi theta, which branches into three gripper closures at different horizontal positions around a fixed cup. Black triangular markers show the three sampled action positions relative to the cup-centre guide." caption="The camera view and task stay the same, but the stochastic policy πθ does not return one guaranteed closing position. Each run samples a new action. The dashed guide marks the cup centre; the three black triangles mark three possible samples." label="fig:the-task" />

## Start with one noisy robot action

Imagine that the robot repeats the same reach many times. The camera image, instruction, cup, and all but one joint stay fixed. The only number we record is where the gripper closes along the horizontal axis. Call this position $a$, measured in centimetres from the cup's centre.

We model the action as

$$
a = \mu_\theta + \sigma \epsilon, \qquad \epsilon \sim \mathcal{N}(0,1).
$$

Here $\mu_\theta$ is the mean position chosen by policy parameters $\theta$, $\sigma$ is the amount of exploration noise, and $\epsilon$ is one draw from a standard normal distribution. In plain language: the policy aims at $\mu_\theta$, but each attempt lands a little to the left or right.

The robot receives one point if the gripper closes inside the success interval $G = [\ell,r]$, and zero otherwise:

$$
R(a) = \mathbf{1}\{a \in G\}.
$$

The symbol $\mathbf{1}\{\cdot\}$ is an indicator: it equals $1$ when the condition is true and $0$ when it is false. Throughout the post, $\sigma = 3$ cm and $G = [-2,2]$ cm. This reward has no gentle slope: crossing either edge changes the reward instantly from $0$ to $1$, or from $1$ to $0$.

<Figure src="/assets/posts/robot-two-gradients-one-flow/hold-or-not.png" alt="Two panels compare nearby grasp positions. On the left, the action marker a lies inside the hatched success interval G, the cup stays upright, and the reward is one. On the right, a lies just outside G, the cup tips, and the reward is zero." caption="The lower strips turn the robot scene into the one-dimensional model: a is the closing position and the hatched band is the success interval G. A small shift moves a across the boundary, so the outcome jumps from R = 1 to R = 0. There is no partial credit at the edge." label="fig:hold-or-not" />

The second figure deliberately tells the same story twice. The robot drawings show the physical consequence; the strips below them show the mathematical abstraction. The next diagram adds the policy's full action distribution on top of that line.

<Figure src="/assets/posts/robot-two-gradients-one-flow/policy-density.png" alt="A Gaussian policy density centred at μ = −3 cm, with the area under the curve inside the hatched success interval G = [−2, 2] cm shaded as the success probability J(μ)." caption="The policy is centred at μ = −3 cm. The hatched band is the fixed success interval G = [−2, 2] cm, and the shaded area under the density is the success probability J(μ). Moving μ changes that area." label="fig:slice" />

## A policy update moves probability

Before deriving either gradient estimator, we need a precise description of what learning changes. Let $\lambda$ run continuously from $0$ to $1$ as the policy moves from its old parameters to its new ones: $\theta(0)=\theta_{\text{old}}$ and $\theta(1)=\theta_{\text{new}}$. This is not the robot's physical time $t$. The index $t$ moves the robot through one rollout; $\lambda$ moves the policy through one update.<Sidenote>In the one-action example, we can simply set $\lambda=\mu$. Increasing $\lambda$ then means sliding the Gaussian policy to the right.</Sidenote>

A rollout $\tau=(s_0,a_0,\ldots,s_T)$ is one complete sequence of states and actions. Let $\rho_\lambda(\tau)$ be the probability density that the current policy and environment assign to that rollout, and let $R(\tau)$ be its total reward. The objective is the expected reward:

$$
J(\lambda) = \int R(\tau)\, \rho_\lambda(\tau) \, d\tau
\label{eq:objective}
$$

Equation [@eq:objective] says: consider every possible rollout, multiply its reward by its probability, and add the results. The integration variable $d\tau$ means that the integral ranges over all possible rollouts. For a fixed task, the scoring rule $R$ stays the same. The policy update changes $\rho_\lambda$, so it changes which rollouts are likely.

We can describe this change as a flow of probability. Let $v_\lambda(\tau)$ be a velocity field: it tells us the direction and speed at which probability near rollout $\tau$ moves as $\lambda$ increases. Conservation of probability gives the continuity equation:

$$
\partial_\lambda \rho_\lambda(\tau) + \nabla_\tau \cdot \big(\rho_\lambda(\tau) v_\lambda(\tau)\big) = 0
\label{eq:continuity}
$$

The first term, $\partial_\lambda\rho_\lambda$, is the rate at which the density changes at a fixed location. The second term measures net probability flowing out of that location: $\rho_\lambda v_\lambda$ is the probability flux, $\nabla_\tau\!\cdot$ is its divergence, and the minus sign says that net outflow lowers the local density. In short, **density changes because probability flows in or out; probability is not created or destroyed.**

The one-dimensional robot example makes [@eq:continuity] easy to check. Set $\lambda=\mu$ and shift the Gaussian mean to the right. Every sampled action moves right at unit speed, so $v=1$. The equation becomes $\partial_\mu\rho_\mu(a)=-\partial_a\rho_\mu(a)$: changing the mean has the same effect as translating the whole density in the opposite coordinate direction.

> Probability in a region falls not because probability was destroyed, but because it went somewhere else.

<Figure src="/assets/posts/robot-two-gradients-one-flow/probability-flux.png" alt="Two Gaussian curves before and after a rightward policy shift cross a fixed hatched interval G, with inflow at the left boundary ℓ and outflow at the right boundary r." caption="The success interval stays fixed while the policy density moves right. Probability enters through ℓ and leaves through r, so the change in success probability is ΔJ = inflow − outflow." label="fig:box" />

<Callout kind="note">
Two technical details matter later. First, [@eq:continuity] describes how the *trajectory distribution changes during a policy update*. It is not a Fokker–Planck equation describing how a physical state distribution changes over time, although both are conservation laws. Second, the velocity field $v_\lambda$ need not be unique: several flows can produce the same density change. That freedom leads to different pathwise estimators with the same expected value [@jankowiak2018pathwise].
</Callout>

## View one: keep the outcomes fixed

The first view keeps each possible rollout $\tau$ fixed and asks how its probability changes. Differentiate [@eq:objective] with respect to $\lambda$. Because the task reward $R$ does not depend on the policy update, the derivative acts only on $\rho_\lambda$:

$$
\frac{dJ}{d\lambda} = \int R\, \partial_\lambda \rho_\lambda \, d\tau = \mathbb{E}_{\tau \sim \rho_\lambda}\big[R(\tau)\, \partial_\lambda \log \rho_\lambda(\tau)\big]
\label{eq:score}
$$

The last expression is the score-function, or likelihood-ratio, gradient. The reward $R(\tau)$ says how good the sampled rollout was. The score $\partial_\lambda\log\rho_\lambda(\tau)$ says how quickly the update changes that rollout's *relative* probability. Their product is the rollout's contribution to the gradient, and the expectation averages that contribution over sampled rollouts.

Why can this estimator work with a black-box environment? The probability of a rollout factors into the initial-state distribution $p_0$, the policy $\pi_\theta$, and the environment dynamics $P$:

$$
\rho_\theta(\tau)=p_0(s_0)\prod_t \pi_\theta(a_t\mid s_t)P(s_{t+1}\mid s_t,a_t),
\qquad
\nabla_\theta\log\rho_\theta(\tau)=\sum_t\nabla_\theta\log\pi_\theta(a_t\mid s_t).
$$

The product on the left assigns a probability to the whole rollout. After taking the logarithm, that product becomes a sum. If the environment does not itself depend on $\theta$, the derivatives of $p_0$ and $P$ are zero, so only derivatives of the policy remain on the right. We need sampled states, actions, and rewards, plus $\nabla_\theta\log\pi_\theta$ from the policy network. We do **not** need derivatives of the environment or reward. This is the key idea behind REINFORCE [@williams1992simple] and the policy gradient theorem [@sutton2000policy].

This route is tolerant of black boxes, but it is not assumption-free. Moving the derivative through the integral must be valid, and the set of possible rollouts—the support of $\rho_\theta$—must behave regularly. If the support itself changes shape with $\theta$, the usual formula needs extra care [@mohamed2020monte].

<Figure src="/assets/posts/robot-two-gradients-one-flow/fixed-outcomes.png" alt="Two aligned rows of six fixed bins. The bins remain in the same positions while their probability weights shift right into the success interval G." caption="Score-function view: the possible outcomes—the bins—stay fixed. A policy update changes only how much probability weight each bin receives; it does not match an old sample to a new one." label="fig:fixed-camera" />

## View two: keep the randomness fixed

The second view matches an outcome under the old policy with an outcome under the new one. Write $\epsilon\sim q$ for a random draw from a fixed noise distribution, and write $\tau=T_\lambda(\epsilon)$ for the rollout produced by passing that draw through the policy and environment. In the robot example, this is simply $a_\lambda=\mu_{\theta(\lambda)}+\sigma\epsilon$.

The important step is to give the old and new policies the **same** $\epsilon$. Think of $\epsilon$ as a numbered lottery ticket. If both policies receive the same ticket, we can say how that particular sampled outcome moved. If they receive independent tickets, we can compare two distributions, but we cannot match one old sample to one new sample.

With shared noise, $\partial T_\lambda(\epsilon)/\partial\lambda$ is the velocity of the sampled rollout. If the reward is differentiable along this path, the chain rule gives the pathwise gradient:

$$
\frac{dJ}{d\lambda} = \mathbb{E}_{\epsilon \sim q}\left[\nabla_\tau R\big(T_\lambda(\epsilon)\big)^{\!\top} \frac{\partial T_\lambda(\epsilon)}{\partial \lambda}\right]
\label{eq:pathwise}
$$

The vector $\nabla_\tau R$ is the local reward slope: it points toward nearby rollouts with higher reward. The vector $\partial T_\lambda/\partial\lambda$ says how the sampled rollout moves as the policy changes. Their dot product measures how quickly the reward changes along that motion, and the expectation averages over noise draws.

For a multi-step rollout, we can write $a_t=g_\theta(s_t,\epsilon_t)$ for the policy and $s_{t+1}=f(s_t,a_t,\xi_t)$ for the environment. Here $g_\theta$ maps a state and policy noise to an action; $f$ maps the current state, action, and environment noise $\xi_t$ to the next state. Once every $\epsilon_t$ and $\xi_t$ is held fixed, the rollout becomes an ordinary computation graph, $\theta\rightarrow a_0\rightarrow s_1\rightarrow\cdots\rightarrow R$, and backpropagation can differentiate it. Stochastic computation graphs formalize how score-function and pathwise estimators can be mixed at different random nodes [@schulman2015gradient].

<Figure src="/assets/posts/robot-two-gradients-one-flow/fixed-noise.png" alt="Eight action samples before and after an update are connected one to one by equal rightward arrows. One bold arrow carries a sample across the left boundary into G." caption="Pathwise view: keeping the same noise ε pairs every old sample with its new location. Each action moves by the same amount; the bold arrow shows a sample crossing into G. The reward changes at that crossing, not along the flat regions on either side." label="fig:moving-camera" />

## Integration by parts connects the two views

The two formulas look unrelated: [@eq:score] differentiates probability, while [@eq:pathwise] differentiates reward. The continuity equation shows that they describe the same flow. Substitute $\partial_\lambda\rho_\lambda=-\nabla_\tau\cdot(\rho_\lambda v_\lambda)$ into the score-function derivation, then use integration by parts to move the derivative from the probability flux onto the reward:

$$
\underbrace{\mathbb{E}\big[R\, \partial_\lambda \log \rho\big]}_{\text{watch the weights change}}
\;=\;
\underbrace{-\int R\, \nabla \cdot (\rho v)}_{\text{weights change because probability flows}}
\;=\;
\underbrace{\mathbb{E}\big[v^{\!\top} \nabla R\big]}_{\text{follow probability toward higher reward}}
$$

The left expression is the score-function view: keep locations fixed and watch their probability weights change. The middle expression replaces that weight change with probability flow. The right expression is the pathwise view: follow the moving probability and measure the reward slope along its path. Integration by parts is the operation that moves the derivative from $\rho$ to $R$.

<Theorem kind="claim" label="thm:cameras" title="two views of one gradient">Suppose [@eq:continuity] holds, the reward is differentiable along the flow, and the boundary term at infinity is zero. Then the score-function estimator [@eq:score] and the pathwise estimator [@eq:pathwise] have the same expectation. They are two representations of the same change in the rollout distribution.</Theorem>

This transport interpretation of pathwise derivatives is developed by Jankowiak and Obermeyer [@jankowiak2018pathwise]. Parmas and Sugiyama give a unified treatment of likelihood-ratio and reparameterization gradients as estimators of the same probability movement [@parmas2021unified].

Three caveats are important. First, equal expectations do not imply equal finite-sample variance. Second, the velocity field is not unique, so different pathwise estimators can have the same mean. Third, [@thm:cameras] assumes that the ordinary reward gradient $\nabla_\tau R$ exists. Our hard $0/1$ reward breaks that last assumption.

## Where common algorithms fit

We can now sort familiar algorithms by one practical question: **what derivative can the algorithm obtain outside the policy itself?** The “view” column follows from that requirement.

<Table caption="The derivative each algorithm family needs outside the policy, and what a hard zero-or-one reward means for it. The last column refers to the robot example in this post." label="tab:methods">

| family | view | derivative needed outside the policy | hard $R=\mathbf{1}\{a\in G\}$ |
|---|---|---|---|
| REINFORCE, PPO, GRPO [@williams1992simple; @schulman2017ppo; @shao2024deepseekmath] | fixed outcomes | none; only sampled returns | unbiased, but can be noisy |
| DDPG, TD3, SAC [@silver2014deterministic; @lillicrap2016continuous; @haarnoja2018sac] | fixed noise | $\nabla_a Q(s,a)$ from a learned critic | depends on the critic's learned smoothness |
| SVG(1), SVG($\infty$) [@heess2015svg] | fixed noise | derivatives through a learned dynamics model | depends on the learned model and reward |
| differentiable simulation [@xu2022shac] | fixed noise | dynamics and reward derivatives from the simulator | naive autodiff returns zero in this example |
| Gumbel-Softmax, straight-through [@jang2017gumbel] | fixed noise | a differentiable relaxation of a discrete choice | produces a biased surrogate gradient |

</Table>

The actor-critic row is easy to misunderstand. DDPG and SAC do not differentiate the real environment. They differentiate a learned critic $Q(s,a)$ with respect to the action. The smoothness requirement therefore falls on a neural network the algorithm controls, not directly on the world's reward or dynamics. By contrast, methods that backpropagate through a learned dynamics model or a differentiable simulator need a longer differentiable path. Evolution strategies and random search sit outside both columns: they estimate changes in performance without forming $\nabla_\theta\log\pi_\theta$ at all [@salimans2017evolution; @mania2018random].

## One boundary is enough to fool autodiff

Return to the robot's actual $0/1$ reward and set the update coordinate $\lambda$ equal to the policy mean $\mu$. The objective $J(\mu)$ is now simply the probability that a Gaussian action lands between $\ell$ and $r$. Let $\Phi$ denote the standard normal cumulative distribution function and $\phi$ its probability density. Then both the objective and its exact derivative have closed forms:

$$
J(\mu) = \Phi\!\left(\frac{r - \mu}{\sigma}\right) - \Phi\!\left(\frac{\ell - \mu}{\sigma}\right),
\qquad
\frac{\partial J}{\partial \mu} = \frac{1}{\sigma}\left[\phi\!\left(\frac{\ell - \mu}{\sigma}\right) - \phi\!\left(\frac{r - \mu}{\sigma}\right)\right].
$$

The formula for $J$ subtracts the Gaussian probability to the left of $\ell$ from the probability to the left of $r$, leaving exactly the probability inside the success interval. The derivative has an equally concrete meaning: $\phi((\ell-\mu)/\sigma)/\sigma$ is probability density entering through the left edge as the policy moves right, while $\phi((r-\mu)/\sigma)/\sigma$ is density leaving through the right edge. Their difference is **inflow minus outflow**.

The same conclusion follows directly from the continuity equation. Integrating [@eq:continuity] over $G=[\ell,r]$ gives

$$
\frac{dJ}{d\lambda}=\rho_\lambda(\ell)v_\lambda(\ell)-\rho_\lambda(r)v_\lambda(r).
$$

The two terms are the probability flux through the interval's left and right edges. In our example $v=1$, so the gradient is $\rho_\mu(\ell)-\rho_\mu(r)$, matching the formula above and [@fig:box].<Sidenote>This boundary is not the boundary term at infinity in [@thm:cameras]. It appears because we integrate over the finite success interval $G$.</Sidenote>

The score-function view has no problem with this boundary because it never differentiates the reward. For one sampled action, its gradient estimate is

$$
\widehat g_{\mathrm{score}}=R(a)\frac{a-\mu}{\sigma^2}.
$$

The reward $R(a)$ says whether the sample succeeded. The factor $(a-\mu)/\sigma^2$ is the derivative of the Gaussian log-probability with respect to its mean. Averaging these products gives an unbiased estimate: over repeated sample sets, its mean equals the exact gradient.

Now apply ordinary autodiff through the same sampled action. The chain rule produces

$$
\widehat g_{\mathrm{path}}=\frac{\partial R}{\partial a}\frac{\partial a}{\partial\mu}=\frac{\partial R}{\partial a}\cdot 1=0 \quad \text{almost surely}.
$$

The action moves one-for-one with $\mu$, which explains $\partial a/\partial\mu=1$. But the indicator reward is flat everywhere except at its two edges, so ordinary autodiff sees $\partial R/\partial a=0$ for every sample that does not land exactly on an edge. A continuous distribution hits either exact edge with probability zero. Therefore the estimator is not merely noisy: it returns **exactly zero for every practical sample set, at every $\mu$**.

Why is the true gradient nonzero? Because success improves when probability crosses the boundary, not because reward has a local slope inside either flat region. In [@fig:moving-camera], all particles shift by the same amount. The change in success count comes entirely from particles crossing into or out of $G$. Pointwise autodiff examines the neighborhood around each sampled particle, so it misses a contribution concentrated exactly at the edges.

In generalized-function notation, the missing derivative is

$$
\frac{d}{da}\mathbf{1}\{a\in[\ell,r]\}=\delta(a-\ell)-\delta(a-r),
$$

where the Dirac delta $\delta$ represents a unit contribution concentrated at one point. Substituting this generalized derivative into [@eq:pathwise] with $v=1$ recovers $\rho_\mu(\ell)-\rho_\mu(r)$. So the pathwise identity is not mathematically wrong. The problem is that ordinary autodiff differentiates the program's pointwise operations and does not automatically recover this boundary contribution.

<Callout kind="warning">
Saying only “the reward is not differentiable” hides the important distinction. The expected objective $J(\mu)$ is smooth because Gaussian averaging smooths the two jumps. The sampled reward $R(a)$ is discontinuous at $a=\ell$ and $a=r$. The function we want to optimize can therefore have a useful derivative even when the pointwise computation passed to autodiff does not reveal it.
</Callout>

<Table caption="Four estimates of the same gradient with σ = 3 cm and G = [−2, 2] cm. Monte Carlo values use 256 samples per estimate and 400 repeated estimates. The score-function mean follows the exact derivative, while naive pathwise autodiff stays at zero." label="tab:estimators">

| μ (cm) | J(μ) | exact ∂J/∂μ | score, MC mean | ± s.e. | naive pathwise | smoothed, β = 2 |
|---|---|---|---|---|---|---|
| −9.0 | 0.0097 | 0.00858 | 0.00836 | 0.00027 | 0 | 0.01046 |
| −6.0 | 0.0874 | 0.05087 | 0.05088 | 0.00052 | 0 | 0.05169 |
| −3.0 | 0.3217 | 0.09264 | 0.09208 | 0.00046 | 0 | 0.08558 |
| 0.0 | 0.4950 | 0.00000 | 0.00031 | 0.00025 | 0 | 0.00092 |
| 3.0 | 0.3217 | −0.09264 | −0.09178 | 0.00046 | 0 | −0.08516 |
| 9.0 | 0.0097 | −0.00858 | −0.00874 | 0.00026 | 0 | −0.01016 |

</Table>

For example, when $\mu=-3$ cm, the robot succeeds about $32\%$ of the time and the exact gradient is $0.09264$. The score-function Monte Carlo mean is $0.09208$, close to the exact answer. Naive pathwise autodiff still reports $0$, while the smoothed reward reports $0.08558$—a useful direction, but a biased value for the original objective.

A common workaround is to replace each hard step with a smooth sigmoid:

$$
R_\beta(a)=\varsigma\!\left(\beta(a-\ell)\right)-\varsigma\!\left(\beta(a-r)\right).
$$

Here $\varsigma(x)=1/(1+e^{-x})$ is the logistic sigmoid, and $\beta$ controls how sharp the two softened edges are. This surrogate reward has nonzero slopes near $\ell$ and $r$, so autodiff can see a pathwise signal. But it is a different objective. At $\beta=2$, the resulting gradient is about $8\%$ too small at $\mu=-3$ and about $19\%$ too large at $\mu=-9$; the bias even changes sign. Increasing $\beta$ makes the surrogate closer to the hard reward, but it also concentrates the useful gradient into narrower regions near the edges, which tends to increase variance.

<Figure src="/assets/posts/robot-two-gradients-one-flow/gradient-field.png" alt="Gradient against policy mean: an exact positive lobe on the left and negative lobe on the right, score Monte Carlo markers on that curve, a nearby smoothed curve, and a flat zero line for naive autodiff." caption="Across policy means, the score-function Monte Carlo markers follow the exact gradient. The smoothed reward gives a nearby but biased curve. Naive pathwise autodiff remains zero everywhere and therefore misses both lobes." label="fig:gradient-field" />

<Callout kind="warning">
Reward smoothing changes more than the gradient estimate: it changes the objective itself. That can move the optimum and change which behavior the policy prefers. Suh and colleagues analyze when differentiable simulators yield better policy gradients and propose an estimator that interpolates between score-function and pathwise information [@suh2022differentiable].
</Callout>

The exact objective, exact derivative, score-function estimate, and variance calculation fit in a few lines of dependency-free Python. The code below makes the numerical claims reproducible.[^1]

```python title="boundary.py"
import math

SIGMA, L, R = 3.0, -2.0, 2.0

def phi(x): return math.exp(-0.5 * x * x) / math.sqrt(2.0 * math.pi)
def Phi(x): return 0.5 * (1.0 + math.erf(x / math.sqrt(2.0)))

def J(mu):
    """Success probability in closed form."""
    return Phi((R - mu) / SIGMA) - Phi((L - mu) / SIGMA)

def dJ(mu):
    """Its exact derivative: density at the left edge minus density at the right."""
    return (phi((L - mu) / SIGMA) - phi((R - mu) / SIGMA)) / SIGMA

def score(mu, n, rng):
    """Unbiased at every mu; needs only sampled successes and failures."""
    xs = [mu + SIGMA * rng.gauss(0.0, 1.0) for _ in range(n)]
    return sum((1.0 if L <= a <= R else 0.0) * (a - mu) / SIGMA**2 for a in xs) / n

def score_sd(mu):
    """Exact standard deviation of one score sample, via int u^2 phi(u) du."""
    ul, ur = (L - mu) / SIGMA, (R - mu) / SIGMA
    m2 = (ul * phi(ul) - ur * phi(ur) + Phi(ur) - Phi(ul)) / SIGMA**2
    return math.sqrt(m2 - dJ(mu) ** 2)

# The naive pathwise estimator is a function that sums zeros, so it is not written here:
# dR/da is 0 wherever it exists, and the mean of n zeros is 0 for every n.
```

## What each view costs

The score-function estimator can see a hard boundary, but rare successes make it noisy. If one sample has standard deviation $s$ and the true gradient has magnitude $|g|$, then the number of independent samples needed for a standard error equal to $10\%$ of $|g|$ is approximately

$$
n \approx \left(\frac{s}{0.1|g|}\right)^2.
$$

This equation comes from the usual standard error $s/\sqrt{n}$. As the policy mean moves away from the cup, success becomes rare, $|g|$ becomes small relative to the estimator's noise, and the required sample count grows rapidly:

<Table caption="The cost of the score-function estimator as the policy mean moves away from the success interval. The gradient and one-sample standard deviation are exact. An unbiased estimator can still be too noisy to use cheaply." label="tab:variance">

| μ (cm) | J(μ) | exact ∂J/∂μ | s.d. of one sample | ratio to gradient | samples for 10% error |
|---|---|---|---|---|---|
| −3.0 | 3.2e−01 | 9.26e−02 | 0.151 | 1.6 | 265 |
| −6.0 | 8.7e−02 | 5.09e−02 | 0.168 | 3.3 | 1084 |
| −9.0 | 9.7e−03 | 8.58e−03 | 0.087 | 10.2 | 10331 |
| −12.0 | 4.3e−04 | 5.12e−04 | 0.025 | 48.5 | 234806 |
| −15.0 | 7.3e−06 | 1.11e−05 | 0.004 | 369.6 | 13657178 |

</Table>

The two views therefore fail in opposite ways on the same task:

- The score-function view sees the correct boundary flux without differentiating the reward, but it needs many samples when successes are rare.
- The pathwise view often provides a low-variance local signal when the computation is smooth, but ordinary autodiff sees no signal at all from this hard boundary.

This does not contradict [@thm:cameras]. The theorem equates the ideal expectations under its regularity assumptions; it does not promise that every finite-sample implementation is useful, or that ordinary autodiff will construct the generalized derivative of a discontinuity.

Read [@tab:methods] with that tradeoff in mind. PPO and GRPO can learn from a $0/1$ reward, but may pay heavily in samples. A differentiable simulator given the same raw indicator reward can return zero because the gradient it computes contains no boundary term.<Sidenote>One practical compromise is to use an accurate simulator for the forward pass and a smoothed surrogate for the backward pass [@song2024quadruped].</Sidenote>

**The main lesson is simple: a policy gradient measures how probability moves between outcomes. You can estimate that movement by watching probability weights change or by following samples through the computation. Hard reward boundaries make the choice visible: the true gradient lives in probability crossing the boundary, and an estimator is useful only if it can see that crossing.**

[^1]: The closed-form standard deviation in `score_sd` uses $\int u^2 \phi(u)\,du = -u\phi(u) + \Phi(u)$; it agrees with a 400,000-sample Monte Carlo estimate to three digits at $\mu = -3$, which is the check I would want before believing the last column of [@tab:variance].