Why VAEs use reparameterisation instead of REINFORCE
A VAE does not need the reparameterisation trick in order to have an encoder gradient. We could train its encoder with REINFORCE instead. The estimate would still be unbiased; it would simply be much noisier.
The first structural difference appears before latent dimension enters the story. With respect to the objective , REINFORCE is zeroth-order: it observes the value at a random sample and combines that value with the sample’s log-probability gradient. Reparameterisation is first-order: it observes the local slope and follows that slope through the sampling map. Even with one latent variable, a function value and a derivative contain different information.
High dimension adds a second problem. REINFORCE broadcasts one scalar value to every latent coordinate, whereas reparameterisation supplies a separate partial derivative to each coordinate. In the factorised -dimensional setting analysed below, this cross-coordinate interference can make matching one pathwise sample require roughly score-function samples. The factor of is an amplification of the information gap, not its origin.
Policy gradients and VAEs share this mathematical shape: their parameters change the distribution that produces the samples. Imagine trying to improve the average exam score in a room when your only knob changes who enters, not any person’s score. The population moves as you turn the knob. Score-function and pathwise estimators are two ways to differentiate that moving average [1].
The reparameterisation trick is therefore not part of the VAE model. It is one way to estimate the model’s gradient. This post removes that estimator, replaces it with REINFORCE, and separates the two things we lose: local slope information already at , and coordinate-local credit when .
One expectation, two gradient estimators
Start with the common objective
Here is the distribution that produces a random sample , controls that distribution, and assigns a value to the sample. For the moment, assume that itself does not contain . We want : the direction in which changing the sampler improves the expected objective.
There are two standard ways to obtain this gradient. The first differentiates the probability of each possible sample. Using gives
This is the score-function estimator; in reinforcement learning it is REINFORCE [2]. Here “score function” names the density derivative , not the objective value . The factor says how good the sampled outcome was; the density score says how to change to make that outcome more likely. Multiplying them turns “this sample was good” into a parameter update.
The important advantage is that we never differentiate . It may be a simulator, a discrete test, or any other black box. We only need to evaluate and differentiate the log-probability of drawing . In optimisation language, the estimator is zeroth-order with respect to —even though it still differentiates the density .
The second method rewrites the random draw. Suppose we can first sample parameter-free noise and then produce through a differentiable map . The noise distribution stays fixed when changes, so ordinary chain-rule differentiation applies:
This is the pathwise estimator. The right-hand side multiplies two quantities: tells us how the objective would change if the sample moved, and tells us how changing the parameters would move that sample. For a Gaussian, ; this special case is the reparameterisation trick used by VAEs.
The pathwise estimator asks for more structure. The map and the objective must both be differentiable. In return, it is first-order with respect to : it gets the local slope instead of only the value .
Under the usual regularity conditions, (1) and (2) are exact expressions for the same gradient. Neither is an approximation to the other. In practice we approximate either expectation with Monte Carlo samples, and those sample estimates can have very different variance.Other gradient estimators exist, including measure-valued derivatives and finite-difference methods, but score-function and pathwise estimators dominate this setting.
That distinction matters because an optimiser never sees the exact expectation. It sees a noisy estimate from one or a few samples and takes a step. So the useful question is not “which estimator has a gradient?” Both do. It is “which estimator gives the lower-variance gradient, and why?”
The gap already exists at d = 1
Consider the smallest possible example. Let with , and let the objective be . There is one latent variable and one parameter. The exact gradient is as simple as it can be:
For REINFORCE, the Gaussian mean score is . At any fixed , the variance-minimising baseline that is constant with respect to is . Treat as a detached control variate when forming the estimator; one sample then gives
The estimate is unbiased, but it fluctuates because it must infer the direction from two random quantities: whether the sampled value was above its baseline, and which side of the Gaussian mean produced it.
The pathwise estimator differentiates the same sample directly:
It receives the exact answer from every sample because this particular objective has a constant local slope. No cross-coordinate credit assignment exists—the entire difference comes from objective-value information versus objective-derivative information.
This example does not prove that every pathwise estimator has lower variance than every score-function estimator. Variance depends on , , the parameterisation and the available control variates. It proves the narrower point we need: setting removes cross-coordinate interference, but it does not turn a zeroth-order observation of into a first-order one.
Applying both estimators to a VAE
A VAE introduces an unobserved latent variable to explain an observed example . Its decoder defines the joint model , while its encoder approximates the posterior over . Training maximises the evidence lower bound, or ELBO:
The expectation is over latent samples drawn from the encoder. Inside the brackets, rewards a latent value that explains the data under the model, while is the encoder-entropy contribution. Their difference is the sampled learning signal; write it as .
The decoder parameters appear only inside this signal, so their derivative passes through the expectation in the usual way. The encoder parameters are harder: they control both the distribution that draws and the term inside the signal. This is the moving-distribution problem from the previous section, with one extra detail.
The direct derivative of averages to zero. Indeed, . Therefore a valid REINFORCE estimator for the encoder is
The baseline is any value that does not depend on the sampled . Subtracting it changes the estimator’s variance but not its expectation, because the expected score is zero. Equation (4) is the precise sense in which a VAE can be trained with REINFORCE.
One encoder serves many inputs. This is amortised inference. Because the typical ELBO can differ greatly from one input to another, one global baseline for the whole dataset is usually too crude. The natural baseline is a function .
A standard VAE also satisfies the pathwise estimator’s requirements. Its Gaussian latent can be written as with fixed noise , and its decoder is differentiable. This makes the VAE a particularly favourable case for reparameterisation.
These conditions are useful, not automatic. Score-function estimators for variational inference came first, and black-box variational inference was built on them [3]. Reparameterised autoencoders did not make (3) differentiable for the first time. They supplied a much lower-variance way to estimate a gradient that already existed [4], [5].
High dimension adds a second penalty
The one-dimensional example isolates the information gap. Now suppose the encoder factorises across latent coordinates, as a diagonal Gaussian encoder does:
Think of as the encoder parameters that control coordinate . Its density score depends only on . But in REINFORCE it is multiplied by the full objective value . Every coordinate therefore receives the same scalar assessment of the entire latent vector.
The pathwise estimator uses instead. Coordinate receives the derivative specific to moving . This creates the additional high-dimensional difference shown in figure 2: REINFORCE broadcasts one global value, while reparameterisation assigns coordinate-local derivatives.
We can make the cost of this additional broadcast exact. Define : fix coordinate , average over all the other coordinates, and keep only the part of the objective value predictable from . The following claim separates the one-coordinate score-function variance from the extra variance contributed by other coordinates. The proof is included for completeness; the paragraph after it gives the plain-language reading.
claim 1 (what the broadcast costs).
Let factorise as above, let be square-integrable, and let be any baseline that does not depend on . Then the variance of coordinate ‘s score-function estimator splits as
The second term is non-negative and is unchanged by every such baseline. It vanishes when is a function of alone; if almost surely, this condition is also necessary.
proof.
Write , where is the residual left after predicting from . By construction, . Split the estimator into , with and .
Because depends only on , conditioning on gives and . The two variances therefore add. Finally,
which proves the split.
∎
The first term in claim 1 is the variance that can remain with one coordinate. When , and the second term is zero, but REINFORCE must still convert the value into a direction by multiplying it by the density score . A baseline can reduce this variance, but it cannot supply the missing derivative .
The second term is the genuinely high-dimensional penalty. It is the variation left after is fixed: in a factorised encoder, that variation comes from the other latent coordinates. It is noise charged to coordinate even though coordinate did not cause it. A baseline that cannot see cannot remove this term.
Now suppose the learning signal is a sum of similarly scaled coordinate contributions. This is exact in the appendix’s orthogonal linear model and is a useful approximation in some larger models. The conditional variance then contains roughly unrelated contributions, so it grows like . The gradient signal itself stays for each coordinate, which makes the signal-to-noise ratio fall like .
Averaging independent Monte Carlo samples improves signal-to-noise by . Recovering a factor of therefore takes about samples. This is where the “one pathwise sample costs roughly REINFORCE samples” rule comes from; it is a scaling statement for this factorised setting, not a universal constant.
The pathwise estimator does not broadcast the scalar . For a coordinatewise sampling map, when , so (2) reduces to . Other coordinates may still affect the value of when the decoder couples them, but they do not enter as an additive global score. If is separable, they disappear completely.
Reparameterisation uses locality twice: a local slope already at , and coordinate-specific credit when .
What a baseline can and cannot buy
A baseline is a reference value for the objective. Instead of multiplying the density score by the raw learning signal , REINFORCE multiplies it by the centred signal . Since , subtracting leaves the expected gradient unchanged. A good baseline says, in effect, “do not reward a sample merely for being typically good; reward it for being better than expected.”
But centring a function value does not turn it into a derivative. Even the best constant baseline leaves the first term of claim 1 unless the particular objective and parameterisation make that term degenerate. This is the limit already visible at : a baseline can remove an irrelevant offset from , but it cannot reveal the local slope that REINFORCE never evaluated.
This centring removes a large but avoidable source of variance. If the global objective has a mean of size , failing to subtract it creates variance. An optimal scalar baseline removes that offset, but the other coordinates still fluctuate, leaving the cross-coordinate contribution at in the quadratic Gaussian example.1
A coordinate-specific baseline can go further. Let denote every latent coordinate except . Under the factorised encoder, a baseline is independent of , so and the estimator remains unbiased. Choosing removes the other coordinates’ additive contribution when is separable and can reduce it more generally. Rao–Blackwellised and local-expectation estimators use this idea [3], [6], but they typically require work for each coordinate. The trade is therefore about times as many samples or about times as much per-coordinate computation. Reparameterisation obtains the local derivative in one backward pass.
Amortisation makes the target move. One encoder serves inputs whose typical ELBO values may be very different, and the baseline accuracy required to expose a small gradient can tighten as the encoder approaches the posterior. A single global baseline can therefore look adequate early and become useless later. Learned input-dependent baselines are a requirement for a practical score-function VAE rather than a minor refinement [7]; the appendix gives the corresponding numbers in the solvable model.
Two VAE-specific control variates
Beyond a generic baseline, the VAE objective offers two useful pieces of estimator structure.
Keep the full learning signal. VAEs often compute (3) as a reconstruction term minus an analytic . For a score-function estimator, applying REINFORCE only to reconstruction may increase variance. At the exact posterior, the full learning signal equals the constant ; the term is acting as a control variate. Splitting it away keeps the noisy part and discards its cancellation.
Drop the zero-mean pathwise term. The standard pathwise loss differentiates both through the sampled and directly through . The direct component has zero expectation, so dropping it remains unbiased. This is the “sticking the landing” (STL) estimator [8]. Its variance becomes exactly zero at the exact posterior, whereas the standard pathwise estimator retains a variance floor. Near convergence, the meaningful comparison is therefore against this corrected pathwise estimator, not only the textbook version.
What to remember
Yes, a VAE can be trained with REINFORCE. The estimator in (4) is unbiased; it simply gives the optimiser a noisier update.
The reason begins at . REINFORCE infers a direction from objective values, while reparameterisation reads the local derivative. High dimension then adds a second cost: one global value is broadcast across coordinates, so each coordinate inherits fluctuations caused by the others. In the solvable factorised model, matching one pathwise sample eventually costs roughly score-function samples.
Score-function estimators remain essential when the latent is discrete, the objective is a black box, or the sampling process cannot be differentiated. In those settings, baselines and stronger control variates are the price of working with less local information [7], [9], [10], [11].
The companion question is what happens when reinforcement learning does satisfy the pathwise requirements. First-order objective information should help even before dimension enters, and coordinate-local credit should add another advantage in high-dimensional action spaces. In practice neither gain is as automatic as it sounds, which is where the other post begins.
Appendix: the model behind the numbers
To put constants behind the scaling argument, we need a VAE whose gradient variance can be computed exactly. A linear Gaussian decoder provides one: probabilistic principal component analysis with a Gaussian encoder [12], [13]. The model is
Here is the latent dimension, is the data dimension, and the encoder parameters are . Choose decoder columns satisfying . This orthogonality makes the exact posterior diagonal and isotropic: with precision . The encoder family therefore contains the exact posterior; the calculation can follow training all the way to a known optimum.
Write the encoder’s mean error as and reparameterise a sample as , where . The sampled learning signal separates into the constant log evidence plus independent coordinate terms:
In (5), each summand depends only on one standard-normal variable . That is the separable case of claim 1, and it reduces every variance calculation to Gaussian moments of order at most eight. The figure and table use , , , and 64 inputs drawn from the model; except when latent dimension is swept.
The three curves in figure 3 show the predicted scalings. Pathwise signal-to-noise stays approximately constant as grows. A baselined score estimator loses a factor of because of the conditional-variance term. Without a baseline it loses roughly another factor of because the unremoved mean offset also grows. The final column of table 1 squares the pathwise-to-score signal-to-noise ratio to convert it into the number of score samples needed for equal precision.
| score, no | score, | score, optimal | pathwise | pathwise + STL | samples needed | |
|---|---|---|---|---|---|---|
| 1 | 0.0469 | 0.262 | 0.330 | 0.649 | 0.693 | 3.9 |
| 8 | 0.0333 | 0.183 | 0.199 | 0.664 | 0.708 | 11.1 |
| 32 | 0.0170 | 0.109 | 0.112 | 0.665 | 0.710 | 35.2 |
| 128 | 0.0058 | 0.057 | 0.058 | 0.669 | 0.714 | 133.8 |
| 512 | 0.0016 | 0.029 | 0.029 | 0.670 | 0.715 | 527.3 |
What the numbers add
The row is the control that separates the two stories. With no other coordinate available to create assignment noise, the optimally baselined score estimator still needs samples per pathwise sample at this parameter point. That residual gap is model-dependent, not a universal constant, but it confirms that the estimator difference begins before dimension does.
The remaining rows measure the high-dimensional amplification. For , the required score-function sample counts are . Once is moderate, the exchange rate tracks latent dimension to within a few percent, as predicted by the conditional-variance term in claim 1.
The same calculation separates offset removal from credit assignment. At , the unbaselined score estimator has about 1400 times the pathwise variance. A perfect scalar baseline cuts this to 33 times, but 88% of what remains is conditional variance from the other coordinates. Changing the data dimension from 64 to 3136 leaves the baselined signal-to-noise ratio unchanged in this model: data dimension changes an offset the baseline can remove, while latent dimension changes how many coordinates share one value.
A separate sweep from encoder initialisation toward the exact posterior exposes the amortisation problem. The baseline may initially miss the per-input ELBO by about 75 nats without dominating variance, but the tolerated error falls below one nat near the posterior, while ELBO values vary by about 50 nats across inputs. These constants belong to this model; the qualitative lesson is that the baseline target can become stricter as inference improves.
The script below implements the closed-form variance calculation for one sampled input. It then checks one coordinate against two million direct Monte Carlo draws; the two estimates agree within the relative errors reported below.2
import numpy as np
D, d, alpha, sigma = 784, 32, 1.0, 0.3
beta = 1.0 + alpha / sigma**2 # posterior precision, per coordinate
rw, rx = np.random.default_rng(0), np.random.default_rng(1)
W = np.linalg.qr(rw.standard_normal((D, d)))[0] * np.sqrt(alpha) # so W.T @ W = alpha I
x = rx.standard_normal(d) @ W.T + sigma * rx.standard_normal(D)
mu, s = np.zeros(d), np.ones(d) # the encoder at initialisation
m = mu - (x @ W) / (beta * sigma**2) # mismatch against the exact posterior
p1, p2 = -beta * m * s, 0.5 * (1.0 - beta * s**2) # A = const + sum_i (p1 e_i + p2 e_i^2)
grad = np.sqrt((beta * m) @ (beta * m) + ((1 - beta * s**2) ** 2).sum())
def score_var(delta):
"""Per-coordinate exact variance of the score estimator; delta = b - ELBO(x), in nats.
Returns the mean and scale blocks, each as (signal-carrying term, conditional variance) —
the two halves of the claim. Only the first depends on the baseline.
"""
q0 = -delta - p2
cond = (p1**2 + 2 * p2**2).sum() - (p1**2 + 2 * p2**2) # Var(f | z_i)
return (((q0 + 3 * p2) ** 2 + 2 * p1**2 + 6 * p2**2) / s**2, cond / s**2,
2 * (q0 + 5 * p2) ** 2 + 10 * p1**2 + 24 * p2**2, 2 * cond)
def path_var(stl):
"""Exact variance of the pathwise estimator, summed over every parameter."""
g = (beta * s - 1 / s) ** 2 if stl else beta**2 * s**2
h = 2 * (1 - beta * s**2) ** 2 if stl else 2 * beta**2 * s**4
return float(g.sum() + (beta**2 * m**2 * s**2 + h).sum())
for name, v in [("score, b = 0", sum(t.sum() for t in score_var(480.63))),
("score, perfect b", sum(t.sum() for t in score_var(0.0))),
("pathwise", path_var(False)), ("pathwise + STL", path_var(True))]:
print(f"{name:>18} variance {v:10.4g} SNR {grad / np.sqrt(v):.4f}")
signal, cond, _, _ = score_var(50.0) # the closed form, put in front of samples
eps = np.random.default_rng(7).standard_normal((2_000_000, d))
adv = eps @ p1 + (eps**2) @ p2 - (50.0 + p2.sum())
print(f"\nmean coordinate 0: sampled {(adv * eps[:, 0] / s[0]).var(ddof=1):.6g}, "
f"exact {signal[0] + cond[0]:.6g}")
Footnotes
-
For under at , the split reads , whose first term is the unsubtracted mean and whose last is . At that is , so the celebrated blow-up is 94% a missing subtraction. ↩
-
The mean-parameter block agrees to a relative error of about at two million draws, the scale block to about ; the scale block converges more slowly because its score is rather than , so its own variance depends on eighth moments. ↩
references
cite
Found this useful or inspiring? Consider citing it — and follow along on X.
@article{feng2026vaewith,
title = {Why VAEs use reparameterisation instead of REINFORCE},
author = {Feng, Aosong},
journal = {asfeng.dev},
year = {2026},
month = {Aug},
url = {https://asfeng.dev/blog/vae-with-reinforce/}
}© 2026 Aosong Feng · text licensed CC BY 4.0 — quote and reuse with attribution.