Aosong Feng 冯傲松

the world, rebuilt small enough to run

← ../writing/

Why VAEs use reparameterisation instead of REINFORCE

A VAE does not need the reparameterisation trick in order to have an encoder gradient. We could train its encoder with REINFORCE instead. The estimate would still be unbiased; it would simply be much noisier.

The first structural difference appears before latent dimension enters the story. With respect to the objective ff, REINFORCE is zeroth-order: it observes the value f(z)f(z) at a random sample and combines that value with the sample’s log-probability gradient. Reparameterisation is first-order: it observes the local slope ∇zf(z)\nabla_z f(z) and follows that slope through the sampling map. Even with one latent variable, a function value and a derivative contain different information.

High dimension adds a second problem. REINFORCE broadcasts one scalar value to every latent coordinate, whereas reparameterisation supplies a separate partial derivative to each coordinate. In the factorised dd-dimensional setting analysed below, this cross-coordinate interference can make matching one pathwise sample require roughly dd score-function samples. The factor of dd is an amplification of the information gap, not its origin.

Policy gradients and VAEs share this mathematical shape: their parameters change the distribution that produces the samples. Imagine trying to improve the average exam score in a room when your only knob changes who enters, not any person’s score. The population moves as you turn the knob. Score-function and pathwise estimators are two ways to differentiate that moving average [1].

The reparameterisation trick is therefore not part of the VAE model. It is one way to estimate the model’s gradient. This post removes that estimator, replaces it with REINFORCE, and separates the two things we lose: local slope information already at d=1d=1, and coordinate-local credit when d>1d>1.

One expectation, two gradient estimators

Start with the common objective

F(ϕ)=Ez∼qϕ[f(z)].F(\phi) = \mathbb{E}_{z \sim q_\phi}[f(z)].

Here qϕq_\phi is the distribution that produces a random sample zz, ϕ\phi controls that distribution, and f(z)f(z) assigns a value to the sample. For the moment, assume that ff itself does not contain ϕ\phi. We want ∇ϕF\nabla_\phi F: the direction in which changing the sampler improves the expected objective.

There are two standard ways to obtain this gradient. The first differentiates the probability of each possible sample. Using ∇ϕqϕ=qϕ∇ϕlog⁡qϕ\nabla_\phi q_\phi = q_\phi \nabla_\phi \log q_\phi gives

∇ϕ Eqϕ[f(z)]  =  ∫f(z) ∇ϕqϕ(z) dz  =  Eqϕ[f(z) ∇ϕlog⁡qϕ(z)].(1)\nabla_\phi \, \mathbb{E}_{q_\phi}\bigl[f(z)\bigr] \;=\; \int f(z) \, \nabla_\phi q_\phi(z) \, dz \;=\; \mathbb{E}_{q_\phi}\bigl[f(z) \, \nabla_\phi \log q_\phi(z)\bigr]. \tag{1}

This is the score-function estimator; in reinforcement learning it is REINFORCE [2]. Here “score function” names the density derivative ∇ϕlog⁡qϕ(z)\nabla_\phi \log q_\phi(z), not the objective value f(z)f(z). The factor f(z)f(z) says how good the sampled outcome was; the density score says how to change ϕ\phi to make that outcome more likely. Multiplying them turns “this sample was good” into a parameter update.

The important advantage is that we never differentiate ff. It may be a simulator, a discrete test, or any other black box. We only need to evaluate f(z)f(z) and differentiate the log-probability of drawing zz. In optimisation language, the estimator is zeroth-order with respect to ff—even though it still differentiates the density qϕq_\phi.

The second method rewrites the random draw. Suppose we can first sample parameter-free noise ε∼q0\varepsilon \sim q_0 and then produce zz through a differentiable map z=Tϕ(ε)z = T_\phi(\varepsilon). The noise distribution stays fixed when ϕ\phi changes, so ordinary chain-rule differentiation applies:

∇ϕ Eqϕ[f(z)]  =  ∇ϕ Eε∼q0[f(Tϕ(ε))]  =  Eε∼q0[∇zf(Tϕ(ε)) ⁣⊤∂Tϕ(ε)∂ϕ].(2)\nabla_\phi \, \mathbb{E}_{q_\phi}\bigl[f(z)\bigr] \;=\; \nabla_\phi \, \mathbb{E}_{\varepsilon \sim q_0}\bigl[f(T_\phi(\varepsilon))\bigr] \;=\; \mathbb{E}_{\varepsilon \sim q_0}\Bigl[\nabla_z f\bigl(T_\phi(\varepsilon)\bigr)^{\!\top} \frac{\partial T_\phi(\varepsilon)}{\partial \phi}\Bigr]. \tag{2}

This is the pathwise estimator. The right-hand side multiplies two quantities: ∇zf\nabla_z f tells us how the objective would change if the sample moved, and ∂Tϕ/∂ϕ\partial T_\phi / \partial \phi tells us how changing the parameters would move that sample. For a Gaussian, Tϕ(ε)=μϕ+σϕ⊙εT_\phi(\varepsilon) = \mu_\phi + \sigma_\phi \odot \varepsilon; this special case is the reparameterisation trick used by VAEs.

The pathwise estimator asks for more structure. The map TϕT_\phi and the objective ff must both be differentiable. In return, it is first-order with respect to ff: it gets the local slope ∇zf\nabla_z f instead of only the value f(z)f(z).

Under the usual regularity conditions, (1) and (2) are exact expressions for the same gradient. Neither is an approximation to the other. In practice we approximate either expectation with Monte Carlo samples, and those sample estimates can have very different variance.Other gradient estimators exist, including measure-valued derivatives and finite-difference methods, but score-function and pathwise estimators dominate this setting.

That distinction matters because an optimiser never sees the exact expectation. It sees a noisy estimate from one or a few samples and takes a step. So the useful question is not “which estimator has a gradient?” Both do. It is “which estimator gives the lower-variance gradient, and why?”

The gap already exists at d = 1

Consider the smallest possible example. Let z=μ+εz=\mu+\varepsilon with ε∼N(0,1)\varepsilon\sim\mathcal{N}(0,1), and let the objective be f(z)=zf(z)=z. There is one latent variable and one parameter. The exact gradient is as simple as it can be:

∂∂μE[f(z)]=∂∂μE[μ+ε]=1.\frac{\partial}{\partial\mu}\mathbb{E}[f(z)] = \frac{\partial}{\partial\mu}\mathbb{E}[\mu+\varepsilon] = 1.

For REINFORCE, the Gaussian mean score is ∂μlog⁡qμ(z)=z−μ=ε\partial_\mu\log q_\mu(z)=z-\mu=\varepsilon. At any fixed μ\mu, the variance-minimising baseline that is constant with respect to zz is b=E[f(z)]=μb=\mathbb{E}[f(z)]=\mu. Treat bb as a detached control variate when forming the estimator; one sample then gives

gscore=(f(z)−b) ∂μlog⁡qμ(z)=ε2,E[gscore]=1,Var⁡(gscore)=2.g_{\mathrm{score}}=(f(z)-b)\,\partial_\mu\log q_\mu(z)=\varepsilon^2,\qquad \mathbb{E}[g_{\mathrm{score}}]=1,\qquad \operatorname{Var}(g_{\mathrm{score}})=2.

The estimate is unbiased, but it fluctuates because it must infer the direction from two random quantities: whether the sampled value was above its baseline, and which side of the Gaussian mean produced it.

The pathwise estimator differentiates the same sample directly:

gpath=f′(z)∂z∂μ=1⋅1=1,Var⁡(gpath)=0.g_{\mathrm{path}}=f'(z)\frac{\partial z}{\partial\mu}=1\cdot1=1,\qquad \operatorname{Var}(g_{\mathrm{path}})=0.

It receives the exact answer from every sample because this particular objective has a constant local slope. No cross-coordinate credit assignment exists—the entire difference comes from objective-value information versus objective-derivative information.

Zeroth-order objective information against first-order objective informationIn one dimension, the score estimate remains random while the pathwise estimate is exact for a linear objective.z = μ + ε, ε ~ N(0, 1), f(z) = z, true gradient = 1score function · zeroth-order in ff(z) − b = ε∂μ log q = ε×g = ε²correct on average · variance 2pathwise · first-order in f∂f/∂z = 1∂z/∂μ = 1×g = 1exact on every sample · variance 0
figure 1. Even with one latent coordinate, the estimators see different information. Here the optimally baselined score estimator returns the random value ε²; pathwise returns the exact gradient 1.

This example does not prove that every pathwise estimator has lower variance than every score-function estimator. Variance depends on ff, qϕq_\phi, the parameterisation and the available control variates. It proves the narrower point we need: setting d=1d=1 removes cross-coordinate interference, but it does not turn a zeroth-order observation of ff into a first-order one.

Applying both estimators to a VAE

A VAE introduces an unobserved latent variable zz to explain an observed example xx. Its decoder defines the joint model pθ(x,z)=p(z)pθ(x∣z)p_\theta(x,z) = p(z)p_\theta(x \mid z), while its encoder qϕ(z∣x)q_\phi(z \mid x) approximates the posterior over zz. Training maximises the evidence lower bound, or ELBO:

L(θ,ϕ;x)  =  Ez∼qϕ(z∣x)[log⁡pθ(x,z)−log⁡qϕ(z∣x)].(3)\mathcal{L}(\theta, \phi; x) \;=\; \mathbb{E}_{z \sim q_\phi(z \mid x)}\bigl[\log p_\theta(x, z) - \log q_\phi(z \mid x)\bigr]. \tag{3}

The expectation is over latent samples drawn from the encoder. Inside the brackets, log⁡pθ(x,z)\log p_\theta(x,z) rewards a latent value that explains the data under the model, while −log⁡qϕ(z∣x)-\log q_\phi(z \mid x) is the encoder-entropy contribution. Their difference is the sampled learning signal; write it as Aθ,ϕ(x,z)A_{\theta,\phi}(x,z).

The decoder parameters θ\theta appear only inside this signal, so their derivative passes through the expectation in the usual way. The encoder parameters ϕ\phi are harder: they control both the distribution that draws zz and the −log⁡qϕ-\log q_\phi term inside the signal. This is the moving-distribution problem from the previous section, with one extra detail.

The direct derivative of −log⁡qϕ-\log q_\phi averages to zero. Indeed, Eq[∇ϕ(−log⁡qϕ)]=−∫∇ϕqϕ dz=−∇ϕ1=0\mathbb{E}_q[\nabla_\phi(-\log q_\phi)] = -\int \nabla_\phi q_\phi \, dz = -\nabla_\phi 1 = 0. Therefore a valid REINFORCE estimator for the encoder is

∇ϕL=Ez∼qϕ(z∣x)[(Aθ,ϕ(x,z)−b(x))∇ϕlog⁡qϕ(z∣x)].(4)\nabla_\phi \mathcal{L} = \mathbb{E}_{z \sim q_\phi(z \mid x)}\Bigl[(A_{\theta,\phi}(x,z)-b(x))\nabla_\phi \log q_\phi(z \mid x)\Bigr]. \tag{4}

The baseline b(x)b(x) is any value that does not depend on the sampled zz. Subtracting it changes the estimator’s variance but not its expectation, because the expected score Eq[∇ϕlog⁡qϕ]\mathbb{E}_q[\nabla_\phi \log q_\phi] is zero. Equation (4) is the precise sense in which a VAE can be trained with REINFORCE.

One encoder serves many inputs. This is amortised inference. Because the typical ELBO can differ greatly from one input xx to another, one global baseline for the whole dataset is usually too crude. The natural baseline is a function b(x)b(x).

A standard VAE also satisfies the pathwise estimator’s requirements. Its Gaussian latent can be written as z=μϕ(x)+σϕ(x)⊙εz = \mu_\phi(x) + \sigma_\phi(x) \odot \varepsilon with fixed noise ε∼N(0,I)\varepsilon \sim \mathcal{N}(0,I), and its decoder is differentiable. This makes the VAE a particularly favourable case for reparameterisation.

These conditions are useful, not automatic. Score-function estimators for variational inference came first, and black-box variational inference was built on them [3]. Reparameterised autoencoders did not make (3) differentiable for the first time. They supplied a much lower-variance way to estimate a gradient that already existed [4], [5].

High dimension adds a second penalty

The one-dimensional example isolates the information gap. Now suppose the encoder factorises across dd latent coordinates, as a diagonal Gaussian encoder does:

qϕ(z)=∏i=1dqi(zi;ϕi).q_\phi(z) = \prod_{i=1}^{d}q_i(z_i;\phi_i).

Think of ϕi\phi_i as the encoder parameters that control coordinate ziz_i. Its density score si=∂ϕilog⁡qi(zi;ϕi)s_i = \partial_{\phi_i}\log q_i(z_i;\phi_i) depends only on ziz_i. But in REINFORCE it is multiplied by the full objective value f(z1,…,zd)f(z_1,\ldots,z_d). Every coordinate therefore receives the same scalar assessment of the entire latent vector.

The pathwise estimator uses ∂f/∂zi\partial f/\partial z_i instead. Coordinate ii receives the derivative specific to moving ziz_i. This creates the additional high-dimensional difference shown in figure 2: REINFORCE broadcasts one global value, while reparameterisation assigns coordinate-local derivatives.

The extra cross-coordinate cost in more than one dimensionOne global objective value is broadcast across coordinates, while pathwise derivatives remain coordinate-specific.score function · global valuepathwise · local derivativesA(z) − bone number, broadcast d timesthe d parameters of q∂f/∂z₁∂f/∂z₂∂f/∂z₃∂f/∂z₄∂f/∂z₅d derivatives, one eachthe d parameters of q
figure 2. The extra $d>1$ credit-assignment cost. One objective value is broadcast to every coordinate; pathwise supplies one partial derivative per coordinate.

We can make the cost of this additional broadcast exact. Define fˉ(zi)=E[f(z)∣zi]\bar f(z_i)=\mathbb{E}[f(z)\mid z_i]: fix coordinate ziz_i, average over all the other coordinates, and keep only the part of the objective value predictable from ziz_i. The following claim separates the one-coordinate score-function variance from the extra variance contributed by other coordinates. The proof is included for completeness; the paragraph after it gives the plain-language reading.

claim 1 (what the broadcast costs).

Let qϕq_\phi factorise as above, let ff be square-integrable, and let bb be any baseline that does not depend on zz. Then the variance of coordinate ii‘s score-function estimator splits as

Var⁡[(f−b)si]=Var⁡[(fˉ−b)si]+E ⁣[si2Var⁡(f∣zi)].\operatorname{Var}[(f-b)s_i] = \operatorname{Var}[(\bar f-b)s_i] + \mathbb{E}\!\left[s_i^2\operatorname{Var}(f\mid z_i)\right].

The second term is non-negative and is unchanged by every such baseline. It vanishes when ff is a function of ziz_i alone; if si2>0s_i^2>0 almost surely, this condition is also necessary.

proof.

Write f=fˉ+rf=\bar f+r, where rr is the residual left after predicting ff from ziz_i. By construction, E[r∣zi]=0\mathbb{E}[r\mid z_i]=0. Split the estimator into X+YX+Y, with X=(fˉ−b)siX=(\bar f-b)s_i and Y=rsiY=rs_i.

Because sis_i depends only on ziz_i, conditioning on ziz_i gives E[Y]=0\mathbb{E}[Y]=0 and Cov⁡(X,Y)=0\operatorname{Cov}(X,Y)=0. The two variances therefore add. Finally,

Var⁡(Y)=E[si2r2]=E ⁣[si2E[r2∣zi]]=E ⁣[si2Var⁡(f∣zi)],\operatorname{Var}(Y)=\mathbb{E}[s_i^2r^2]=\mathbb{E}\!\left[s_i^2\mathbb{E}[r^2\mid z_i]\right]=\mathbb{E}\!\left[s_i^2\operatorname{Var}(f\mid z_i)\right],

which proves the split.

∎

The first term in claim 1 is the variance that can remain with one coordinate. When d=1d=1, fˉ=f\bar f=f and the second term is zero, but REINFORCE must still convert the value f(z)f(z) into a direction by multiplying it by the density score sis_i. A baseline can reduce this variance, but it cannot supply the missing derivative ∂f/∂zi\partial f/\partial z_i.

The second term is the genuinely high-dimensional penalty. It is the variation left after ziz_i is fixed: in a factorised encoder, that variation comes from the other latent coordinates. It is noise charged to coordinate ii even though coordinate ii did not cause it. A baseline that cannot see zz cannot remove this term.

Now suppose the learning signal is a sum of dd similarly scaled coordinate contributions. This is exact in the appendix’s orthogonal linear model and is a useful approximation in some larger models. The conditional variance Var⁡(f∣zi)\operatorname{Var}(f\mid z_i) then contains roughly d−1d-1 unrelated contributions, so it grows like dd. The gradient signal itself stays O(1)O(1) for each coordinate, which makes the signal-to-noise ratio fall like d−1/2d^{-1/2}.

Averaging NN independent Monte Carlo samples improves signal-to-noise by N\sqrt{N}. Recovering a factor of d\sqrt d therefore takes about N=dN=d samples. This is where the “one pathwise sample costs roughly dd REINFORCE samples” rule comes from; it is a scaling statement for this factorised setting, not a universal constant.

The pathwise estimator does not broadcast the scalar f(z)f(z). For a coordinatewise sampling map, ∂zk/∂ϕi=0\partial z_k/\partial\phi_i=0 when k≠ik\ne i, so (2) reduces to ∂zif(z) ∂ϕiTi\partial_{z_i}f(z)\,\partial_{\phi_i}T_i. Other coordinates may still affect the value of ∂zif\partial_{z_i}f when the decoder couples them, but they do not enter as an additive global score. If ff is separable, they disappear completely.

Reparameterisation uses locality twice: a local slope already at d=1d=1, and coordinate-specific credit when d>1d>1.

What a baseline can and cannot buy

A baseline is a reference value for the objective. Instead of multiplying the density score sis_i by the raw learning signal ff, REINFORCE multiplies it by the centred signal f−bf-b. Since E[si]=0\mathbb{E}[s_i]=0, subtracting bb leaves the expected gradient unchanged. A good baseline says, in effect, “do not reward a sample merely for being typically good; reward it for being better than expected.”

But centring a function value does not turn it into a derivative. Even the best constant baseline leaves the first term of claim 1 unless the particular objective and parameterisation make that term degenerate. This is the limit already visible at d=1d=1: a baseline can remove an irrelevant offset from f(z)f(z), but it cannot reveal the local slope ∇zf(z)\nabla_z f(z) that REINFORCE never evaluated.

This centring removes a large but avoidable source of variance. If the global objective has a mean of size O(d)O(d), failing to subtract it creates O(d2)O(d^2) variance. An optimal scalar baseline removes that offset, but the other coordinates still fluctuate, leaving the cross-coordinate contribution at O(d)O(d) in the quadratic Gaussian example.1

A coordinate-specific baseline can go further. Let z−iz_{-i} denote every latent coordinate except ziz_i. Under the factorised encoder, a baseline bi(x,z−i)b_i(x,z_{-i}) is independent of sis_i, so E[bisi]=0\mathbb{E}[b_i s_i]=0 and the estimator remains unbiased. Choosing bi≈E[f∣z−i]b_i\approx\mathbb{E}[f\mid z_{-i}] removes the other coordinates’ additive contribution when ff is separable and can reduce it more generally. Rao–Blackwellised and local-expectation estimators use this idea [3], [6], but they typically require work for each coordinate. The trade is therefore about dd times as many samples or about dd times as much per-coordinate computation. Reparameterisation obtains the local derivative in one backward pass.

Amortisation makes the target move. One encoder serves inputs whose typical ELBO values may be very different, and the baseline accuracy required to expose a small gradient can tighten as the encoder approaches the posterior. A single global baseline can therefore look adequate early and become useless later. Learned input-dependent baselines are a requirement for a practical score-function VAE rather than a minor refinement [7]; the appendix gives the corresponding numbers in the solvable model.

Two VAE-specific control variates

Beyond a generic baseline, the VAE objective offers two useful pieces of estimator structure.

Keep the full learning signal. VAEs often compute (3) as a reconstruction term minus an analytic KL(qϕ∥p)\mathrm{KL}(q_\phi\Vert p). For a score-function estimator, applying REINFORCE only to reconstruction may increase variance. At the exact posterior, the full learning signal Aθ,ϕ(x,z)A_{\theta,\phi}(x,z) equals the constant log⁡pθ(x)\log p_\theta(x); the −log⁡qϕ-\log q_\phi term is acting as a control variate. Splitting it away keeps the noisy part and discards its cancellation.

Drop the zero-mean pathwise term. The standard pathwise loss differentiates −log⁡qϕ-\log q_\phi both through the sampled zz and directly through ϕ\phi. The direct component has zero expectation, so dropping it remains unbiased. This is the “sticking the landing” (STL) estimator [8]. Its variance becomes exactly zero at the exact posterior, whereas the standard pathwise estimator retains a variance floor. Near convergence, the meaningful comparison is therefore against this corrected pathwise estimator, not only the textbook version.

What to remember

Yes, a VAE can be trained with REINFORCE. The estimator in (4) is unbiased; it simply gives the optimiser a noisier update.

The reason begins at d=1d=1. REINFORCE infers a direction from objective values, while reparameterisation reads the local derivative. High dimension then adds a second cost: one global value is broadcast across coordinates, so each coordinate inherits fluctuations caused by the others. In the solvable factorised model, matching one pathwise sample eventually costs roughly dd score-function samples.

Score-function estimators remain essential when the latent is discrete, the objective is a black box, or the sampling process cannot be differentiated. In those settings, baselines and stronger control variates are the price of working with less local information [7], [9], [10], [11].

The companion question is what happens when reinforcement learning does satisfy the pathwise requirements. First-order objective information should help even before dimension enters, and coordinate-local credit should add another advantage in high-dimensional action spaces. In practice neither gain is as automatic as it sounds, which is where the other post begins.

Appendix: the model behind the numbers

To put constants behind the scaling argument, we need a VAE whose gradient variance can be computed exactly. A linear Gaussian decoder provides one: probabilistic principal component analysis with a Gaussian encoder [12], [13]. The model is

p(z)=N(0,Id),pθ(x∣z)=N(Wz+c,σ2ID),qϕ(z∣x)=N ⁣(μ,diag⁡(s2)).p(z)=\mathcal{N}(0,I_d),\qquad p_\theta(x\mid z)=\mathcal{N}(Wz+c,\sigma^2I_D),\qquad q_\phi(z\mid x)=\mathcal{N}\!\left(\mu,\operatorname{diag}(s^2)\right).

Here dd is the latent dimension, DD is the data dimension, and the encoder parameters are ϕ=(μ,log⁡s)\phi=(\mu,\log s). Choose decoder columns satisfying W⊤W=αIdW^\top W=\alpha I_d. This orthogonality makes the exact posterior diagonal and isotropic: p(z∣x)=N(m⋆,β−1Id)p(z\mid x)=\mathcal{N}(m^\star,\beta^{-1}I_d) with precision β=1+α/σ2\beta=1+\alpha/\sigma^2. The encoder family therefore contains the exact posterior; the calculation can follow training all the way to a known optimum.

Write the encoder’s mean error as m=μ−m⋆m=\mu-m^\star and reparameterise a sample as z=μ+s⊙εz=\mu+s\odot\varepsilon, where ε∼N(0,Id)\varepsilon\sim\mathcal{N}(0,I_d). The sampled learning signal separates into the constant log evidence log⁡pθ(x)\log p_\theta(x) plus independent coordinate terms:

A(ε)=log⁡pθ(x)+d2log⁡β+∑i=1d[−β2(mi+siεi)2+12εi2+log⁡si],(5)A(\varepsilon) = \log p_\theta(x) + \tfrac{d}{2}\log\beta + \sum_{i=1}^{d} \Bigl[-\tfrac{\beta}{2}\bigl(m_i + s_i\varepsilon_i\bigr)^2 + \tfrac{1}{2}\varepsilon_i^2 + \log s_i\Bigr], \tag{5}

In (5), each summand depends only on one standard-normal variable εi\varepsilon_i. That is the separable case of claim 1, and it reduces every variance calculation to Gaussian moments of order at most eight. The figure and table use D=784D=784, α=1\alpha=1, σ=0.3\sigma=0.3, and 64 inputs drawn from the model; d=32d=32 except when latent dimension is swept.

Gradient signal-to-noise ratio against latent dimensionThree curves on logarithmic axes. The pathwise curve is horizontal at about 0.67 from latent dimension one to five hundred and twelve. The baselined score-function curve falls from about 0.26 to 0.029, parallel to a reference line of slope minus one half. The unbaselined score-function curve falls from about 0.047 to 0.0016, steepening toward a reference line of slope minus one.10.10.010.001141664256512latent dimension dgradient signal-to-noise ratiopathwisescore function, with a baselinescore function, no baseline
figure 3. Exact gradient signal-to-noise ratio against latent dimension at initialisation, averaged over 64 inputs. Three power laws: the pathwise estimator is flat, a baselined score estimator loses half a power of the latent dimension, an unbaselined one a full power. The dotted lines are slopes −1/2 and −1.

The three curves in figure 3 show the predicted scalings. Pathwise signal-to-noise stays approximately constant as dd grows. A baselined score estimator loses a factor of d1/2d^{1/2} because of the conditional-variance term. Without a baseline it loses roughly another factor of d1/2d^{1/2} because the unremoved mean offset also grows. The final column of table 1 squares the pathwise-to-score signal-to-noise ratio to convert it into the number of score samples needed for equal precision.

ddscore, no bbscore, b(x)b(x)score, optimal b⋆b^\starpathwisepathwise + STLsamples needed
10.04690.2620.3300.6490.6933.9
80.03330.1830.1990.6640.70811.1
320.01700.1090.1120.6650.71035.2
1280.00580.0570.0580.6690.714133.8
5120.00160.0290.0290.6700.715527.3
table 1. Exact gradient signal-to-noise ratio at initialisation, and the number of latent samples an optimally-baselined score estimator needs per pathwise sample to match it. The d = 1 row retains the zeroth-order information gap; the growth with d adds cross-coordinate credit noise.

What the numbers add

The d=1d=1 row is the control that separates the two stories. With no other coordinate available to create assignment noise, the optimally baselined score estimator still needs 3.93.9 samples per pathwise sample at this parameter point. That residual gap is model-dependent, not a universal constant, but it confirms that the estimator difference begins before dimension does.

The remaining rows measure the high-dimensional amplification. For d=8,32,128,512d=8,32,128,512, the required score-function sample counts are 11.1,35.2,133.8,527.311.1,35.2,133.8,527.3. Once dd is moderate, the exchange rate tracks latent dimension to within a few percent, as predicted by the conditional-variance term in claim 1.

The same calculation separates offset removal from credit assignment. At d=32d=32, the unbaselined score estimator has about 1400 times the pathwise variance. A perfect scalar baseline cuts this to 33 times, but 88% of what remains is conditional variance from the other coordinates. Changing the data dimension from 64 to 3136 leaves the baselined signal-to-noise ratio unchanged in this model: data dimension changes an offset the baseline can remove, while latent dimension changes how many coordinates share one value.

A separate sweep from encoder initialisation toward the exact posterior exposes the amortisation problem. The baseline may initially miss the per-input ELBO by about 75 nats without dominating variance, but the tolerated error falls below one nat near the posterior, while ELBO values vary by about 50 nats across inputs. These constants belong to this model; the qualitative lesson is that the baseline target can become stricter as inference improves.

The script below implements the closed-form variance calculation for one sampled input. It then checks one coordinate against two million direct Monte Carlo draws; the two estimates agree within the relative errors reported below.2

import numpy as np

D, d, alpha, sigma = 784, 32, 1.0, 0.3
beta = 1.0 + alpha / sigma**2                        # posterior precision, per coordinate
rw, rx = np.random.default_rng(0), np.random.default_rng(1)
W = np.linalg.qr(rw.standard_normal((D, d)))[0] * np.sqrt(alpha)    # so W.T @ W = alpha I
x = rx.standard_normal(d) @ W.T + sigma * rx.standard_normal(D)

mu, s = np.zeros(d), np.ones(d)                      # the encoder at initialisation
m = mu - (x @ W) / (beta * sigma**2)                 # mismatch against the exact posterior
p1, p2 = -beta * m * s, 0.5 * (1.0 - beta * s**2)    # A = const + sum_i (p1 e_i + p2 e_i^2)
grad = np.sqrt((beta * m) @ (beta * m) + ((1 - beta * s**2) ** 2).sum())

def score_var(delta):
    """Per-coordinate exact variance of the score estimator; delta = b - ELBO(x), in nats.

    Returns the mean and scale blocks, each as (signal-carrying term, conditional variance) —
    the two halves of the claim. Only the first depends on the baseline.
    """
    q0 = -delta - p2
    cond = (p1**2 + 2 * p2**2).sum() - (p1**2 + 2 * p2**2)         # Var(f | z_i)
    return (((q0 + 3 * p2) ** 2 + 2 * p1**2 + 6 * p2**2) / s**2, cond / s**2,
            2 * (q0 + 5 * p2) ** 2 + 10 * p1**2 + 24 * p2**2, 2 * cond)

def path_var(stl):
    """Exact variance of the pathwise estimator, summed over every parameter."""
    g = (beta * s - 1 / s) ** 2 if stl else beta**2 * s**2
    h = 2 * (1 - beta * s**2) ** 2 if stl else 2 * beta**2 * s**4
    return float(g.sum() + (beta**2 * m**2 * s**2 + h).sum())

for name, v in [("score, b = 0", sum(t.sum() for t in score_var(480.63))),
                ("score, perfect b", sum(t.sum() for t in score_var(0.0))),
                ("pathwise", path_var(False)), ("pathwise + STL", path_var(True))]:
    print(f"{name:>18}  variance {v:10.4g}   SNR {grad / np.sqrt(v):.4f}")

signal, cond, _, _ = score_var(50.0)                 # the closed form, put in front of samples
eps = np.random.default_rng(7).standard_normal((2_000_000, d))
adv = eps @ p1 + (eps**2) @ p2 - (50.0 + p2.sum())
print(f"\nmean coordinate 0: sampled {(adv * eps[:, 0] / s[0]).var(ddof=1):.6g}, "
      f"exact {signal[0] + cond[0]:.6g}")

Footnotes

  1. For f(z)=−12∥z−a∥2f(z) = -\tfrac{1}{2}\lVert z - a\rVert^2 under N(μ,Id)\mathcal{N}(\mu, I_d) at μ=a\mu = a, the split reads ((d+2)/2)2+3/2+(d−1)/2\bigl((d+2)/2\bigr)^2 + 3/2 + (d-1)/2, whose first term is the unsubtracted mean and whose last is E[si2Var⁡(f∣zi)]\mathbb{E}[s_i^2\operatorname{Var}(f \mid z_i)]. At d=32d = 32 that is 289+1.5+15.5=306289 + 1.5 + 15.5 = 306, so the celebrated blow-up is 94% a missing subtraction. ↩

  2. The mean-parameter block agrees to a relative error of about 2×10−52 \times 10^{-5} at two million draws, the scale block to about 2×10−32 \times 10^{-3}; the scale block converges more slowly because its score is εi2−1\varepsilon_i^2 - 1 rather than εi\varepsilon_i, so its own variance depends on eighth moments. ↩

references

[1]
S. Mohamed, M. Rosca, M. Figurnov, and A. Mnih, “Monte Carlo gradient estimation in machine learning,” Journal of Machine Learning Research, vol. 21, no. 132, pp. 1–62, 2020.
[2]
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, no. 3–4, pp. 229–256, 1992.
[3]
R. Ranganath, S. Gerrish, and D. M. Blei, “Black box variational inference,” in Proceedings of the 17th International Conference on Artificial Intelligence and Statistics, in PMLR, vol. 33. 2014, pp. 814–822.
[4]
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” in Proceedings of the 2nd International Conference on Learning Representations, 2014.
[5]
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in Proceedings of the 31st International Conference on Machine Learning, in PMLR, vol. 32. 2014, pp. 1278–1286.
[6]
M. K. Titsias and M. Lázaro-Gredilla, “Local expectation gradients for black box variational inference,” in Advances in Neural Information Processing Systems, 2015.
[7]
A. Mnih and K. Gregor, “Neural variational inference and learning in belief networks,” in Proceedings of the 31st International Conference on Machine Learning, in PMLR, vol. 32. 2014, pp. 1791–1799.
[8]
G. Roeder, Y. Wu, and D. Duvenaud, “Sticking the landing: simple, lower-variance gradient estimators for variational inference,” in Advances in Neural Information Processing Systems, 2017.
[9]
A. Mnih and D. J. Rezende, “Variational inference for Monte Carlo objectives,” in Proceedings of the 33rd International Conference on Machine Learning, in PMLR, vol. 48. 2016, pp. 2188–2196.
[10]
G. Tucker, A. Mnih, C. J. Maddison, D. Lawson, and J. Sohl-Dickstein, “REBAR: low-variance, unbiased gradient estimates for discrete latent variable models,” in Advances in Neural Information Processing Systems, 2017.
[11]
W. Grathwohl, D. Choi, Y. Wu, G. Roeder, and D. Duvenaud, “Backpropagation through the void: optimizing control variates for black-box gradient estimation,” in Proceedings of the 6th International Conference on Learning Representations, 2018.
[12]
M. E. Tipping and C. M. Bishop, “Probabilistic principal component analysis,” Journal of the Royal Statistical Society: Series B, vol. 61, no. 3, pp. 611–622, 1999.
[13]
J. Lucas, G. Tucker, R. B. Grosse, and M. Norouzi, “Don’t blame the ELBO! A linear VAE perspective on posterior collapse,” in Advances in Neural Information Processing Systems, 2019.

cite

Found this useful or inspiring? Consider citing it — and follow along on X.

@article{feng2026vaewith,
  title   = {Why VAEs use reparameterisation instead of REINFORCE},
  author  = {Feng, Aosong},
  journal = {asfeng.dev},
  year    = {2026},
  month   = {Aug},
  url     = {https://asfeng.dev/blog/vae-with-reinforce/}
}

© 2026 Aosong Feng · text licensed CC BY 4.0 — quote and reuse with attribution.