Aosong Feng 冯傲松

the world, rebuilt small enough to run

← ../writing/

The Ideal RL Problem Isn't RL

A five-axis framework for understanding why different RL algorithms exist—and why language-model training is limited more by reward than by the environment.

·18 min·rlreward-designllm

cite, share, .md

This map will help us answer three puzzles from the LLM era:

  • Group Relative Policy Optimization (GRPO) removed the value network that Proximal Policy Optimization (PPO) usually relies on, yet training still worked [1]. Why?
  • LLM training often keeps the policy close to two different reference points. The constraints look similar, but they solve different problems. Why do we need both?
  • Reinforcement learning with verifiable rewards (RLVR) replaced a learned reward model with a checker, such as a test harness. Why did that apparently small change matter so much [2]?

This is not a survey of RL algorithms. It is a way to see why whole families of algorithms exist. We will begin with an ideal world, then remove one useful piece of information at a time. Each loss forces us to pay a different price: more variance, more data, more modeling, or more distrust of the reward.

The equations provide the technical spine, but you can skim them without losing the main argument. A policy is the decision-maker. A trajectory is one sequence of states and actions, and its return is the total reward it earns.

The goal never changes. We still want to maximize

J(θ)=Eτ∼pθ[R(τ)],J(\theta)=\mathbb{E}_{\tau\sim p_\theta}[R(\tau)],

the average return RR of trajectories τ\tau sampled from a policy with parameters θ\theta. What changes is what we are allowed to know, observe, and differentiate. By the end, the three puzzles above will be simple consequences of where LLM training sits on the map.

The ideal world is a computing problem

Start with a world in which we know everything we could reasonably ask for. In that world, there is no learning problem left. There is only a computation.

Imagine controlling a robot arm inside a perfect simulator. We know the physics, every movement has a clear cost, and small changes to the motors produce smooth, predictable changes in the arm. We can try as many movements as we like. This is the ideal corner of our map.

Our ideal world has five properties. The letters will label the five axes of the map:

  • D — Differentiability. The environment and reward are smooth enough to differentiate.
  • K — Dynamics knowledge. We know exactly how the environment changes after an action.
  • F — Feedback density. Every step tells us how well it went.
  • B — Interaction budget. We can collect as much fresh experience as we need.
  • R — Reward fidelity. The reward exactly expresses the goal we care about.

When all five properties hold, we can write the entire process explicitly. The policy chooses an action, the known dynamics ff move the world to its next state, and rewards add up over time:

at=πθ(st),st+1=f(st,at),J(θ)=∑t=0Tγ t r(st,at)(1)a_t=\pi_\theta(s_t),\qquad s_{t+1}=f(s_t,a_t),\qquad J(\theta)=\sum_{t=0}^{T}\gamma^{\,t}\,r(s_t,a_t) \tag{1}

Every arrow is a function we know and can differentiate. The whole rollout is one computation graph: θ\theta affects the first action, that action affects the next state, and so on until the final reward. The chain rule gives us the exact gradient ∇θJ\nabla_\theta J. In other words, we backpropagate through the world.

This is trajectory optimization, a standard problem in optimal control. Notice what we do not need: exploration, a value function, or sampled estimates of the gradient. Nothing must be learned because nothing is unknown.1

That is the first key idea: the ideal RL problem is not RL. RL is what remains when we can no longer solve the problem this directly.

The ideal world is a point; RL is everything that grows outward from it.

Every radar chart below follows the same rule: the center means more information and an easier optimization problem; moving outward means that condition has degraded. The axes are qualitative, not measurements on a shared numerical scale. Their purpose is to compare which limitations dominate in different settings.

The empty map: a five-axis radar with the ideal world as a center pointRadar frame with spokes D, R, F, B, K and a filled dot at the center.D · DifferentiabilityR · Reward fidelityF · Feedback densityB · Interaction budgetK · Dynamics knowledgeideal worldcenter = more informed · outward = more degraded
The five-axis map. The center is ideal; moving outward means a condition gets worse. The scale is conceptual rather than numerical.

The whole argument can be previewed in one table. Each row starts with a lost capability and ends with the family of methods that compensates for it:

axiswhat becomes unavailableproblem createdtypical response
D — Differentiabilityderivatives through the environmentexact backpropagation failsscore-function gradients such as REINFORCE
K — Dynamics knowledgea usable model of what happens nextcounterfactual futures are hard to evaluatesearch or learned world models
F — Feedback densityfrequent signals about progresscredit for a final result is hard to assigncritics, GAE, process rewards, or group baselines
B — Interaction budgetunlimited fresh experiencedata becomes scarce or stalePPO clipping, off-policy correction, or offline RL
R — Reward fidelitya reward that exactly matches the real goalthe policy can exploit an imperfect proxyreference anchors, preference learning, or verifiers
The five axes as a reader's guide: what becomes unavailable, what problem that creates, and which methods respond.

This ideal case suggests a useful first question for any new problem: Can I simply backpropagate through it? The rest of the post is about what to do when the answer is no.2

Only your policy must be differentiable

Now remove differentiability (D). We can still run the world as often as we like, but we cannot differentiate through it. Perhaps the actions are discrete, the reward contains hard branches, or the simulator is a compiled program we cannot inspect.

The previous strategy no longer works. But we can make one important shift: instead of differentiating through the world, differentiate the probability of visiting different trajectories. The log-derivative trick gives us

∇θJ=∇θ ⁣∫pθ(τ) R(τ) dτ=∫pθ(τ) ∇θlog⁡pθ(τ) R(τ) dτ=Eτ∼pθ ⁣[R(τ) ∇θlog⁡pθ(τ)](2)\nabla_\theta J=\nabla_\theta\!\int p_\theta(\tau)\,R(\tau)\,d\tau=\int p_\theta(\tau)\,\nabla_\theta\log p_\theta(\tau)\,R(\tau)\,d\tau=\mathbb{E}_{\tau\sim p_\theta}\!\big[R(\tau)\,\nabla_\theta\log p_\theta(\tau)\big] \tag{2}

In plain English: sample a trajectory, see how much reward it earns, and make rewarded trajectories more likely. Here R(τ)R(\tau) scores the trajectory, while ∇θlog⁡pθ(τ)\nabla_\theta\log p_\theta(\tau) measures how a parameter update would change its probability.Moving ∇θ\nabla_\theta through the integral requires regularity assumptions. They are harmless here; Mohamed and colleagues survey the exceptions [3].

Why does this work even when the environment is a black box? A trajectory’s probability splits into three pieces: where it starts, what actions the policy chooses, and how the environment responds.

log⁡pθ(τ)=log⁡p(s0)+∑tlog⁡πθ(at∣st)+∑tlog⁡p(st+1∣st,at)(3)\log p_\theta(\tau)=\log p(s_0)+\sum_t \log \pi_\theta(a_t\mid s_t)+\sum_t \log p(s_{t+1}\mid s_t,a_t) \tag{3}

Only the policy depends on θ\theta. When we take the derivative, the initial-state and environment terms disappear, leaving ∑t∇θlog⁡πθ(at∣st)\sum_t\nabla_\theta\log\pi_\theta(a_t\mid s_t).

This is the surprising part: the environment disappears from the gradient formula. Its uncertainty returns as noise in our estimate. This result underlies REINFORCE and the policy-gradient theorem [4], [5]. The environment can be jagged, discrete, or hidden. Only the policy—the part we built—must be differentiable.

For a slower derivation with a robot and concrete numbers, see Two Gradients, One Flow and the companion VAE post.

There is another route: a pathwise gradient holds the random noise fixed and differentiates the reward along the sampled path. It usually has lower variance because it uses local slope information, but it requires derivatives from the environment.

The score-function estimator uses only sampled rewards, so it works with black boxes. Its higher variance is the price of that generality [6], [7]. Much of modern RL is an attempt to reduce this variance without giving up the black-box advantage.

The first improvement is free: subtract a baseline bb from every return. This does not change the expected gradient because the expected policy score is zero—the derivative of “all probabilities sum to one” is also zero:

Ea∼πθ(⋅∣s)[∇θlog⁡πθ(a∣s)]=∇θ ⁣∫πθ(a∣s) da=∇θ1=0(4)\mathbb{E}_{a\sim\pi_\theta(\cdot\mid s)}\big[\nabla_\theta\log\pi_\theta(a\mid s)\big]=\nabla_\theta\!\int \pi_\theta(a\mid s)\,da=\nabla_\theta 1=0 \tag{4}

A baseline changes the variance of the estimate, but not its average. A state-dependent baseline b(st)b(s_t) works too. This simple observation leads to several familiar techniques: assign an action only the rewards that follow it, learn a value function VV as the baseline, and update on the advantage A=Q−VA=Q-V—how much better an action was than expected.

We will meet three more variance-reduction tools later: temporal credit assignment, PPO’s clipping rule, and GRPO’s group baseline. They look related, but each responds to a different missing piece of information.

This is exactly what makes policy gradients useful for LLMs. Tokens are discrete, reward models can be treated as black boxes, and sampling itself is not differentiable. None of that matters because log⁡πθ\log\pi_\theta comes from our own neural network. Policy gradients became a universal adapter: they ask almost nothing of the environment.

Atari is the classical example [8]. Its emulator can be run but not differentiated through. It also teaches a broader lesson: real problems usually lose several ideal conditions at once.

The map with Atari pinnedOne cool-colored polygon labeled Atari on the five-axis radar.D · DifferentiabilityR · Reward fidelityF · Feedback densityB · Interaction budgetK · Dynamics knowledgeAtaricenter = more informed · outward = more degraded
Atari is far from ideal on both differentiability (D) and dynamics knowledge (K): we can run the emulator, but we cannot differentiate through it or use it as a known model. Real problems often degrade on several axes at once.

Perfect knowledge, zero gradients

Differentiability (D) and dynamics knowledge (K) are different. We can know exactly how a world works without being able to differentiate through it.

Knowledge comes in levels. At the best end, we have a closed-form equation. One step down, we have an executable simulator: we can ask “what if?” but cannot use calculus. Below that, we can learn an approximate model f^\hat f, gaining the ability to simulate at the cost of model errors. At the worst end, we know nothing beyond the trajectories we have already observed.

When a simulator is known but not differentiable, search can replace gradients. The rules of Go are exact, yet there is no useful derivative of a move. Monte Carlo tree search (MCTS) instead explores many possible futures. AlphaZero combines that search with policy and value networks, gradually compressing expensive lookahead into learned intuition [9].

AlphaZero is not ideal on every other axis: its reward is only a win or loss at the end of a game. That means it has perfect knowledge of the rules but very sparse feedback. Real problems occupy combinations of axes, not neat stops on a single line.

Single-turn LLM generation lands at an unusual extreme. Its “dynamics” are just f=concatf=\mathrm{concat}: append the next token to the existing prefix. This operation is deterministic, known, and cheap. No one needs to learn a world model for ordinary text generation because the world model is string concatenation.

This also explains why tree search can reappear in LLM reasoning. A learned scorer evaluates partial solutions while search explores possible next steps [10]. It is the AlphaZero pattern replayed in a much simpler environment.

The map with Atari and AlphaZero pinnedTwo cool-colored polygons, one solid and one dashed, on the five-axis radar.D · DifferentiabilityR · Reward fidelityF · Feedback densityB · Interaction budgetK · Dynamics knowledgeAtariAlphaZerocenter = more informed · outward = more degraded
Atari (solid) and AlphaZero (dashed) are nearly mirror images. Atari has unknown dynamics and frequent scores; AlphaZero knows the rules exactly but sees only the final win or loss.

When reward comes last, credit becomes inference

Now reduce feedback density (F). Suppose the only reward arrives at the end. The policy-gradient estimator still works, but it gives every action in the trajectory the same final score. A brilliant move in a lost game is punished; a blunder in a won game is rewarded. These mistakes cancel out in expectation, but each individual update is noisy.

This is the credit-assignment problem: if an entire trajectory earned reward RR, which actions actually deserved credit?

There are two classic answers:

  • Monte Carlo: use the reward that actually followed the action. This is unbiased, but every random event before the end adds noise.
  • Temporal difference (TD): use a learned value function to estimate what will happen next. This reduces noise, but it is biased whenever the value estimate is wrong.

Generalized advantage estimation (GAE) provides a dial between these two choices:

A^tGAE=∑l≥0(γλ)l δt+l(5)\hat A^{\mathrm{GAE}}_t=\sum_{l\ge 0}(\gamma\lambda)^l\,\delta_{t+l} \tag{5}

At λ=0\lambda=0, GAE behaves like one-step TD. At λ=1\lambda=1, it becomes Monte Carlo return minus a baseline [11]. The important point is simpler than the formula: a critic—a learned value estimator—becomes useful when actions and rewards are far apart. The longer and more uncertain the path between them, the more variance the critic can remove.

LLMs often sit at the worst end of feedback density (F): they produce hundreds or thousands of tokens, then receive one score. There are two broad responses. Outcome-based training accepts the final score. Process reward models instead score intermediate reasoning steps, repairing the missing feedback by hand [12], [13]. The same process scorer can also guide tree search.

GRPO takes a different route. For each prompt, it samples a group of GG answers and compares each reward with the group’s average:

A^i=Ri−mean⁡(R1,…,RG)std⁡(R1,…,RG)(6)\hat A_i=\frac{R_i-\operatorname{mean}(R_1,\dots,R_G)}{\operatorname{std}(R_1,\dots,R_G)} \tag{6}

The group mean takes the place of a learned critic. Subtracting it is a valid baseline; dividing by the group’s standard deviation normalizes the update, although it also slightly changes how prompts of different difficulty are weighted. Later variants remove that effect [14], while leave-one-out baselines avoid it altogether [15].

GRPO therefore pays for sparse feedback with parallel samples instead of a learned value model. Process rewards make a different trade: they add intermediate supervision. Both are attempts to answer the same question—what part of a long answer earned the final reward?

The classical example is Montezuma’s Revenge, an Atari game in which the player can travel through long stretches without earning points. It is difficult because sparse feedback is combined with unknown dynamics. For an LLM, sts_t is the current prefix, ata_t is the next token, and rr may be a single score at the end of the answer.

PPO makes sample reuse safer

Next, shrink the interaction budget (B) by making fresh experience expensive. The available data may come from the previous policy πold\pi_{\text{old}}, from another policy μ\mu, or from a fixed dataset D\mathcal{D} that cannot be extended.

This creates a mismatch. The policy-gradient formula expects examples from the current policy, but our examples came from an older or different one. Importance sampling corrects for that mismatch:

∇θJ=Eτ∼μ ⁣[(∏tπθ(at∣st)μ(at∣st)) R(τ) ∑t∇θlog⁡πθ(at∣st)](7)\nabla_\theta J=\mathbb{E}_{\tau\sim\mu}\!\Big[\Big(\prod_t \tfrac{\pi_\theta(a_t\mid s_t)}{\mu(a_t\mid s_t)}\Big)\,R(\tau)\,\sum_t\nabla_\theta\log\pi_\theta(a_t\mid s_t)\Big] \tag{7}

The ratio asks: how much more or less likely was the new policy to produce this old trajectory? Unfortunately, multiplying one ratio for every step can create enormous variance, especially as the two policies drift apart [16].

PPO uses a practical safeguard. It considers the ratio one action at a time, ρt=πθ(at∣st)/πold(at∣st)\rho_t=\pi_\theta(a_t\mid s_t)/\pi_{\text{old}}(a_t\mid s_t), and clips updates that move it too far:

Lclip=E[min⁡(ρt A^t, clip⁡(ρt, 1−ϵ, 1+ϵ) A^t)](8)\mathcal{L}^{\text{clip}}=\mathbb{E}\big[\min\big(\rho_t\,\hat A_t,\ \operatorname{clip}(\rho_t,\,1-\epsilon,\,1+\epsilon)\,\hat A_t\big)\big] \tag{8}

The clipping rule is best understood as a tool for safer sample reuse. TRPO, PPO’s predecessor, constrained the distance from πold\pi_{\text{old}} directly. PPO replaces that expensive constraint with a simpler approximation [17], [18]. The clip acts like a seat belt: it limits the damage stale data can cause.

At the far end of the interaction-budget axis (B), interaction stops completely. We have only a frozen dataset D\mathcal{D}. Offline RL studies how to learn without drifting into parts of the world that the dataset does not cover [21]. One important LLM method lives at this endpoint; we will meet it in the next section.Value-based methods such as Q-learning are off-policy by design. This post follows the policy-gradient branch because it leads most directly to modern LLM training.

Modern LLM systems generate answers in large, asynchronous batches. By the time an update runs, those answers often came from a slightly older policy. This is why clipping remains useful even in methods described as GRPO rather than PPO: it solves a data-staleness problem, not a critic problem.

Real-world robotics is the classical example of an expensive budget. Every rollout consumes time and energy and can damage hardware. Its hand-designed reward functions also preview the final axis: a reward can be easy to calculate and still fail to express what we truly want.

When no true reward function exists

So far, we have assumed that a true reward exists. It might be hidden, delayed, or expensive to observe, but it is there.

Low reward fidelity (R) breaks that assumption. What number measures a good poem? What formula captures a helpful answer? The reward is not merely hidden—it does not exist as a ready-made function. This is the defining problem of RL for LLMs.

We can arrange possible reward sources from strongest to weakest:

  1. An analytic reward, such as a mathematical control cost, directly states the objective.
  2. A queryable checker, such as a compiler, unit test, or game score, is exact where it applies.
  3. An expensive judge, usually a human, can evaluate open-ended outputs but cannot score every training example [22].
  4. A learned proxy rϕr_\phi imitates those judgments cheaply, but becomes unreliable outside its training data.
  5. At the weakest end, we have no scores at all—only preferences or demonstrations.3

A common approach starts with pairwise preferences: for the same prompt, a person chooses the better of two answers. The Bradley–Terry model turns those comparisons into a learned reward rϕr_\phi [23]:

P(yw≻yl∣x)=σ(rϕ(x,yw)−rϕ(x,yl))(9)P(y_w \succ y_l \mid x)=\sigma\big(r_\phi(x,y_w)-r_\phi(x,y_l)\big) \tag{9}

Only the difference between the two scores matters. Adding the same constant to both would change nothing. That small fact will let us remove the explicit reward model in a moment.

Optimizing a proxy too aggressively creates a familiar Goodhart’s-law failure: once a measure becomes the target, it stops being a good measure. Experiments make this pattern visible [24]. As training pressure increases, the learned reward keeps rising while a stronger held-out evaluator eventually rates the answers worse.

Reinforcement learning from human feedback (RLHF) responds by limiting how far the policy may move from a trusted reference model. It maximizes learned reward while charging a KL-divergence penalty—a distance-like measure for probability distributions—for moving away from that reference:

max⁡π Ex∼D, y∼π(⋅∣x)[rϕ(x,y)]−β Ex∼D[KL(π(⋅∣x) ∥ πref(⋅∣x))](10)\max_{\pi}\ \mathbb{E}_{x\sim\mathcal{D},\,y\sim\pi(\cdot\mid x)}\big[r_\phi(x,y)\big]-\beta\,\mathbb{E}_{x\sim\mathcal{D}}\big[\mathrm{KL}\big(\pi(\cdot\mid x)\,\|\,\pi_{\text{ref}}(\cdot\mid x)\big)\big] \tag{10}

This objective has a closed-form solution: reweight the reference model toward high-reward answers [25].

π∗(y∣x)=1Z(x) πref(y∣x) erϕ(x,y)/β(11)\pi^*(y\mid x)=\frac{1}{Z(x)}\,\pi_{\text{ref}}(y\mid x)\,e^{r_\phi(x,y)/\beta} \tag{11}

Here Z(x)Z(x) is a normalization term that sums over every possible response, so we cannot compute it directly. But we can rearrange the equation to express the reward in terms of the policy:

rϕ(x,y)=βlog⁡π∗(y∣x)πref(y∣x)+βlog⁡Z(x)r_\phi(x,y)=\beta\log\frac{\pi^*(y\mid x)}{\pi_{\text{ref}}(y\mid x)}+\beta\log Z(x)

When we substitute this expression into the pairwise preference model, both answers share the same prompt. The impossible βlog⁡Z(x)\beta\log Z(x) term therefore cancels. We are left with a loss on the policy itself:

LDPO=− E(x,yw,yl)∼Dlog⁡σ ⁣(βlog⁡πθ(yw∣x)πref(yw∣x)−βlog⁡πθ(yl∣x)πref(yl∣x))(12)\mathcal{L}_{\mathrm{DPO}}=-\,\mathbb{E}_{(x,y_w,y_l)\sim\mathcal{D}}\log\sigma\!\Big(\beta\log\tfrac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}-\beta\log\tfrac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\Big) \tag{12}

This is Direct Preference Optimization (DPO). Its key insight is that we do not need to train a separate reward model first; the preference rankings already contain the information the policy needs [25].

DPO sits on two axes at once. On reward fidelity (R), it represents reward implicitly through preferences. On interaction budget (B), it is fully offline: (12) trains on a fixed dataset and never samples from the evolving policy. That makes DPO simple, but it also means the training data cannot correct the policy once it moves beyond the comparisons the dataset contains.

RLVR moves one rung up the reward ladder. On problems with checkable answers, it replaces the learned proxy with a verifier such as a test suite [2]. The change sounds small, but it improves the axis that matters most: the model can no longer win merely by fooling a learned reward model.

The improvement has limits. Feedback still arrives at the end, so sparse credit remains. Verifiers cover only checkable domains. And a checker is exact only about what it checks—a model can still game incomplete tests. Rubric graders and LLM judges try to extend the verifier to open-ended tasks, but doing so moves us back toward learned proxies.

We can now see why the fixed reference anchor exists. The moving leash protects an estimate based on stale data. The fixed anchor limits how far the model may chase a reward we do not fully trust.

β is the exchange rate between reward and trust.

The map with RLHF and RLVR pinned: one spoke movesTwo warm-colored polygons differing only on the R spoke.D · DifferentiabilityR · Reward fidelityF · Feedback densityB · Interaction budgetK · Dynamics knowledgeRLHFRLVRcenter = more informed · outward = more degraded
RLHF (dashed) and RLVR (solid) differ only in reward fidelity (R). A verifier gives RLVR a more trustworthy reward.

The corner classical RL never studied

We now have enough machinery. Where does single-turn LLM training sit on the five axes?

  • D — Differentiability: poor, but not important. Token sampling is discrete, yet policy gradients only need log⁡πθ\log\pi_\theta to be differentiable.The learned reward model rϕr_\phi is itself differentiable, and continuous relaxations of token sampling exist. Few systems use them, which is evidence that D is not the main constraint.
  • K — Dynamics knowledge: nearly ideal. The dynamics are string concatenation, known exactly and cheap to simulate.
  • F — Feedback density: poor. A long answer may receive only one score at the end.
  • B — Interaction budget: fairly good. Generating text costs compute, but not broken hardware or risky real-world interaction.
  • R — Reward fidelity: poor. For open-ended answers, the reward is usually a learned approximation of human judgment.

Classical RL usually assumed the reward was given and treated the environment as the mystery. Single-turn LLM training flips that picture: the environment is trivial, while the reward is the hard part.

Classical RL was constrained by dynamics. LLM training is constrained by reward.

Small multiples of six worlds on the five-axis mapSix mini radars comparing where each world is degraded.D Differentiability · K Dynamics knowledge · F Feedback densityB Interaction budget · R Reward fidelity · outward = more degradedDRFBKtrajectory optimizationAtariAlphaZeroreal roboticsRLHF, single turnRLVR
The assembled map. Classical problems (cool) tend to struggle with knowledge of the dynamics, while LLM problems (warm) tend to struggle with feedback and reward.

This placement answers the three opening puzzles and suggests one broader lesson.

1. A complete response can be treated as one action

Because nothing external happens between tokens, single-turn generation can be viewed as a contextual bandit—a one-step decision problem in which the prompt supplies context and the full response is the action [15], [25]. We can still choose a token-level view when we want per-token credit assignment or process rewards. Both descriptions are valid; they support different tools.When a per-token KL penalty is added to the reward, the token-level view gains a genuine dense reward signal—even though that factorization began as a modeling choice.

2. GRPO can remove the critic because sampling is a viable substitute

In long, stochastic environments, a critic is valuable because it estimates which actions led to a distant reward. In single-turn generation, the external world does not evolve between tokens, and many responses can be sampled in parallel. GRPO uses the group’s average reward as a Monte Carlo baseline instead of learning a value model.

A token-level critic can still reduce variance, so PPO-based RLHF is not mistaken; it is simply one of two ways to pay the same cost.

3. The two policy tethers protect against different failures

The moving leash—PPO’s clip relative to πold\pi_{\text{old}}—handles stale data. The fixed anchor—a KL penalty relative to πref\pi_{\text{ref}}—limits exploitation of an unreliable reward. One belongs to the estimator; the other belongs to the objective. That is why the anchor is often folded directly into the reward.

4. The most valuable improvements target reward and feedback

Verifiable rewards improve reward fidelity (R). Process rewards increase feedback density (F). Execution feedback supplies stronger rewards for code. AI feedback makes judgment cheaper [26], while LLM judges try to cover tasks without exact checkers.

RLVR’s one-rung improvement mattered because it moved the binding constraint. A small improvement at the bottleneck can matter more than a large improvement elsewhere.

Agents make the problem hard again

Everything above describes single-turn generation. Once an LLM can browse, call tools, run code, or talk with people over many turns, the problem becomes difficult again.

The dynamics are no longer string concatenation alone. They now include a browser, compiler, market, or person—systems that may be unknown, random, and constantly changing. Dynamics knowledge (K) gets worse.4 The interaction budget (B) gets worse too: rollouts consume time and money and can have real side effects. Feedback density (F) falls as success stretches across many tool calls. Reward fidelity (R) remains imperfect because most useful agent behavior still lacks an exact checker.

Single-turn RLHF is roughly a bandit with a synthetic reward. Agentic RL combines classical uncertainty about the environment with the LLM-era problem of an untrusted reward. It is hard in both old and new ways at once.

The map therefore predicts that classical RL tools will return as LLMs become more agentic. We already see tree search over agent actions [27] and value-guided search over reasoning steps [10]. World models of tools and users, along with conservative methods for expensive rollouts, are natural next steps.

The map with the agentic LLM pinned: the largest shapeA large warm-colored polygon dwarfing the dashed RLHF shape.D · DifferentiabilityR · Reward fidelityF · Feedback densityB · Interaction budgetK · Dynamics knowledgeagentic LLMRLHFcenter = more informed · outward = more degraded
An agentic LLM (solid) compared with single-turn RLHF (dashed). Tool use makes dynamics less known, rollouts more expensive, and feedback more delayed, while reward remains imperfect.

RL for agents feels harder than RLHF because the classical difficulties have returned while the reward remains weak.

How to use the map

The thesis in one line: every RL method pays for something the optimizer does not know or cannot access. The goal stays the same; the available information changes.

Use the map as a reading tool. For any new method, ask:

  • Which axis does it move closer to ideal?
  • What cost does that save?
  • What new cost or assumption does it introduce?

Try the questions on four examples. The footnotes give answers for the first three:

  1. Replacing human raters with AI feedback.5
  2. Dreamer-style world models [28].6
  3. Self-play.7
  4. Open: place test-time search — a reasoning model spending inference compute on a tree before answering — on the map yourself. Both of its coordinates appeared in this post.

For technical reference, here is the notation used throughout the post:

symbolfixed meaningwhich axis touches it
st, ats_t,\ a_tstate and action—
πθ(at∣st)\pi_\theta(a_t\mid s_t)the policy: the only object we always own—
ffdynamics, st+1=f(st,at)s_{t+1}=f(s_t,a_t) or p(st+1∣st,at)p(s_{t+1}\mid s_t,a_t)D (differentiability), K (dynamics knowledge)
r, R(τ)r,\ R(\tau)per-step reward; return R(τ)=∑tγtrtR(\tau)=\sum_t\gamma^t r_tF (feedback density), R (reward fidelity)
J(θ)=Eτ∼pθ[R(τ)]J(\theta)=\mathbb{E}_{\tau\sim p_\theta}[R(\tau)]the goal, which stays the same in every section—
V, AV,\ Avalue and advantage estimatesuseful when feedback density (F) is low
rϕr_\phia learned reward modelreward fidelity (R)
πold\pi_{\text{old}}last iteration’s policy — the leash’s tetherinteraction budget (B)
πref\pi_{\text{ref}}the frozen reference policy — the anchor’s tetherreward fidelity (R)
μ, D\mu,\ \mathcal{D}someone else’s policy; a frozen datasetinteraction budget (B)
β, γ, λ, ϵ\beta,\ \gamma,\ \lambda,\ \epsilonKL weight, discount, credit dial, clip width—
table 1. Notation used throughout the post.

Research effort did not disappear when RL moved from games and robots to language models. It moved to different axes: away from learning the dynamics and toward improving feedback and reward.

Reward is the new dynamics.

Footnotes

  1. Ideal means maximally informed, not solved: pathwise gradients through long or chaotic rollouts can explode even when every derivative exists [6], [7], and a hard if-branch can hide the true gradient from autodiff entirely — the subject of the robot post. ↩

  2. Exploration is deliberately not a sixth axis. It is not an independent dial you set; it is the tax collected when K or R are degraded under a finite budget B — if you knew ff or the true rr, you would not need to poke the world to find out. ↩

  3. The floor is inverse RL: recover a reward from behavior alone [29], [30]. Deliberately out of scope here — this post stops one rung above, where preferences are at least labeled. ↩

  4. Strictly, the context window is now an observation of a larger world state, not the state itself. Partial observability deserves to be a sixth axis someday; in this post it stays a footnote because every LLM world on the map shares it equally. ↩

  5. Axis R: it replaces the expensive oracle with a cheap proxy of the oracle — one rung down in fidelity to buy several rungs of scale. The subtle debt is that it is a proxy of a proxy: the judge model was itself trained on human preferences. ↩

  6. Axis K: learn f^\hat f to recover queryability, then spend it twice — planning (search against the model) and pathwise gradients through the model’s smooth latent dynamics, a partial repair of D purchased with model bias. ↩

  7. Axis B: an opponent that is always exactly as strong as you is a data generator with no marginal cost and a built-in curriculum — B pushed toward ideal, which is precisely the purchase AlphaZero used to afford its terminal-only F. ↩

references

[1]
Z. Shao et al., “DeepSeekMath: Pushing the limits of mathematical reasoning in open language models,” arXiv preprint arXiv:2402.03300, 2024.
[2]
D. Guo et al., “DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning,” Nature, vol. 645, no. 8081, pp. 633–638, 2025, doi: 10.1038/s41586-025-09422-z.
[3]
S. Mohamed, M. Rosca, M. Figurnov, and A. Mnih, “Monte Carlo gradient estimation in machine learning,” Journal of Machine Learning Research, vol. 21, no. 132, pp. 1–62, 2020.
[4]
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, no. 3–4, pp. 229–256, 1992.
[5]
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems, 2000, pp. 1057–1063.
[6]
L. Metz, C. D. Freeman, S. S. Schoenholz, and T. Kachman, “Gradients are not all you need,” arXiv preprint arXiv:2111.05803, 2021.
[7]
H. J. T. Suh, M. Simchowitz, K. Zhang, and R. Tedrake, “Do differentiable simulators give better policy gradients?,” in Proceedings of the 39th International Conference on Machine Learning, in PMLR, vol. 162. 2022, pp. 20668–20696.
[8]
V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540, pp. 529–533, 2015.
[9]
D. Silver et al., “A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play,” Science, vol. 362, no. 6419, pp. 1140–1144, 2018.
[10]
X. Feng et al., “Alphazero-like tree-search can guide large language model decoding and training,” arXiv preprint arXiv:2309.17179, 2023.
[11]
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in International Conference on Learning Representations, 2016.
[12]
J. Uesato et al., “Solving math word problems with process- and outcome-based feedback,” arXiv preprint arXiv:2211.14275, 2022.
[13]
H. Lightman et al., “Let’s verify step by step,” in International Conference on Learning Representations, 2024.
[14]
Z. Liu et al., “Understanding R1-Zero-like training: A critical perspective,” arXiv preprint arXiv:2503.20783, 2025.
[15]
A. Ahmadian et al., “Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024.
[16]
D. Precup, R. S. Sutton, and S. Singh, “Eligibility traces for off-policy policy evaluation,” in Proceedings of the 17th International Conference on Machine Learning, 2000, pp. 759–766.
[17]
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, in PMLR, vol. 37. 2015, pp. 1889–1897.
[18]
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
[19]
D. M. Ziegler et al., “Fine-tuning language models from human preferences,” arXiv preprint arXiv:1909.08593, 2019.
[20]
L. Ouyang et al., “Training language models to follow instructions with human feedback,” in Advances in Neural Information Processing Systems, 2022, pp. 27730–27744.
[21]
S. Levine, A. Kumar, G. Tucker, and J. Fu, “Offline reinforcement learning: Tutorial, review, and perspectives on open problems,” arXiv preprint arXiv:2005.01643, 2020.
[22]
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Advances in Neural Information Processing Systems, 2017.
[23]
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. The method of paired comparisons,” Biometrika, vol. 39, no. 3/4, pp. 324–345, 1952.
[24]
L. Gao, J. Schulman, and J. Hilton, “Scaling laws for reward model overoptimization,” in Proceedings of the 40th International Conference on Machine Learning, in PMLR, vol. 202. 2023, pp. 10835–10866.
[25]
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” in Advances in Neural Information Processing Systems, 2023.
[26]
Y. Bai et al., “Constitutional AI: Harmlessness from AI feedback,” arXiv preprint arXiv:2212.08073, 2022.
[27]
A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang, “Language agent tree search unifies reasoning, acting, and planning in language models,” in Proceedings of the 41st International Conference on Machine Learning, in PMLR, vol. 235. 2024.
[28]
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in International Conference on Learning Representations, 2020.
[29]
A. Y. Ng and S. Russell, “Algorithms for inverse reinforcement learning,” in Proceedings of the 17th International Conference on Machine Learning, 2000, pp. 663–670.
[30]
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning,” in Proceedings of the 23rd AAAI Conference on Artificial Intelligence, 2008, pp. 1433–1438.

cite

Found this useful or inspiring? Consider citing it — and follow along on X.

@article{feng2026taxonomyof,
  title   = {The Ideal RL Problem Isn't RL},
  author  = {Feng, Aosong},
  journal = {asfeng.dev},
  year    = {2026},
  month   = {Aug},
  url     = {https://asfeng.dev/blog/taxonomy-of-ignorance/}
}

© 2026 Aosong Feng · text licensed CC BY 4.0 — quote and reuse with attribution.