The Ideal RL Problem Isn't RL
A five-axis framework for understanding why different RL algorithms exist—and why language-model training is limited more by reward than by the environment.
This map will help us answer three puzzles from the LLM era:
- Group Relative Policy Optimization (GRPO) removed the value network that Proximal Policy Optimization (PPO) usually relies on, yet training still worked [1]. Why?
- LLM training often keeps the policy close to two different reference points. The constraints look similar, but they solve different problems. Why do we need both?
- Reinforcement learning with verifiable rewards (RLVR) replaced a learned reward model with a checker, such as a test harness. Why did that apparently small change matter so much [2]?
This is not a survey of RL algorithms. It is a way to see why whole families of algorithms exist. We will begin with an ideal world, then remove one useful piece of information at a time. Each loss forces us to pay a different price: more variance, more data, more modeling, or more distrust of the reward.
The equations provide the technical spine, but you can skim them without losing the main argument. A policy is the decision-maker. A trajectory is one sequence of states and actions, and its return is the total reward it earns.
The goal never changes. We still want to maximize
the average return of trajectories sampled from a policy with parameters . What changes is what we are allowed to know, observe, and differentiate. By the end, the three puzzles above will be simple consequences of where LLM training sits on the map.
The ideal world is a computing problem
Start with a world in which we know everything we could reasonably ask for. In that world, there is no learning problem left. There is only a computation.
Imagine controlling a robot arm inside a perfect simulator. We know the physics, every movement has a clear cost, and small changes to the motors produce smooth, predictable changes in the arm. We can try as many movements as we like. This is the ideal corner of our map.
Our ideal world has five properties. The letters will label the five axes of the map:
- D — Differentiability. The environment and reward are smooth enough to differentiate.
- K — Dynamics knowledge. We know exactly how the environment changes after an action.
- F — Feedback density. Every step tells us how well it went.
- B — Interaction budget. We can collect as much fresh experience as we need.
- R — Reward fidelity. The reward exactly expresses the goal we care about.
When all five properties hold, we can write the entire process explicitly. The policy chooses an action, the known dynamics move the world to its next state, and rewards add up over time:
Every arrow is a function we know and can differentiate. The whole rollout is one computation graph: affects the first action, that action affects the next state, and so on until the final reward. The chain rule gives us the exact gradient . In other words, we backpropagate through the world.
This is trajectory optimization, a standard problem in optimal control. Notice what we do not need: exploration, a value function, or sampled estimates of the gradient. Nothing must be learned because nothing is unknown.1
That is the first key idea: the ideal RL problem is not RL. RL is what remains when we can no longer solve the problem this directly.
The ideal world is a point; RL is everything that grows outward from it.
Every radar chart below follows the same rule: the center means more information and an easier optimization problem; moving outward means that condition has degraded. The axes are qualitative, not measurements on a shared numerical scale. Their purpose is to compare which limitations dominate in different settings.
The whole argument can be previewed in one table. Each row starts with a lost capability and ends with the family of methods that compensates for it:
| axis | what becomes unavailable | problem created | typical response |
|---|---|---|---|
| D — Differentiability | derivatives through the environment | exact backpropagation fails | score-function gradients such as REINFORCE |
| K — Dynamics knowledge | a usable model of what happens next | counterfactual futures are hard to evaluate | search or learned world models |
| F — Feedback density | frequent signals about progress | credit for a final result is hard to assign | critics, GAE, process rewards, or group baselines |
| B — Interaction budget | unlimited fresh experience | data becomes scarce or stale | PPO clipping, off-policy correction, or offline RL |
| R — Reward fidelity | a reward that exactly matches the real goal | the policy can exploit an imperfect proxy | reference anchors, preference learning, or verifiers |
This ideal case suggests a useful first question for any new problem: Can I simply backpropagate through it? The rest of the post is about what to do when the answer is no.2
Only your policy must be differentiable
Now remove differentiability (D). We can still run the world as often as we like, but we cannot differentiate through it. Perhaps the actions are discrete, the reward contains hard branches, or the simulator is a compiled program we cannot inspect.
The previous strategy no longer works. But we can make one important shift: instead of differentiating through the world, differentiate the probability of visiting different trajectories. The log-derivative trick gives us
In plain English: sample a trajectory, see how much reward it earns, and make rewarded trajectories more likely. Here scores the trajectory, while measures how a parameter update would change its probability.Moving through the integral requires regularity assumptions. They are harmless here; Mohamed and colleagues survey the exceptions [3].
Why does this work even when the environment is a black box? A trajectory’s probability splits into three pieces: where it starts, what actions the policy chooses, and how the environment responds.
Only the policy depends on . When we take the derivative, the initial-state and environment terms disappear, leaving .
This is the surprising part: the environment disappears from the gradient formula. Its uncertainty returns as noise in our estimate. This result underlies REINFORCE and the policy-gradient theorem [4], [5]. The environment can be jagged, discrete, or hidden. Only the policy—the part we built—must be differentiable.
For a slower derivation with a robot and concrete numbers, see Two Gradients, One Flow and the companion VAE post.
There is another route: a pathwise gradient holds the random noise fixed and differentiates the reward along the sampled path. It usually has lower variance because it uses local slope information, but it requires derivatives from the environment.
The score-function estimator uses only sampled rewards, so it works with black boxes. Its higher variance is the price of that generality [6], [7]. Much of modern RL is an attempt to reduce this variance without giving up the black-box advantage.
The first improvement is free: subtract a baseline from every return. This does not change the expected gradient because the expected policy score is zero—the derivative of “all probabilities sum to one” is also zero:
A baseline changes the variance of the estimate, but not its average. A state-dependent baseline works too. This simple observation leads to several familiar techniques: assign an action only the rewards that follow it, learn a value function as the baseline, and update on the advantage —how much better an action was than expected.
We will meet three more variance-reduction tools later: temporal credit assignment, PPO’s clipping rule, and GRPO’s group baseline. They look related, but each responds to a different missing piece of information.
This is exactly what makes policy gradients useful for LLMs. Tokens are discrete, reward models can be treated as black boxes, and sampling itself is not differentiable. None of that matters because comes from our own neural network. Policy gradients became a universal adapter: they ask almost nothing of the environment.
Atari is the classical example [8]. Its emulator can be run but not differentiated through. It also teaches a broader lesson: real problems usually lose several ideal conditions at once.
Perfect knowledge, zero gradients
Differentiability (D) and dynamics knowledge (K) are different. We can know exactly how a world works without being able to differentiate through it.
Knowledge comes in levels. At the best end, we have a closed-form equation. One step down, we have an executable simulator: we can ask “what if?” but cannot use calculus. Below that, we can learn an approximate model , gaining the ability to simulate at the cost of model errors. At the worst end, we know nothing beyond the trajectories we have already observed.
When a simulator is known but not differentiable, search can replace gradients. The rules of Go are exact, yet there is no useful derivative of a move. Monte Carlo tree search (MCTS) instead explores many possible futures. AlphaZero combines that search with policy and value networks, gradually compressing expensive lookahead into learned intuition [9].
AlphaZero is not ideal on every other axis: its reward is only a win or loss at the end of a game. That means it has perfect knowledge of the rules but very sparse feedback. Real problems occupy combinations of axes, not neat stops on a single line.
Single-turn LLM generation lands at an unusual extreme. Its “dynamics” are just : append the next token to the existing prefix. This operation is deterministic, known, and cheap. No one needs to learn a world model for ordinary text generation because the world model is string concatenation.
This also explains why tree search can reappear in LLM reasoning. A learned scorer evaluates partial solutions while search explores possible next steps [10]. It is the AlphaZero pattern replayed in a much simpler environment.
When reward comes last, credit becomes inference
Now reduce feedback density (F). Suppose the only reward arrives at the end. The policy-gradient estimator still works, but it gives every action in the trajectory the same final score. A brilliant move in a lost game is punished; a blunder in a won game is rewarded. These mistakes cancel out in expectation, but each individual update is noisy.
This is the credit-assignment problem: if an entire trajectory earned reward , which actions actually deserved credit?
There are two classic answers:
- Monte Carlo: use the reward that actually followed the action. This is unbiased, but every random event before the end adds noise.
- Temporal difference (TD): use a learned value function to estimate what will happen next. This reduces noise, but it is biased whenever the value estimate is wrong.
Generalized advantage estimation (GAE) provides a dial between these two choices:
At , GAE behaves like one-step TD. At , it becomes Monte Carlo return minus a baseline [11]. The important point is simpler than the formula: a critic—a learned value estimator—becomes useful when actions and rewards are far apart. The longer and more uncertain the path between them, the more variance the critic can remove.
LLMs often sit at the worst end of feedback density (F): they produce hundreds or thousands of tokens, then receive one score. There are two broad responses. Outcome-based training accepts the final score. Process reward models instead score intermediate reasoning steps, repairing the missing feedback by hand [12], [13]. The same process scorer can also guide tree search.
GRPO takes a different route. For each prompt, it samples a group of answers and compares each reward with the group’s average:
The group mean takes the place of a learned critic. Subtracting it is a valid baseline; dividing by the group’s standard deviation normalizes the update, although it also slightly changes how prompts of different difficulty are weighted. Later variants remove that effect [14], while leave-one-out baselines avoid it altogether [15].
GRPO therefore pays for sparse feedback with parallel samples instead of a learned value model. Process rewards make a different trade: they add intermediate supervision. Both are attempts to answer the same question—what part of a long answer earned the final reward?
The classical example is Montezuma’s Revenge, an Atari game in which the player can travel through long stretches without earning points. It is difficult because sparse feedback is combined with unknown dynamics. For an LLM, is the current prefix, is the next token, and may be a single score at the end of the answer.
PPO makes sample reuse safer
Next, shrink the interaction budget (B) by making fresh experience expensive. The available data may come from the previous policy , from another policy , or from a fixed dataset that cannot be extended.
This creates a mismatch. The policy-gradient formula expects examples from the current policy, but our examples came from an older or different one. Importance sampling corrects for that mismatch:
The ratio asks: how much more or less likely was the new policy to produce this old trajectory? Unfortunately, multiplying one ratio for every step can create enormous variance, especially as the two policies drift apart [16].
PPO uses a practical safeguard. It considers the ratio one action at a time, , and clips updates that move it too far:
The clipping rule is best understood as a tool for safer sample reuse. TRPO, PPO’s predecessor, constrained the distance from directly. PPO replaces that expensive constraint with a simpler approximation [17], [18]. The clip acts like a seat belt: it limits the damage stale data can cause.
At the far end of the interaction-budget axis (B), interaction stops completely. We have only a frozen dataset . Offline RL studies how to learn without drifting into parts of the world that the dataset does not cover [21]. One important LLM method lives at this endpoint; we will meet it in the next section.Value-based methods such as Q-learning are off-policy by design. This post follows the policy-gradient branch because it leads most directly to modern LLM training.
Modern LLM systems generate answers in large, asynchronous batches. By the time an update runs, those answers often came from a slightly older policy. This is why clipping remains useful even in methods described as GRPO rather than PPO: it solves a data-staleness problem, not a critic problem.
Real-world robotics is the classical example of an expensive budget. Every rollout consumes time and energy and can damage hardware. Its hand-designed reward functions also preview the final axis: a reward can be easy to calculate and still fail to express what we truly want.
When no true reward function exists
So far, we have assumed that a true reward exists. It might be hidden, delayed, or expensive to observe, but it is there.
Low reward fidelity (R) breaks that assumption. What number measures a good poem? What formula captures a helpful answer? The reward is not merely hidden—it does not exist as a ready-made function. This is the defining problem of RL for LLMs.
We can arrange possible reward sources from strongest to weakest:
- An analytic reward, such as a mathematical control cost, directly states the objective.
- A queryable checker, such as a compiler, unit test, or game score, is exact where it applies.
- An expensive judge, usually a human, can evaluate open-ended outputs but cannot score every training example [22].
- A learned proxy imitates those judgments cheaply, but becomes unreliable outside its training data.
- At the weakest end, we have no scores at all—only preferences or demonstrations.3
A common approach starts with pairwise preferences: for the same prompt, a person chooses the better of two answers. The Bradley–Terry model turns those comparisons into a learned reward [23]:
Only the difference between the two scores matters. Adding the same constant to both would change nothing. That small fact will let us remove the explicit reward model in a moment.
Optimizing a proxy too aggressively creates a familiar Goodhart’s-law failure: once a measure becomes the target, it stops being a good measure. Experiments make this pattern visible [24]. As training pressure increases, the learned reward keeps rising while a stronger held-out evaluator eventually rates the answers worse.
Reinforcement learning from human feedback (RLHF) responds by limiting how far the policy may move from a trusted reference model. It maximizes learned reward while charging a KL-divergence penalty—a distance-like measure for probability distributions—for moving away from that reference:
This objective has a closed-form solution: reweight the reference model toward high-reward answers [25].
Here is a normalization term that sums over every possible response, so we cannot compute it directly. But we can rearrange the equation to express the reward in terms of the policy:
When we substitute this expression into the pairwise preference model, both answers share the same prompt. The impossible term therefore cancels. We are left with a loss on the policy itself:
This is Direct Preference Optimization (DPO). Its key insight is that we do not need to train a separate reward model first; the preference rankings already contain the information the policy needs [25].
DPO sits on two axes at once. On reward fidelity (R), it represents reward implicitly through preferences. On interaction budget (B), it is fully offline: (12) trains on a fixed dataset and never samples from the evolving policy. That makes DPO simple, but it also means the training data cannot correct the policy once it moves beyond the comparisons the dataset contains.
RLVR moves one rung up the reward ladder. On problems with checkable answers, it replaces the learned proxy with a verifier such as a test suite [2]. The change sounds small, but it improves the axis that matters most: the model can no longer win merely by fooling a learned reward model.
The improvement has limits. Feedback still arrives at the end, so sparse credit remains. Verifiers cover only checkable domains. And a checker is exact only about what it checks—a model can still game incomplete tests. Rubric graders and LLM judges try to extend the verifier to open-ended tasks, but doing so moves us back toward learned proxies.
We can now see why the fixed reference anchor exists. The moving leash protects an estimate based on stale data. The fixed anchor limits how far the model may chase a reward we do not fully trust.
β is the exchange rate between reward and trust.
The corner classical RL never studied
We now have enough machinery. Where does single-turn LLM training sit on the five axes?
- D — Differentiability: poor, but not important. Token sampling is discrete, yet policy gradients only need to be differentiable.The learned reward model is itself differentiable, and continuous relaxations of token sampling exist. Few systems use them, which is evidence that D is not the main constraint.
- K — Dynamics knowledge: nearly ideal. The dynamics are string concatenation, known exactly and cheap to simulate.
- F — Feedback density: poor. A long answer may receive only one score at the end.
- B — Interaction budget: fairly good. Generating text costs compute, but not broken hardware or risky real-world interaction.
- R — Reward fidelity: poor. For open-ended answers, the reward is usually a learned approximation of human judgment.
Classical RL usually assumed the reward was given and treated the environment as the mystery. Single-turn LLM training flips that picture: the environment is trivial, while the reward is the hard part.
Classical RL was constrained by dynamics. LLM training is constrained by reward.
This placement answers the three opening puzzles and suggests one broader lesson.
1. A complete response can be treated as one action
Because nothing external happens between tokens, single-turn generation can be viewed as a contextual bandit—a one-step decision problem in which the prompt supplies context and the full response is the action [15], [25]. We can still choose a token-level view when we want per-token credit assignment or process rewards. Both descriptions are valid; they support different tools.When a per-token KL penalty is added to the reward, the token-level view gains a genuine dense reward signal—even though that factorization began as a modeling choice.
2. GRPO can remove the critic because sampling is a viable substitute
In long, stochastic environments, a critic is valuable because it estimates which actions led to a distant reward. In single-turn generation, the external world does not evolve between tokens, and many responses can be sampled in parallel. GRPO uses the group’s average reward as a Monte Carlo baseline instead of learning a value model.
A token-level critic can still reduce variance, so PPO-based RLHF is not mistaken; it is simply one of two ways to pay the same cost.
3. The two policy tethers protect against different failures
The moving leash—PPO’s clip relative to —handles stale data. The fixed anchor—a KL penalty relative to —limits exploitation of an unreliable reward. One belongs to the estimator; the other belongs to the objective. That is why the anchor is often folded directly into the reward.
4. The most valuable improvements target reward and feedback
Verifiable rewards improve reward fidelity (R). Process rewards increase feedback density (F). Execution feedback supplies stronger rewards for code. AI feedback makes judgment cheaper [26], while LLM judges try to cover tasks without exact checkers.
RLVR’s one-rung improvement mattered because it moved the binding constraint. A small improvement at the bottleneck can matter more than a large improvement elsewhere.
Agents make the problem hard again
Everything above describes single-turn generation. Once an LLM can browse, call tools, run code, or talk with people over many turns, the problem becomes difficult again.
The dynamics are no longer string concatenation alone. They now include a browser, compiler, market, or person—systems that may be unknown, random, and constantly changing. Dynamics knowledge (K) gets worse.4 The interaction budget (B) gets worse too: rollouts consume time and money and can have real side effects. Feedback density (F) falls as success stretches across many tool calls. Reward fidelity (R) remains imperfect because most useful agent behavior still lacks an exact checker.
Single-turn RLHF is roughly a bandit with a synthetic reward. Agentic RL combines classical uncertainty about the environment with the LLM-era problem of an untrusted reward. It is hard in both old and new ways at once.
The map therefore predicts that classical RL tools will return as LLMs become more agentic. We already see tree search over agent actions [27] and value-guided search over reasoning steps [10]. World models of tools and users, along with conservative methods for expensive rollouts, are natural next steps.
RL for agents feels harder than RLHF because the classical difficulties have returned while the reward remains weak.
How to use the map
The thesis in one line: every RL method pays for something the optimizer does not know or cannot access. The goal stays the same; the available information changes.
Use the map as a reading tool. For any new method, ask:
- Which axis does it move closer to ideal?
- What cost does that save?
- What new cost or assumption does it introduce?
Try the questions on four examples. The footnotes give answers for the first three:
- Replacing human raters with AI feedback.5
- Dreamer-style world models [28].6
- Self-play.7
- Open: place test-time search — a reasoning model spending inference compute on a tree before answering — on the map yourself. Both of its coordinates appeared in this post.
For technical reference, here is the notation used throughout the post:
| symbol | fixed meaning | which axis touches it |
|---|---|---|
| state and action | — | |
| the policy: the only object we always own | — | |
| dynamics, or | D (differentiability), K (dynamics knowledge) | |
| per-step reward; return | F (feedback density), R (reward fidelity) | |
| the goal, which stays the same in every section | — | |
| value and advantage estimates | useful when feedback density (F) is low | |
| a learned reward model | reward fidelity (R) | |
| last iteration’s policy — the leash’s tether | interaction budget (B) | |
| the frozen reference policy — the anchor’s tether | reward fidelity (R) | |
| someone else’s policy; a frozen dataset | interaction budget (B) | |
| KL weight, discount, credit dial, clip width | — |
Research effort did not disappear when RL moved from games and robots to language models. It moved to different axes: away from learning the dynamics and toward improving feedback and reward.
Reward is the new dynamics.
Footnotes
-
Ideal means maximally informed, not solved: pathwise gradients through long or chaotic rollouts can explode even when every derivative exists [6], [7], and a hard if-branch can hide the true gradient from autodiff entirely — the subject of the robot post. ↩
-
Exploration is deliberately not a sixth axis. It is not an independent dial you set; it is the tax collected when K or R are degraded under a finite budget B — if you knew or the true , you would not need to poke the world to find out. ↩
-
The floor is inverse RL: recover a reward from behavior alone [29], [30]. Deliberately out of scope here — this post stops one rung above, where preferences are at least labeled. ↩
-
Strictly, the context window is now an observation of a larger world state, not the state itself. Partial observability deserves to be a sixth axis someday; in this post it stays a footnote because every LLM world on the map shares it equally. ↩
-
Axis R: it replaces the expensive oracle with a cheap proxy of the oracle — one rung down in fidelity to buy several rungs of scale. The subtle debt is that it is a proxy of a proxy: the judge model was itself trained on human preferences. ↩
-
Axis K: learn to recover queryability, then spend it twice — planning (search against the model) and pathwise gradients through the model’s smooth latent dynamics, a partial repair of D purchased with model bias. ↩
-
Axis B: an opponent that is always exactly as strong as you is a data generator with no marginal cost and a built-in curriculum — B pushed toward ideal, which is precisely the purchase AlphaZero used to afford its terminal-only F. ↩
references
cite
Found this useful or inspiring? Consider citing it — and follow along on X.
@article{feng2026taxonomyof,
title = {The Ideal RL Problem Isn't RL},
author = {Feng, Aosong},
journal = {asfeng.dev},
year = {2026},
month = {Aug},
url = {https://asfeng.dev/blog/taxonomy-of-ignorance/}
}© 2026 Aosong Feng · text licensed CC BY 4.0 — quote and reuse with attribution.