Research & Technique 29 September 2026 8 min read 1,650 words

A reward is a sieve, not a lesson

Google Cloud now sells reinforcement learning fine-tuning as a managed service: you bring prompts and a reward function, it handles the rest. Read the mechanism underneath and the purchase changes shape, because the only thing a reward can do is choose between answers your model already gives.

The argument

Reinforcement learning fine-tuning re-ranks the samples a model already produces rather than teaching it your task, so its only usable prompts are the ones your base model already gets right some of the time.

On 25 September, two Google engineers published a best-practices guide for the company's managed reinforcement learning fine-tuning service, and a few paragraphs in they name the limit of the thing they are selling. RLFT "amplifies existing competence," the guide says. "It makes occasional success reliable, but it can't teach a skill the model never demonstrates."

That is the most useful sentence on the page, and it is in a bullet. The pitch above it is simpler: "you bring prompts and a reward function; we handle the infrastructure and the proprietary model internals." Reinforcement learning used to require a training cluster and access to weights nobody outside the lab has. Now it is a form you fill in. What the form does not say, and what the bullet says quietly, is that the service cannot give your model an ability it does not already have. It can only make one of its existing habits more likely.

So the reward function you write is not a description of your task. It is a sieve. It sits over the model's own output distribution and decides which of the things already coming through are to come through more often. That reframing changes what you should measure before you buy anything, and it is measurable for the price of some inference.

What the loop actually computes

The managed service will not show you its optimiser: "the reinforcement learning that makes this work is fully managed, you never configure it." But the shape is not secret, because the open implementations of the same family of algorithms are readable. Hugging Face's TRL library ships GRPO, group relative policy optimisation, and its documentation states the property that matters first: "GRPO is an online learning algorithm, meaning it improves iteratively by using the data generated by the trained model itself during training."

Concretely, one training step goes like this. Take a prompt. Sample \( G \) completions from the current model. Score each one with the reward function you wrote, giving rewards \( r_1 \dots r_G \). Then compute, for each completion, an advantage:

\[ \hat{A}_i = \frac{r_i - \operatorname{mean}(\mathbf{r})}{\operatorname{std}(\mathbf{r})} \]

That is the line in TRL's trainer, near enough verbatim: advantages = rewards - mean_grouped_rewards, then a division by the group's standard deviation. The loss raises the probability of tokens in above-average completions and lowers it for below-average ones, with a penalty term to keep the updated model near where it started.

Read the numerator. The reward never enters the gradient on its own. Only its distance from the average of the group does. A reward of 0.9 teaches nothing if the other seven samples also scored 0.9. A reward of zero teaches nothing if everything scored zero. The signal is not "this answer is good", it is "this answer is better than what you otherwise said", and it exists only where the model disagrees with itself.

TRL takes this seriously enough to log it. Among its training metrics is frac_reward_zero_std, "the fraction of samples in the generation batch with a reward std of zero, implying there is little diversity for that prompt (all answers are correct or incorrect)." It is a counter for how much of each step was wasted. I know of no more honest number in RL practice, and it does not appear in any launch material.

The arithmetic of a dead prompt

Put the crudest possible model on it. Suppose your reward is binary, your prompt has some per-sample success probability \( p \), and the \( G \) samples are independent. The prompt produces a usable gradient only when it is neither all successes nor all failures, with probability \( 1 - p^G - (1-p)^G \).

With the common choice of \( G = 8 \), that function is brutal at the edges. At \( p = 0.5 \), 99 per cent of prompts carry signal. At \( p = 0.1 \), 57 per cent do. At \( p = 0.05 \), 34 per cent. At \( p = 0.01 \), under 8 per cent, meaning that ninety-two times in a hundred the service generates eight completions, runs your grader eight times, and moves the weights not at all.

The same cliff sits at the top. At \( p = 0.9 \) you are back to 57 per cent, which is why a task your model nearly always gets right is also a poor RL target: there is nothing left to select between. Real sampling is not independent and real rewards are not binary, both of which soften the curve, but the shape survives. The most informative prompt in your dataset is the one your model gets right about half the time.

This is exactly why Google's guide tells you to "direct RLFT when the base model already succeeds part of the time, enough for the reward to tell better answers from worse ones," and to use a light supervised warm start otherwise, because "the base success rate is too low for RL to gain traction." Those are not preferences. They are the denominator of the advantage, written as advice.

The measurement you should take before you pay

Here is the part practitioners keep skipping. The quantity that decides whether RLFT can work on your task is a per-prompt count: out of \( n \) samples, how many passed. That is not a new instrument. It is what the HumanEval harness has always computed. Its estimator, in human_eval/evaluation.py, is

\[ \text{pass@}k = 1 - \frac{\binom{n-c}{k}}{\binom{n}{k}} \]

for \( c \) correct out of \( n \) samples per problem. Everyone quotes the aggregate. The number that predicts RL trainability is the intermediate one, \( c \), and its distribution across your prompts.

So: take two hundred of your own prompts, sample sixteen completions each, score them with the reward function you were going to write anyway, and plot the histogram of \( c \). If the mass sits at \( c = 0 \), reinforcement learning has nothing to select and you need demonstrations, better retrieval, a stronger base model or a different decomposition of the task. If it sits at \( c = n \), you do not need tuning, you need a cheaper model. RLFT earns its cost on the middle of that histogram, and you can see the middle before you commission a training run. The guide tells you to watch the reward curve in the console once you have started. The sharper number is available the day before.

The sieve framing also predicts which rewards fail. A reward must discriminate between the near misses your model actually emits, not between excellence and catastrophe, because catastrophe is not in the group. And a sieve will pass anything shaped to fit it. Google's own content moderation case study says the plain failure was models that "reward-hack with invalid formats to dodge evaluation", fixed by pairing format validation with a deterministic grader. Their reward hygiene list, "ensemble judges, penalize length, floor degenerate outputs, and prefer a verifiable check over a model's opinion", is not a style guide. It is maintenance on a filter that an optimiser is actively probing for holes. TRL's documentation flags the same known hole from the other side, noting work that found GRPO's sample-level loss under-penalises long responses, so length quietly becomes the cheapest thing to optimise.

The strongest case against this

Sharpening is not a consolation prize, and that is the serious objection. Reliability is the product. A system that writes an executable query against an unseen schema two times in five is a demonstration; one that does it nineteen times in twenty is a feature, and no amount of prompt revision closes that gap. Google's guide reports precisely that kind of win, in code graded by execution, in entity extraction where supervised tuning had plateaued, in slide layouts scored by rendering them. If the base model's occasional competence is the raw material, then converting occasional into dependable is most of the value in applied machine learning.

I accept all of it, and it strengthens the argument rather than answering it. If what you are buying is a conversion of occasional into reliable, then the quantity you must know before buying is how occasional, on your prompts, under your grader. The guide's own two-stage recommendation, supervised fine-tuning as a warm start and then RL, is an admission that base success rate is the gating variable. Those are two different purchases with two different data requirements, and the histogram is what tells you which one you are in. Skip it and the likely outcome is a flat reward curve read as "this task is too hard for RL", when the real finding was that your prompts were all-or-nothing.

One limit on the evidence: the case studies in the guide report direction and not magnitude. Accuracy "rose", false positives were "sharply cut", queries came out "at closed-frontier quality". There are no baselines, no dataset sizes, no scores. They are a vendor's account of its early adopters, and nothing in them can be replicated by a reader.

What a learner should take is smaller than a technique and more durable. When a method learns from the model's own samples, the sampling distribution is the budget and the reward is only the spending rule. You cannot select your way to a behaviour that never appears in the group, and the fraction of your prompts where the model already disagrees with itself is the entire amount RL has to work with. An open-source trainer counts that fraction every step and writes it to the log. No console I can read shows it, and no announcement has a line for it, which is a reasonable description of where the interesting numbers in this field tend to live.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Best practices guide for customizing Gemini models via Reinforcement Learning (RL) Google Cloud · 2026-09-25
  2. GRPO Trainer documentation, main branch Hugging Face TRL · 2026-09-29
  3. GRPOTrainer advantage computation and reward logging, main branch Hugging Face TRL · 2026-09-29
  4. estimate_pass_at_k in the HumanEval evaluation harness OpenAI · 2026-09-29

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

reinforcement learningreward designpost-trainingpass@kreward hacking