Safety & Alignment 3 October 2026 8 min read 1,819 words

The waypoints never read the sentence

NVIDIA's open Alpamayo driving models write out the reason for each manoeuvre and the documentation calls the result auditable. The module that chooses the manoeuvre cannot read that text, and the alignment between the two is something the training produced.

The argument

A Chain-of-Causation trace is optimised to agree with the trajectory rather than to have produced it, so the property the documentation sells as auditability is the one this design removes: only an explanation free to contradict the action could ever expose a wrong reason.

Alpamayo 2 Super is a 34-billion parameter driving model that writes down why it is about to do something. Thirty-two billion of those parameters are a vision-language backbone that reads multi-camera video and emits a passage of text NVIDIA calls Chain-of-Causation: the cyclist drifting out, the brake lights three cars ahead, the reason for easing off. The other two billion are a diffusion expert that produces the thing the car would actually execute, a trajectory of 64 waypoints covering the next 6.4 seconds. Both come out of one inference call. NVIDIA's developer hub says what the pairing is for: CoC traces "accompany every predicted trajectory, making each driving decision transparent and auditable."

Read the verb. They accompany.

The two outputs leave the model through different channels, and the text is not the channel the waypoints are computed from. In Alpamayo 1.5, the smaller 10B sibling, the documented inference path is explicit about the order of operations: "the VLM generates chain-of-causation reasoning, then a diffusion expert produces trajectory predictions conditioned on the VLM's hidden states." Conditioned on the hidden states. Not on the sentence. The sentence is a separate decoding of the same activations, and nothing in the architecture requires the words that get sampled to name the part of the state the planner leaned on. My argument is that this makes the trace a narration optimised to agree with the trajectory rather than a record of what produced it, and that auditability is precisely the property such a design cannot deliver.

What the action expert can and cannot read

It helps to be concrete about the plumbing, because the plumbing is the argument. A vision-language-action model is three things stacked: an encoder that turns camera frames into tokens, a language model that consumes those tokens along with navigation input and history, and an action head that turns the resulting internal state into continuous control. Alpamayo 1 and 1.5 use a Cosmos-Reason backbone with an action expert bolted on; Alpamayo 2 Super pairs a 32B backbone with a 2B diffusion expert and, by its own description, generates the CoC text first and then samples trajectories through the expert.

The action head is not a reader of text. An implementation proposal filed against the vLLM-Omni project by a contributor working out how to serve Alpamayo 1.5 describes the module as built "from a deep-copied Qwen3 text config with embed_tokens removed." The token embedding table, the lookup that converts a word id into a vector, has been deleted from it. The expert stage instead has to "consume the KV-cache produced by the AR worker and append n_diffusion_tokens = 64 new queries on top of the fixed prefix." It attends to the keys and values the reasoning pass left in the cache and denoises 64 waypoints against them.

So the headline is literal rather than figurative. The waypoints never read the sentence. They read the cache the sentence was also decoded from. Those are different objects: a hidden state is a few thousand dimensions per position, and the token sampled from it is one choice out of a vocabulary. The decode is lossy and, more importantly, it is underdetermined. Many different sentences are compatible with the same activations, and the sampler picks one. An explanation and a decision drawn from a shared cause can agree, disagree, or agree for the wrong reason, and the architecture does not adjudicate between those cases.

Where the reasons came from in the first place

The second half of the mechanism is the supervision, and it is the part most likely to surprise someone who assumed these traces were human-written. NVIDIA has published the pipeline that manufactures them. The CoC autolabeler runs four stages. First it produces "per-clip high-level motion labels from trajectory data" - the meta-action, derived from what the vehicle did. Second it selects "frames where ego meta-actions change, since these transitions are likely to contain decision-making context." Third it runs a vision-language model over those keyframes "to produce chain-of-causation labels." Fourth, optionally, it runs further VLM passes to flag "potentially lower-quality labels."

Look at the direction of travel. The manoeuvre is extracted from the recorded trajectory. The frames are chosen because the manoeuvre changed there. Then a model is shown those frames and asked what caused the change it already implicitly knows about. The label is a rationalisation of a known outcome, generated after the fact, by a model with no access to whatever actually produced the human driver's behaviour. The labelling model is Qwen 3.5 or 3 locally, or GPT-5 or GPT-5.5 through an API.

NVIDIA is candid about the failure modes. Because the pipeline relies on VLMs, it says, the generated outputs "may contain errors, including incorrect maneuver attribution (for example, right vs. left lane change), hallucinated objects, or inaccurate temporal-causal reasoning about surrounding agents and ego behavior." It recommends human auditing, and states the tool "is not a production-ready system." The model comparison table confirms the recipe reaches the weights: Chain-of-Causation reasoning in both Alpamayo 1 and 1.5 is "hybrid auto-labeling with human in the loop for reasoning traces." A human reviews. A model writes.

Train on that corpus and you get a model fluent in a genre. The genre is plausible driving commentary, scored during data generation on whether it reads well against a keyframe, not on whether it identifies a cause.

The tell is in the feature table

Which brings us to the sentence that gives the whole thing away, and it is NVIDIA's own. The Alpamayo 1.5 comparison table lists a capability the earlier model lacked, and describes it in five words: "Reinforcement learning for reasoning/action consistency." The FAQ repeats it in longer form: the model "has undergone RL post-training, achieving improvements in reasoning quality and reasoning-trajectory alignment as a result."

Consistency between the reasoning and the action is a training objective. Sit with that. You do not train a system for consistency between a cause and its effect, because that relation is not something a reward can improve; it either holds by construction or the thing you called a cause was never one. The presence of an alignment objective is an admission that the trace and the trajectory could come apart, and the reported success of that objective means the gap has been closed from the outside, by optimisation, rather than from the inside, by the text being load-bearing.

This is where the auditability claim fails, and it fails in an unusual way. The problem is not that the explanations are bad. They may well be good, in the sense of being accurate, specific and useful. The problem is that their agreement with the action has been directly optimised, so agreement carries no information. An explanation earns its place in an investigation through its capacity to indict: the value of a flight recorder is that it can contradict the pilot. A trace trained until it stops disagreeing with the trajectory has had exactly that capacity removed. The better the alignment number, the less an investigator can learn from a trace that matches.

The strongest case for the other side

The serious objection is that hidden-state conditioning is more faithful than the alternative, not less, and this is true. A system where the planner read only the emitted text would be worse in an obvious way: it would throw away almost everything the backbone computed, and it would invite the model to write a sentence and then act on a thinner version of its own reasoning. Conditioning on the cache keeps the channel wide. The trace and the trajectory are not independent stories told by separate systems; they are two readouts of one computation, which already rules out the crudest kind of confabulation.

A second objection is practical. Even an imperfect trace is enormously useful. It gives engineers a searchable key over millions of miles of logs, it clusters failures by stated reason, it surfaces hallucinated objects, and it makes a policy's behaviour discussable in review. None of that requires the trace to be causal. NVIDIA also builds the closed-loop pieces, AlpaSim and AlpaGym, which is where a driving policy's claims get tested by intervention rather than narration.

Both points land. Neither rescues the word auditable. A wide channel cuts both ways: because the expert attends to the cache rather than the tokens, the breadth of what it reads is exactly what guarantees the decoded sentence need not name the operative part. And the operational value of traces is the reason the stronger claim gets made on their back. Transparent and auditable is a regulatory register, not an engineering one, and it will be quoted in rooms where someone is deciding whether a stated reason counts as evidence.

One limit on the above is worth stating plainly. Everything here is read off NVIDIA's own published repositories and one contributor's implementation proposal, because the model cards sit behind a gated host I could not open in this session. If a faithfulness measurement exists for these traces, something that tests whether the stated reason moves the trajectory, I did not see it, and its absence from the repositories is part of why the claim deserves pressure rather than the benefit of the doubt.

What to take from this

Faithfulness is a property of a channel, not of a style. When you meet a system that reasons and then acts, ask one question before you ask anything else: what does the action head receive? If the answer is the text, the trace is at least in the causal path. If the answer is hidden states, activations, or a KV cache, then the trace is a sibling of the decision rather than its parent, and no amount of fluency changes that. Ask the question of agents too, where the same shape is now standard.

Then stop treating agreement as evidence and start treating disagreement as the measurement. The informative experiment is interventional: perturb the scene so the stated reason no longer holds and see whether the trajectory moves. Suppress the cyclist and watch whether the swerve survives. A trace that tracks those interventions has earned something. A trace that merely matches the waypoints has told you only that it was trained to.

I would not want the traces removed. They are the most useful thing about these models for anyone trying to understand a fleet. But the documentation has put a load-bearing word on a component that cannot carry it, and the fix is a sentence, not an architecture. Call them commentary and they are a gift. Call them an audit trail and the first serious accident will produce a beautifully written reason that nobody can show had anything to do with the steering.

What this is argued from

Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.

  1. Alpamayo 2 Super NVIDIA (NVlabs) on GitHub · 2026-10-03
  2. NVIDIA Alpamayo Developer Hub NVIDIA (NVlabs) on GitHub · 2026-10-03
  3. Alpamayo 1.5 NVIDIA (NVlabs) on GitHub · 2026-10-03
  4. Chain-of-Causation (CoC) Autolabeling Pipeline NVIDIA (NVlabs) on GitHub · 2026-10-03
  5. Alpamayo 1 Nano NVIDIA (NVlabs) on GitHub · 2026-10-03
  6. "[RFC]: Support NVIDIA Alpamayo 1.5 (Reasoning VLA for Autonomous Driving) in vLLM-Omni" vLLM project on GitHub · 2026-04-17

Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.

vision-language-actionreasoning faithfulnessautolabelingauditabilitydiffusion policy