AI for Software Engineering advanced 8 min read 5 flashcards

LLM-Assisted Code Migration at Scale

The best-documented industrial win for code models is not feature development but mechanical migration, where change location is searchable, correctness is machine-checkable, and the work was previously too tedious to fund.

Google's migration of 32-bit integer IDs to 64-bit across its advertising codebase existed because of an arithmetic deadline: the identifier space was running out, and rollover would have caused outages. It is also the clearest published case of a code model paying for itself. In the reported programme, 80 percent of the code modifications in the resulting change lists were purely model-authored, and the team was tracking against a target of 50 percent or better acceleration relative to the manual estimate (Nikolov et al., 2025, How is Google using AI for internal code migrations?, ICSE-SEIP 2025, arXiv:2501.06972). A companion industry paper covering a different programme reports 39 distinct migrations run by three developers over twelve months, producing 595 submitted changes containing 93,574 edits, of which the model generated 74.45 percent of changes and 69.46 percent of edits, with developers estimating a 50 percent reduction in total migration time (Ziftci et al., 2025, Migrating Code At Scale With LLMs At Google, FSE 2025, arXiv:2504.09691).

Why this task fits the tool

Mechanical migration has three properties that almost no feature work has.

Change location is a search problem. The set of places to edit can be enumerated by static analysis, type information, or a call-graph query, so the model is never asked where the work is. Both Google papers treat location discovery as a separate component from generation, which is why the pipeline scales past what a context window could hold.

Correctness has a mechanical oracle. A migration is validated by compilation, the existing test suite, and often a differential comparison against the previous behaviour. The expensive human judgement that dominates ordinary review is replaced by a build result, which is what collapses the verification cost described in /learn/the-verification-cost-of-generated-code.

The change is repetitive but not uniform. This is the part that defeats a plain refactoring script. The edit is the same in intent and different in detail thousands of times over: a cast here, a signature there, a test fixture that has to change shape. Deterministic tooling handles the uniform cases; the long tail of near-misses is where a model earns its place, and where a human was previously required for every instance.

What the pipeline looks like

The published architecture is unglamorous and worth copying. A change-location pass produces candidate edit sites. A per-site prompt, carrying the surrounding code and migration-specific instructions, produces a candidate edit. A validation stage builds and tests, and discards or retries candidates that fail. Surviving edits are batched into reviewable change lists and routed to human owners for review and rollout. Both papers are explicit that review and rollout remained largely human-driven, and that the project-specific customisation lived in the prompts and the validation steps rather than in the model.

Two numbers from that pipeline deserve attention. The reported acceleration is around 50 percent after review and rollout are counted, which is much lower than the share of model-authored code, because the human stages did not speed up. And on one migration, 87 percent of model-generated code was committed without modification, which is a statement about the narrowness of the task rather than about general code quality.

When it breaks

Semantic migrations have no oracle. Moving from one logging library to another with different ordering guarantees, or from one concurrency model to another, cannot be validated by a build. Where the test suite does not encode the property that the migration might break, the model's output is unverified by construction, and the programme reverts to manual review cost.

Coverage gaps become silent breakage. The validation stage is only as strong as the test suite behind it. A migrated module with 30 percent branch coverage that compiles and passes is evidence of very little, and migrations tend to touch exactly the old, under-tested code that nobody wanted to work on.

The review stage sets the ceiling. Since generation is the part that got cheap, the programme's throughput is bounded by owner review of thousands of mechanical changes, which is the queueing problem in /learn/reviewing-machine-authored-changes. Google's reported pipeline addresses this by batching and routing rather than by asking for faster reviews.

It does not generalise to feature work by itself. The published results are from an environment with a monorepo, strong static analysis, uniform build tooling and dense ownership metadata. Each of those is load-bearing. A team without them is attempting a different and harder task with the same tool, which is the single most common misreading of these papers.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Nikolov et al., 2025, How is Google using AI for internal code migrations?, ICSE-SEIP 2025, arXiv:2501.06972 arxiv.org
  2. Ziftci et al., 2025, Migrating Code At Scale With LLMs At Google, FSE 2025, arXiv:2504.09691 arxiv.org
Check yourself

5 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track