140 prompts live as string literals across nine services. You must move them into a central prompt registry with no quality regression and no big-bang cutover, while every team keeps shipping. What is the sequence, and where can it go wrong?
Show the full answer Hide the answer
The sequence
Each step is independently reversible and each one produces evidence before the next begins.
- Inventory by observation, not by grep. Hash every prompt string at build time and emit the prompt id and hash as attributes on each model call. After two weeks of production traffic you know which prompts are live, how often, and in which code path. Expect a third to be dead. Migrating dead prompts is the cheapest work to cancel.
- Pin the model before touching the prompt. Replace floating model aliases with dated snapshots. OpenAI publishes dated snapshots such as
gpt-4o-2024-08-06alongside a published retirement date, so pinning removes silent drift and gives you a deadline you can plan for instead of a surprise. Migrating prompts while the model underneath can change makes every comparison meaningless. - Capture a golden set per live prompt from real traffic. 100 to 300 inputs, with today's outputs stored as the baseline. This is the expensive step and the one that cannot be skipped: without it, "no quality regression" is an assertion rather than a test.
- Ship the registry as a read-through cache with the in-code string as the fallback default. The compiled-in prompt stays in the binary. If the registry is unavailable or returns an unknown id, the service uses its local copy and emits a metric. A prompt registry that can take down nine services is a worse problem than the one being solved.
- Move one service at a time, byte-identical. No edits during the move. Verify in production that the hash served by the registry equals the hash compiled in, on every call, for a week. Only when that holds for a service does that service earn the right to edit prompts remotely.
- Gate edits. A change runs against the golden set, then canaries at 5 percent of traffic with the previous version still resolvable by id.
Where data and behaviour can diverge, and how you would know
The divergence that hurts is invisible: a prompt changed in the registry while the code that formats its variables did not. A template gaining a {tone} placeholder that the caller never fills renders the literal braces into the prompt, and the model usually copes, badly. Guard it by validating the template's variable set against the caller's supplied keys at resolution time and failing closed to the local copy.
The second divergence is caching. If the registry client caches for five minutes and instances restart at different times, two pods serve two different prompt versions simultaneously, which is fine for a rollout and catastrophic for an A/B measurement. Stamp the resolved version id on every logged interaction so any later analysis can split by it.
The point of no return
Deleting the in-code fallback strings. Until then every step rolls back with a deploy. After it, the registry is a hard runtime dependency and needs the availability posture of one: replication, a read path that survives its own control plane, and a documented break-glass.
How long it really takes
A quarter for 140 prompts across nine teams, and the golden sets dominate the effort, not the plumbing. Teams routinely budget two weeks for the registry and discover the work is in agreeing what "good" means for 90 live prompts.
When not to do this at all
Under roughly 20 prompts owned by one team, a reviewed file in the repository already gives versioning, review, rollback and diffing, with no new runtime dependency. The registry earns its cost when prompts are changed by people who cannot deploy, or when the same prompt is used by several services. Before that, it is a service you now operate in exchange for a feature Git already had.