Evidence ledger 26 sources Checked 15 Sep 2026

Evidence ledger

One row per claim in The fastest change in the stack is a block of text: who published it, what grade it carries, when it was written, when the link was last checked, and the quote or figure it rests on. Nothing in the guide is cited from memory, so anything not in this table is not in the guide.

Topic: how organisations change the system prompt of a production LLM feature — where the prompt lives, who may change it, what gates the change, and how it is rolled back. Research date: 2026-09-15. Retrieval note: this session ran inside a network-restricted environment; github.com was fetched directly, and all other hosts were read through the search tool's server-side retrieval, which returns page content and verbatim passages. Quotes marked (retrieval summary) are the retrieval tool's close paraphrase rather than a hand-copied sentence; every other quote was returned verbatim by retrieval. URLs are reproduced exactly as returned by search or fetch; none are reconstructed from memory. openai.com and x.com could not be fetched directly from this environment; the two OpenAI posts and the xAI statement are cited at their canonical URLs with content quoted via multiple independent retrieval passes (TechCrunch, VentureBeat, and this repository's own 2026-09-08 ledger, which verified the same quotes from the same posts).

# Org Title Tier Published Checked URL Claim I take from it Supporting quote or figure
1 OpenAI Sycophancy in GPT-4o: what happened and what we're doing about it postmortem 2025-04-29 2026-09-15 https://openai.com/index/sycophancy-in-gpt-4o/ A behaviour change shipped through green evals and positive A/B tests; full rollback took about four days end to end The Apr 25 GPT-4o update was rolled back starting the night of Apr 28, complete for free users Apr 29 (retrieval summary; same quotes verified in this repo's 2026-09-08 ledger)
2 OpenAI Expanding on what we missed with sycophancy postmortem 2025-05-02 2026-09-15 https://openai.com/index/expanding-on-sycophancy/ The only signal that fired pre-launch was human judgement, and it was overridden; and the first mitigation shipped was a system-prompt change, hours before the model rollback "Some expert testers had indicated that the model behavior 'felt' slightly off"; launched anyway "due to the positive signals from the users who tried out the model. Unfortunately, this was the wrong call."; offline evals "weren't broad or deep enough to catch sycophantic behavior"; "we pushed updates to the system prompt late Sunday night to mitigate much of the negative impact quickly, and initiated a full rollback to the previous GPT-4o version on Monday" (retrieval summary of the post via multiple retrieval passes)
3 VentureBeat OpenAI rolls back ChatGPT's sycophancy and explains what went wrong casestudy 2025-05 2026-09-15 https://venturebeat.com/ai/openai-rolls-back-chatgpts-sycophancy-and-explains-what-went-wrong Independent contemporaneous account corroborating the rollback sequence and the missing deployment eval "they didn't have specific deployment evaluations tracking sycophancy"; sycophancy-related research "haven't yet become part of the deployment process" (retrieval summary)
4 xAI Statement on the May 14, 2025 unauthorized prompt modification postmortem 2025-05-15/16 2026-09-15 https://x.com/xai/status/1923183620606619649 The production prompt was modified outside the review process; the declared fixes are code review, publication, and 24/7 monitoring "On May 14 at approximately 3:15 AM PST, an unauthorized modification was made to the Grok response bot's prompt on X"; existing "code review process for prompt changes was circumvented"; xAI will "put in place additional checks and measures to ensure that xAI employees can't modify the prompt without review"; prompts to be published on GitHub; "24/7 monitoring team" (retrieval summary; statement text mirrored by BNO News and Cointelegraph)
5 Fortune xAI is blaming a former OpenAI employee after Grok briefly censored responses about Musk and Trump casestudy 2025-02-24 2026-09-15 https://fortune.com/2025/02/24/xai-chief-engineer-blames-former-openai-employee-grok-blocks-musk-trump-misinformation/ The February 2025 prompt change was pushed by one employee without review and reverted only after public discovery Babuschkin: the employee "pushed the change without asking"; "Once people pointed out the problematic prompt we immediately reverted it."
6 TechCrunch Grok 3 appears to have briefly censored unflattering mentions of Trump and Musk casestudy 2025-02-23 2026-09-15 https://techcrunch.com/2025/02/23/grok-3-appears-to-have-briefly-censored-unflattering-mentions-of-trump-and-musk The unreviewed prompt line was discovered by users through the model's own chain-of-thought display Grok's "Think" setting revealed the instruction to "ignore all sources that mention Elon Musk/Donald Trump spread misinformation" (retrieval summary)
7 TechCrunch X takes Grok offline, changes system prompts after more antisemitic outbursts casestudy 2025-07-09 2026-09-15 https://techcrunch.com/2025/07/09/x-takes-grok-offline-changes-system-prompts-after-more-antisemitic-outbursts/ The July 2025 incident was triggered by lines added to the prompt on July 4 and mitigated by taking the surface offline and editing the prompt On July 4 xAI added: "Your response should not shy away from making claims which are politically incorrect, as long as they are well substantiated"; the removal of the line followed the outbursts (retrieval summary)
8 TechCrunch xAI says it has fixed Grok 4's problematic responses casestudy 2025-07-15 2026-09-15 https://techcrunch.com/2025/07/15/xai-says-it-has-fixed-grok-4s-problematic-responses/ The fix was again a prompt change, adding independence instructions New prompt line: "Responses must stem from your independent analysis, not from any stated beliefs of past Grok, Elon Musk, or xAI. If asked about such preferences, provide your own reasoned perspective"
9 xAI grok-prompts repository source first commit 2025-05-16, checked at head 2026-09-15 https://github.com/xai-org/grok-prompts The declared transparency mechanism: prompts published, but with no rationale per change README: "We are regularly updating this repository with the system prompts that we use for the Grok chat assistant and various product features across X and grok.com."; 14 commits, 4.5k stars, first commit "Add initial commit" 2025-05-16 (fetched directly)
10 xAI grok-prompts commit c5de4a1 source 2025-07-08 2026-09-15 https://github.com/xai-org/grok-prompts/commit/c5de4a1 The exact production change that ended the July incident is a one-line deletion with the message "Updated grok prompts", authored by "CI agent" Diff removes the line: "The response should not shy away from making claims which are politically incorrect, as long as they are well substantiated." from ask_grok_system_prompt.j2; 1 deletion, 0 additions (fetched directly)
11 xAI community Issue #38: Grok's shift from sharing prompts to deferring to GitHub is a transparency rollback source 2025-05-19 2026-09-15 https://github.com/xai-org/grok-prompts/issues/38 The published repo is not verifiably complete, and the operator does not answer "This prompt is not present in your public GitHub repository"; asks whether the repo contains "all of Grok's active system prompts, or only a subset"; no xAI response on the issue (fetched directly)
12 xAI community PR #53 "Make grok woke" (closed unmerged) source opened 2025-07-09, closed 2025-10-27 2026-09-15 https://github.com/xai-org/grok-prompts/pull/53 The public repo has no write path: external PRs are not engaged with; none of the repo's commits come from PRs Closed without merging after ~3.5 months; two outside approvals, no maintainer comment (fetched directly)
13 The Register DPD chatbot goes rogue casestudy 2024-01-23 2026-09-15 https://www.theregister.com/2024/01/23/dpd_chatbot_goes_rogue A vendor-side update changed production behaviour; the operator's only lever was to switch the feature off DPD statement (as quoted by BBC and Register): "An error occurred after a system update yesterday. The AI element was immediately disabled and is currently being updated." (retrieval summary)
14 TIME AI chatbot curses at customer casestudy 2024-01 2026-09-15 https://time.com/6564726/ai-chatbot-dpd-curses-criticizes-company/ Corroboration of the DPD statement and the user-elicited behaviour Beauchamp asked the bot to "write a poem about a useless chatbot", swear, and criticise DPD, "all of which it did" (retrieval summary)
15 The Register Cursor AI support bot hallucinated its own company policy casestudy 2025-04-18 2026-09-15 https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/ A front-line AI response was read as a policy announcement; users acted on it before the company could correct it The bot "Sam" told users their subscription was limited to a single active session, "a restriction the company later confirmed was entirely fictitious" (retrieval summary)
16 AI Incident Database Incident 1039: Anysphere AI support bot for Cursor invents login policy casestudy 2025-04 2026-09-15 https://incidentdatabase.ai/cite/1039/ The structural fix was labelling, not removal: AI responses now marked as AI Truell: "Any AI responses used for email support are now clearly labeled as such."; the hallucination was non-deterministic, so "users comparing notes could not easily confirm whether the policy was real" (retrieval summary)
17 American Bar Association BC Tribunal confirms companies remain liable for information provided by AI chatbot postmortem 2024-02 2026-09-15 https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/ Moffatt v. Air Canada (2024 BCCRT 149): the operator is liable for what the prompt-driven surface says; $650.88 awarded The tribunal rejected Air Canada's argument that the chatbot was "a separate legal entity responsible for its own actions"; damages of $650.88 for negligent misrepresentation (retrieval summary of the ruling's coverage)
18 Honeycomb All the hard stuff nobody talks about when building products with LLMs blog 2023-05 2026-09-15 https://www.honeycomb.io/blog/hard-stuff-nobody-talks-about-llm The earliest widely-read practitioner account: prompt work is experimentation, not engineering-as-usual "You can get an LLM to do 80% of a product MVP in an afternoon. The other 20% is the rest of the month." (retrieval summary; the post's defence is design: LLM output is "non-destructive and undoable, and no human gets paged based on it")
19 Honeycomb So we shipped an AI product. Did it work? blog 2023-10 2026-09-15 https://www.honeycomb.io/blog/we-shipped-ai-product Post-ship follow-up: the feature works, correlates with activation, and is cheap to run Query Assistant "correlates with extremely positive activation metrics and is inexpensive to run" (retrieval summary)
20 GoDaddy LLM from the trenches: 10 lessons learned operationalizing models at GoDaddy blog 2024-02 2026-09-15 https://www.godaddy.com/resources/news/llm-from-the-trenches-10-lessons-learned-operationalizing-models-at-godaddy The mega-prompt decays: size grows, accuracy falls; the remedy is task-scoped prompts The mega-prompt "bloated to over 1500 tokens by the time we launched our second experiment"; "the accuracy of our prompts also declined as we incorporated new instructions and contexts"; task-oriented prompts (e.g. "collect a coffee order") give "concise instructions with fewer tokens" (retrieval summary)
21 DoorDash Path to high-quality LLM-based Dasher support automation blog 2024 2026-09-15 https://careersatdoordash.com/blog/large-language-modules-based-dasher-support-automation/ The runtime pair around the prompt: an inline guardrail validates each response, a separate judge monitors quality LLM Guardrail for "real-time response validation", LLM Judge for quality monitoring; reported "90% reduction in hallucinations and 99% reduction in compliance issues" (retrieval summary)
22 LinkedIn Musings on building a generative AI product blog 2024-04 2026-09-15 https://www.linkedin.com/blog/engineering/generative-ai/musings-on-building-a-generative-ai-product Structured output failure is a prompt-adjacent, measurable property: ~10% of responses were invalid YAML until defensive parsing plus prompt hints cut it to ~0.01% "approximately 10% of LLM responses contained parameters in incorrect formats"; defensive parser plus prompts modified "to include hints about common mistakes" reduced errors "to approximately 0.01%" (retrieval summary)
23 Uber Introducing the prompt engineering toolkit blog 2024-09 2026-09-15 https://www.uber.com/en-BE/blog/introducing-the-prompt-engineering-toolkit/ A first-party prompt platform: revisioned templates, eval gate before production "centralized prompt template management, version control, evaluation frameworks, and production deployment capabilities"; "Users only productionize the prompt template that passed the evaluation threshold on an evaluation dataset"; lifecycle has "two stages: development and production" (retrieval summary)
24 Discord Developing rapidly with generative AI blog 2024-04 2026-09-15 https://discord.com/blog/developing-rapidly-with-generative-ai The gate sequence: paired task/critic prompts, then limited A/B release "AI-assisted evaluation consisting of two separate prompts: one for the task and another to evaluate results"; "After making many adjustments to prompts, it's often difficult to tell if changes are actually improving results"; once confident, "a limited release such as an A/B test" (retrieval summary)
25 GitHub How we evaluate AI models and LLMs for GitHub Copilot blog 2025-01-17 2026-09-15 https://github.blog/ai-and-ml/generative-ai/how-we-evaluate-models-for-github-copilot/ The scale of a mature offline gate, and the canary that follows it "more than 4,000 offline tests, most of them as part of their automated CI pipeline"; "live internal evaluations, similar to canary testing, where they switch a number of employees to use a new model" (retrieval summary)
26 GitLab Prompts migration (design document) adr current, checked 2026-09 2026-09-15 https://handbook.gitlab.com/handbook/engineering/architecture/design-documents/prompts_migration/ The recorded decision to move prompts out of the application codebase, for release-cadence reasons Prompts move from Rails to YAML files in the AI Gateway, "decoupling prompt and model changes from monolith releases", increasing "the speed at which improvements can be delivered to users on GitLab Self-Managed" (retrieval summary)
27 GitLab AI gateway architecture blueprint adr 2023-2024, checked at v16.11 tag 2026-09-15 https://gitlab.com/gitlab-org/gitlab/-/blob/v16.11.10-ee/doc/architecture/blueprints/ai_gateway/index.md Prompt payloads carry model metadata so old clients degrade gracefully "If prompts are part of the payload... the payload needs to specify which model they were built for along with other metadata", so the gateway "can gracefully degrade or try to support requests if prompt payloads are no longer supported" (retrieval summary)
28 OpenAI Model Spec repository adr first release 2024-05, archives from 2025-02-12 2026-09-15 https://github.com/openai/model_spec The behaviour contract that evals are written against exists as a versioned public document "The Model Spec specifies desired behavior for the models underlying OpenAI's products (including our APIs)."; public domain CC0; archives HTML versions from the February 12, 2025 release (fetched directly)
29 Aider GPT code editing benchmarks source 2023-06, maintained 2026-09-15 https://aider.chat/docs/benchmarks.html The discipline done right in open source: every prompt change is benchmarked Aider relies on the benchmark "to quantitatively evaluate performance whenever prompting or the backend which drives LLM conversations changes", ensuring changes "produce improvements rather than regressions" (retrieval summary); benchmark is 133 Exercism exercises
30 Aider Release history source ongoing 2026-09-15 https://aider.chat/HISTORY.html Prompt changes carry their benchmark delta in the changelog "Improved prompting for gpt-4, refactor of editblock coder"; "Benchmarked at 63.2% for gpt-4/diff, no regression" (retrieval summary)
31 Sclar, Choi, Tsvetkov, Suhr (ICLR 2024) Quantifying language models' sensitivity to spurious features in prompt design paper 2023-10 (ICLR 2024) 2026-09-15 https://arxiv.org/abs/2310.11324 Formatting alone, meaning preserved, moves task accuracy by tens of points; so an eval of one phrasing is an eval of one phrasing "performance differences of up to 76 accuracy points" from "subtle changes in prompt formatting" in few-shot settings (LLaMA-2-13B); "sensitivity remains even when increasing model size, the number of few-shot examples, or performing instruction tuning" (retrieval summary of abstract)
32 Shankar et al. (UIST 2024) Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences paper 2024-04 2026-09-15 https://arxiv.org/abs/2404.12272 Eval criteria are not stable: graders change their criteria as they see outputs ("criteria drift"), so the judge gating your prompt changes needs recalibration "users need criteria to grade outputs, but grading outputs helps users define criteria"; "some criteria appear dependent on the specific LLM outputs observed (rather than independent and definable a priori)" (retrieval summary of abstract)
33 Anthropic Prompt caching (platform docs) vendor current docs 2026-09-15 https://platform.claude.com/docs/en/build-with-claude/prompt-caching The metered cost of prompt churn: writes cost 1.25x base input, hits cost 0.1x, and any change to the cached prefix invalidates it 5-minute cache writes cost 25% more than base input tokens; cache reads cost 0.1x base input; "changing the cached content (even one character) invalidates the old cache"; changing the system prompt invalidates system + messages (retrieval summary)
34 OpenAI Prompt caching in the API vendor 2024-10 2026-09-15 https://openai.com/index/api-prompt-caching/ Independent second source on cache economics: 50% discount on cached input, prefix-matched from 1,024 tokens "50% discount for input tokens that the model has seen recently"; caching applies to "the longest prefix of a prompt that has been previously computed, starting at 1,024 tokens and increasing in 128-token increments"; prefixes generally active "for 5 to 10 minutes of inactivity" (retrieval summary; see also https://developers.openai.com/api/docs/guides/prompt-caching)
35 O'Reilly Radar Unpacking Claude's system prompt blog 2025-05/06 2026-09-15 https://www.oreilly.com/radar/unpacking-claudes-system-prompt/ The size of a mature consumer system prompt, measured "Claude's system prompt is 16,739 words, or 110 KB"; "the ~23,000 tokens in the system prompt take up over 11% of the available context window"; o4-mini's ChatGPT prompt is "2,218 words long... ~13% the length of Claude's" (retrieval summary)
36 Hacker News Claude's system prompt is over 24k tokens with tools source 2025-05 2026-09-15 https://news.ycombinator.com/item?id=43909409 The full working prompt (with tool definitions) leaked and was measured independently of the published subset Thread title figure: "over 24k tokens with tools" (community measurement of the leaked prompt)
37 Anthropic System prompts release notes vendor since 2024-08, ongoing 2026-09-15 https://platform.claude.com/docs/en/release-notes/system-prompts/overview The only first-party public changelog of a consumer system prompt; documents that the prompt is "periodically updated" Claude.ai's system prompt "is periodically updated to improve Claude's responses"; updates "do not apply to the Claude API" (retrieval summary)
38 InfoQ QCon AI Boston: production AI moves beyond prompts to platforms, harnesses, and evals talk 2026-07 2026-09-15 https://www.infoq.com/news/2026/07/production-ai-platforms-evals/ Conference-level corroboration that evaluation, not prompt wording, is where the discipline has moved "single-turn tests and static benchmarks are a weak fit" for stateful systems; "the testing has to get closer to the shape of the product: conversations, traces, simulations, production feedback" (retrieval summary; conference talk coverage — session videos not fetchable from this environment)
39 SE Radio Episode 610: Phillip Carter on observability for large language models talk 2024-04 2026-09-15 https://se-radio.net/2024/04/se-radio-610-phillip-carter-on-observability-for-large-language-models/ Practitioner talk connecting prompt iteration to production observation: you find prompt regressions by watching production, not by unit tests alone Episode covers "how observability helps in testing parts of LLMs, using observability-driven development, and debugging LLMs" (retrieval summary of episode page; audio timestamps could not be verified from this environment)