A support assistant answers questions by retrieving from an internal wiki and a ticket archive and passing the result to a hosted model. Your model documentation template has a training-data section you cannot fill because you did not train the model. Which section should replace it?
Show the full answer Hide the answer
The deciding property
For a retrieval system, the data that determines behaviour is the corpus, not the weights. Change what is in the index and the assistant's answers change, with no deploy, no model change and no entry in the change log. The documentation section that earns its place is the one describing the thing that can move without anyone noticing.
What a corpus manifest has to carry
Per source, five fields, each of which has an engineering consequence:
- System of record and owner. Who can add a document that the assistant will then repeat to a customer.
- Access model at index time. This is where most designs are quietly wrong. A crawler running as a service account indexes everything it can read, and the retriever then serves chunks to users who could not open the original. Either filter at query time by the asking user's entitlements, or index separate partitions per entitlement set. The manifest has to say which.
- Refresh lag. A nightly crawl plus a one-hour response cache means a corrected wiki page can be contradicted by the assistant for up to about 25 hours. That single number is what a reviewer actually wants, and almost no documentation states it.
- Retraction path. When a document is deleted or legally retracted, the chain is source delete, index delete, cache purge, and the end-to-end lag is measurable. If the answer is "we rebuild the index weekly", the retraction time is a week and somebody senior should know that.
- Chunking and embedding version. Re-chunking or re-embedding changes retrieval behaviour as much as changing the model, so it belongs under change control.
Why the other options fail
- The provider's training sources and licensing. Useful for a procurement file and irrelevant to how this feature behaves. You cannot verify it, you cannot act on it, and it tells a reviewer nothing about the internal documents the assistant will actually quote.
- Fifty correct question and answer pairs. Evidence that the system worked on fifty inputs on one day. It is a test report, not documentation of what the system is, and it ages the moment the corpus changes.
- "Only sees documents the user may already read." This is a control claim, not a section of documentation, and in most implementations it is false, because the index was built by a service account. Stating it without the per-source access model is worse than stating nothing, because it closes the question a reviewer should have asked.
When this is the wrong answer
For a closed corpus of a forty-page handbook that changes twice a year, with no access control and nothing retractable, a manifest is paperwork. The trigger that makes it necessary is a corpus containing anything access-controlled, anything personal, or anything that can be legally withdrawn — because then staleness and retraction lag are obligations with numbers attached, and the only way to answer them in production is to have written the numbers down.