A user uploads a policy document at 09:02 and it appears in the document list immediately. At 09:20 the assistant still says it cannot find anything on the subject. Error rates are flat and the ingest queue is empty. Where do you look first?
Show the full answer Hide the answer
The first three things I would look at, in order
- The per-document ingest state, not the queue. An empty queue is consistent with two opposite worlds: everything was processed, or the message was consumed and dropped. Ask the system "what is the state of document 8841" and expect an answer like
extracted → chunked → embedded → upserted → committed, with a timestamp per stage. - The dead-letter path and the skip counters. Extraction silently produces zero characters for a scanned PDF with no text layer, a password-protected file, or a format the parser does not handle. A pipeline that treats "zero chunks" as success leaves no error and no document.
- The index commit, not the upsert. Many vector stores acknowledge a write into a buffer and make it searchable only after a flush, segment merge or index build. A 200 on upsert is not a promise that a query will see the vector.
The diagnosis
The document list is served from the source-of-truth database and the assistant is served from a derived index with its own asynchronous write path, so "visible" and "findable" are two different states and the UI conflated them. The three candidates above account for nearly all of these reports, and one more deserves a check: the chunks were written with the wrong tenant or ACL metadata, so they exist and the query filter excludes them. That one presents identically and is the most dangerous, because the mirror-image bug exposes a document to the wrong reader.
The misleading signal
Flat error rates. Every stage returned success on the work it believed it had. A pipeline measured in "messages processed per minute" will look perfectly healthy while producing zero chunks per document, because the unit of success was the message rather than the document.
The fix
- Make ingest state visible to the user. "Indexed 09:04" next to the document, and "processing" until then. This removes the support ticket rather than the bug, and it is the cheapest change on the list.
- Instrument the lag that matters:
indexed_at − uploaded_at, p50 and p95, with an SLO. For a corpus of 1.8M documents averaging four chunks each, a p95 under 60 seconds is an ordinary target; alert when p95 crosses it or when the oldest unindexed document exceeds 10 minutes. - Alert on zero-chunk documents as a failure class, with the dead-letter queue depth paged rather than dashboarded.
- Run a canary document per tenant, re-indexed hourly and queried end to end, so filter-and-ACL regressions surface without a customer finding them.
When this is the wrong answer
If the product genuinely promises read-after-write — "upload a contract and ask about it immediately" — then monitoring the lag is not enough. Make the retriever read the source text for documents still in flight, straight from the extraction output for that one document, and accept the cost. Do not try to make an asynchronous index synchronous; that trade only ends in a rebuilt index on the request path.