Enterprise Generative Search — Azure and Open Source  ·  View 26 of 41  ·  Runtime

Grounding and Verification

How a fluent draft becomes claims that either hold, get re-retrieved, or are withdrawn.

Editable source SVG draw.io All views
Draft Model output structured, claim-marked Schema validation reject and retry once Decompose Claim segmentation one assertion per claim Numeric and date extraction checked literally Quote span capture for the citation anchor Align Claim to evidence NLI cross-encoder, per pair Numeric equality check deterministic, not model Cross-source agreement pairwise on claims Judge Supported score above 0.72 Unsupported no evidence covers it Contradicted sources disagree Act Cite and keep evidence id bound Re-retrieve once targeted at the gap Drop or abstain budget exhausted Show both positions dated and attributed Present Claim-level citations anchored to a passage Confidence and as-of stated, not implied Disagreement banner when sources conflict Provenance persisted replayable gap conflict Grounding and Verification — Turning a Draft Into Claims That Hold Application we own Decision point Risk / gap Data store failure / alternate Groundedness is measured against the evidence actually supplied, not against the world. A claim that is true but unsupported here is still dropped. v 1.0 · owner Data and AI Global Practice

Decisions

  • Groundedness is measured against the evidence actually supplied, not against the world. A claim that happens to be true but is unsupported here is still dropped, because the platform can only vouch for what it retrieved.
  • Numbers and dates are checked deterministically, not by a model. An NLI model will accept a transposed figure that a string comparison will not.
  • Contradiction is an output state. Both positions are shown, dated and attributed, rather than blended — which is the requirement that came out of the analyst journey in view 07.

Numbers

  • Groundedness gate 0.95, citation correctness 0.97, measured hallucination rate 1.1% against a 1.5% ceiling.
  • About 14% of drafts contain at least one unsupported claim; 9% are resolved by one targeted re-retrieval, 5% end in a dropped claim or an abstention.
  • Verification costs 180 ms and 0.0009 USD per answer.

Risks

  • Over-refusal is a real cost. False refusal rate is a gate in view 32 precisely so that grounding is not tightened until the product becomes unhelpful.
  • Verification uses a model, and models are wrong. The verifier is a different model family from the synthesiser, so a shared blind spot is less likely, and its disagreements are sampled for human review weekly.