Disclosure Without Overclaiming
What honest capability communication looks like, why published evaluation numbers mislead by default, and the specific claims that most often outrun their evidence.
Transparency documents are written by the people with the strongest interest in the system looking good, which is not a moral failing but a structural bias that shapes every artefact unless something counteracts it. The result is documentation that is accurate sentence by sentence and misleading in aggregate.
Where overclaiming happens
Benchmark numbers presented as capability. A score on a public benchmark is a measurement on that benchmark, and contamination, prompt tuning and metric selection all inflate it relative to what a user will experience. Reporting a benchmark result without the evaluation conditions, and without a result on held-out or private data, is technically true and practically misleading.
Selected operating points. A model evaluated at its best guidance scale, best decoding parameters, or best prompt reports the top of a distribution. Reporting the curve, or the configuration a user actually gets, is the honest version.
Comparisons against weak baselines. Beating a poorly tuned competitor is a fact about the tuning. The comparison a reader assumes is against a well-configured alternative, and papers and product pages routinely provide the other one.
Aggregate metrics hiding subgroup failure. As covered elsewhere, an aggregate number is compatible with a subgroup being badly served, so a headline figure without disaggregation overstates uniformity of performance.
Capability claims from demonstrations. A curated example establishes that the behaviour is possible, not that it is reliable. The distance between "can do" and "does reliably" is where most user disappointment originates.
What honest looks like
State the evaluation conditions, including data version, decoding configuration and whether the benchmark was public. Report the distribution rather than the best case. Disaggregate. Name the failure modes with examples, in the same document and with similar prominence to the capabilities. Distinguish measured results from expected behaviour explicitly. And where a limitation is known and unquantified, say that it is unquantified rather than omitting it.
The test worth applying is whether a reader who relied entirely on the document would be surprised by anything in their first week of use. Surprises are where the disclosure failed.
When it breaks
Nobody reads limitations sections. Placing caveats where they will not be encountered satisfies the letter of disclosure and not its purpose. Limitations that matter belong where the relevant capability is described, not collected in an appendix.
Legal review pushes toward vagueness. Precise limitations create commitments; vague ones create deniability. The pressure is real and it produces documents that are safe and useless, and resisting it requires someone whose job is the document's usefulness.
Competitive pressure sets the norm. If comparable products publish inflated numbers, publishing honest ones reads as a weaker product. This is a collective action problem that individual teams cannot solve, and it is an argument for standardised evaluation reporting rather than for matching the norm.
Overclaiming is discovered by users, not by reviewers. The gap between documented and actual behaviour surfaces in support tickets and public complaints, at which point the cost is trust rather than accuracy. Treating the first month's surprises as feedback on the documentation, rather than only on the product, is what closes the loop.
12 flashcards for this concept
Click a card to reveal the answer.