Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “measurement”

Tagged “measurement”

7 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Safety, Security & Governance 24 min

Algorithmic Audits: What an Outside Examination of an AI System Can Actually Establish

Gender Shades measured a 34.4-point error gap on 1,270 faces and moved three vendors within seven months. New York City's mandatory bias audits produced 18 posted reports from 391 employers. The difference was not auditor skill b…

ai-assurance-and-audit evaluation governance compliance ∑ ◫
Reasoning & Evaluation 23 min

Clustering Has No Ground Truth: Impossibility, Validation, and What a Cluster Can Promise

In 2002 Jon Kleinberg proved that no clustering function can satisfy three properties almost everyone would ask for. Every algorithm is therefore a definition of what a cluster is, and every validation index is another definition…

clustering unsupervised-learning evaluation metrics ∑ ◫
Reasoning & Evaluation 25 min

Error Bars for Evals: Why Most Benchmark Differences Are Noise

A 250-question benchmark carries a standard error of about three percentage points. Most of the model comparisons published on top of such benchmarks cannot distinguish the models they are comparing. Evaluations are experiments, …

evaluation benchmarks statistics mlops ∑ ◫
Reasoning & Evaluation 24 min

The Bayesian Workflow: Why Fitting a Posterior Is the Easy Part

A 700-draw run of the eight-schools model reported R-hat of 1.01 and put the 2.5% quantile of the between-school spread at 0.89, while the exact posterior holds almost 18% of its mass below that value. Calling a sampler takes one…

bayesian-methods statistics uncertainty calibration ∑ ◫
Reasoning & Evaluation 24 min

Twenty Tests, One False Discovery: Multiple Testing From Bonferroni to the False Discovery Rate

A dead Atlantic salmon, scanned in 2009, showed 16 'active' voxels at p below 0.001; every procedure that controlled an error rate across the family found none. This is the argument over what that error rate should be, from Holm …

statistical-inference statistics experimentation ab-testing ∑ ◫
Reasoning & Evaluation 24 min

Your Improvement Is Inside the Noise: Seeds, Nondeterminism and the Reproducibility Problem in ML

Change one bit in one weight of a ResNet and, three epochs later, test accuracy differs by more than ten points. Training is a chaotic process, so a seed is not a control variable but a draw from a distribution. Most published an…

experiment-tracking-and-reproducibility statistics determinism mlops ∑ ◫
Reasoning & Evaluation 24 min

Your Test Set Is Wrong: Label Errors, Annotator Disagreement and the Ceiling on Measured Accuracy

Human reviewers confirmed 2,916 label errors in the ImageNet validation set, and an expert audit suggests the true figure is closer to one image in five. Every benchmark score is computed against an answer key written by people w…

human-data-and-annotation evaluation benchmarks measurement ∑ ◫
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N