Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “statistics”

Tagged “statistics”

17 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Reasoning & Evaluation 24 min

Twenty Tests, One False Discovery: Multiple Testing From Bonferroni to the False Discovery Rate

A dead Atlantic salmon, scanned in 2009, showed 16 'active' voxels at p below 0.001; every procedure that controlled an error rate across the family found none. This is the argument over what that error rate should be, from Holm …

statistical-inference statistics experimentation ab-testing ∑ ◫
Reasoning & Evaluation 24 min

When Did the World Change? Changepoint Detection From Page's CUSUM to Bayesian Online Inference

CUSUM has been provably optimal since 1986, yet on the first human-annotated changepoint benchmark a detector that never reports a change beat most of the field under default settings. Seventy years of changepoint theory, from Pa…

anomaly-and-changepoint changepoint-detection anomaly-detection statistics ∑ ◫
Safety, Security & Governance 24 min

You Cannot Have All Three: COMPAS, Calibration and the Impossibility Theorems of Fair Classification

In 2016 ProPublica showed that COMPAS wrongly flagged 44.9% of Black defendants who never reoffended against 23.5% of white ones, and its vendor showed the scores meant the same thing for both groups. Both were right, and a one-l…

fairness-and-bias fairness bias calibration ∑ ◫
Reasoning & Evaluation 24 min

Your Improvement Is Inside the Noise: Seeds, Nondeterminism and the Reproducibility Problem in ML

Change one bit in one weight of a ResNet and, three epochs later, test accuracy differs by more than ten points. Training is a chaotic process, so a seed is not a control variable but a draw from a distribution. Most published an…

experiment-tracking-and-reproducibility statistics determinism mlops ∑ ◫
Reasoning & Evaluation 24 min

Your Test Set Is Wrong: Label Errors, Annotator Disagreement and the Ceiling on Measured Accuracy

Human reviewers confirmed 2,916 label errors in the ImageNet validation set, and an expert audit suggests the true figure is closer to one image in five. Every benchmark score is computed against an answer key written by people w…

human-data-and-annotation evaluation benchmarks measurement ∑ ◫
← Newer Page 2 of 2 · 17 posts
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N