Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “evaluation”

Tagged “evaluation”

31 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Reasoning & Evaluation 3 min

What the bake-off taught us: classical ML is not dead, it is just under-attended

We pitted twelve sklearn algorithms head-to-head on a tabular dataset. The winner was not the most expensive one. It was not the most modern one. It was the one whose assumptions matched the data.

machine-learning tabular benchmarks evaluation
Reasoning & Evaluation 20 min

When the Judge Is Also a Player: LLM-as-Judge, Contamination, and Why Leaderboards Drift

A strong model grading other models looks like a free lunch for evaluation. It is not. Position, verbosity, and self-preference biases plus quietly leaked test sets mean a leaderboard number can move several points without any mo…

evaluation llm-as-judge benchmarks contamination ∑ ◫
Safety, Security & Governance 24 min

You Cannot Have All Three: COMPAS, Calibration and the Impossibility Theorems of Fair Classification

In 2016 ProPublica showed that COMPAS wrongly flagged 44.9% of Black defendants who never reoffended against 23.5% of white ones, and its vendor showed the scores meant the same thing for both groups. Both were right, and a one-l…

fairness-and-bias fairness bias calibration ∑ ◫
Reasoning & Evaluation 24 min

Your Improvement Is Inside the Noise: Seeds, Nondeterminism and the Reproducibility Problem in ML

Change one bit in one weight of a ResNet and, three epochs later, test accuracy differs by more than ten points. Training is a chaotic process, so a seed is not a control variable but a draw from a distribution. Most published an…

experiment-tracking-and-reproducibility statistics determinism mlops ∑ ◫
Reasoning & Evaluation 24 min

Your Test Set Is Wrong: Label Errors, Annotator Disagreement and the Ceiling on Measured Accuracy

Human reviewers confirmed 2,916 label errors in the ImageNet validation set, and an expert audit suggests the true figure is closer to one image in five. Every benchmark score is computed against an answer key written by people w…

human-data-and-annotation evaluation benchmarks measurement ∑ ◫
Reasoning & Evaluation 14 min

Your eval pipeline is the moat, not your model choice

The model layer is commoditising and the answer flips every six months. The only durable advantage is the ability to A/B a model swap end-to-end in 48 hours and know whether it improved things for your users.

evaluation benchmarks llm-as-judge reliability
Reasoning & Evaluation 28 min

Zero-Shot Forecasting: What Time-Series Foundation Models Actually Learned

A 35-million-parameter model with no knowledge of your business can forecast your demand as well as the pipeline your team spent two quarters building. That is a real result and it is routinely misread. Pretrained forecasters lea…

time-series forecasting foundation-models zero-shot ∑ ◫
← Newer Page 3 of 3 · 31 posts
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N