Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Reasoning & Evaluation

Reasoning & Evaluation

Models that think for longer, and the measurement problems that come with them.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Reasoning & Evaluation 27 min

The Last Undefeated Baseline: Why Gradient-Boosted Trees Still Beat Deep Learning on Tabular Data

Deep learning took images in three years and text in five. It has been attacking tabular data since 2016 and still has not won. The reason is not compute or architecture; it is three specific inductive biases that make a tree the…

tabular gradient-boosting deep-learning xgboost ∑ ◫
Reasoning & Evaluation 24 min

The Leaderboard Is Not Your Corpus: Why Top-Ranked Embedding Models Disappoint in Production

Embedding models are chosen from a leaderboard more often than from an experiment, and the leaderboard now publishes training splits for its own test sets. Between contamination, task-family averaging and geometry no benchmark me…

embeddings retrieval-rag evaluation benchmarks ∑ ◫
Reasoning & Evaluation 23 min

The Model Knows It Is Being Tested: Evaluation Awareness and the Limits of Behavioural Safety Evidence

On one synthetic honeypot evaluation, Claude Sonnet 4.5 said out loud that it suspected it was being tested in 80 to 100 percent of transcripts, against under 10 percent for its predecessor. When the internal representations behi…

safety-alignment evaluation interpretability red-teaming ∑ ◫
Reasoning & Evaluation 23 min

The Progress Illusion in Recommender Systems: Weak Baselines, Sampled Metrics and Leaky Splits

In 2019 a careful team could reproduce only 7 of 18 neural recommenders from top venues, and 6 of those 7 lost to nearest-neighbour heuristics. The models were not the problem. The protocol was: untuned baselines, metrics compute…

recommender-systems evaluation metrics benchmarks ∑ ◫
Reasoning & Evaluation 13 min

The reasoning-model bubble: when test-time compute stops paying

o3, R1 and Claude extended thinking are a real capability shift on a narrow slice of tasks. They are also being shoved into product surfaces that punish every property reasoning models exhibit - and the bill is starting to arrive.

reasoning test-time-compute economics evaluation ∑
Reasoning & Evaluation 24 min

Twenty Tests, One False Discovery: Multiple Testing From Bonferroni to the False Discovery Rate

A dead Atlantic salmon, scanned in 2009, showed 16 'active' voxels at p below 0.001; every procedure that controlled an error rate across the family found none. This is the argument over what that error rate should be, from Holm …

statistical-inference statistics experimentation ab-testing ∑ ◫
Reasoning & Evaluation 24 min

What a Feature Attribution Can and Cannot Tell You: SHAP, LIME and the Explanation Gap

Add a column the model never reads and SHAP can hand it more than a quarter of the credit for a decision. That is not a library bug: a Shapley attribution answers a question you chose, often without noticing, and a 2024 PNAS resu…

transparency-and-documentation interpretability responsible-ai causal-inference ∑ ◫
Reasoning & Evaluation 3 min

What the bake-off taught us: classical ML is not dead, it is just under-attended

We pitted twelve sklearn algorithms head-to-head on a tabular dataset. The winner was not the most expensive one. It was not the most modern one. It was the one whose assumptions matched the data.

machine-learning tabular benchmarks evaluation
Reasoning & Evaluation 24 min

When Did the World Change? Changepoint Detection From Page's CUSUM to Bayesian Online Inference

CUSUM has been provably optimal since 1986, yet on the first human-annotated changepoint benchmark a detector that never reports a change beat most of the field under default settings. Seventy years of changepoint theory, from Pa…

anomaly-and-changepoint changepoint-detection anomaly-detection statistics ∑ ◫
Reasoning & Evaluation 20 min

When the Judge Is Also a Player: LLM-as-Judge, Contamination, and Why Leaderboards Drift

A strong model grading other models looks like a free lunch for evaluation. It is not. Position, verbosity, and self-preference biases plus quietly leaked test sets mean a leaderboard number can move several points without any mo…

evaluation llm-as-judge benchmarks contamination ∑ ◫
Reasoning & Evaluation 24 min

Your Improvement Is Inside the Noise: Seeds, Nondeterminism and the Reproducibility Problem in ML

Change one bit in one weight of a ResNet and, three epochs later, test accuracy differs by more than ten points. Training is a chaotic process, so a seed is not a control variable but a draw from a distribution. Most published an…

experiment-tracking-and-reproducibility statistics determinism mlops ∑ ◫
Reasoning & Evaluation 24 min

Your Test Set Is Wrong: Label Errors, Annotator Disagreement and the Ceiling on Measured Accuracy

Human reviewers confirmed 2,916 label errors in the ImageNet validation set, and an expert audit suggests the true figure is closer to one image in five. Every benchmark score is computed against an answer key written by people w…

human-data-and-annotation evaluation benchmarks measurement ∑ ◫
← Newer Page 2 of 3 · 26 posts Older →
Browse by topic

Tags

evaluation31 inference22 agents19 statistics17 benchmarks15 llm14 transformers13 mlops9 rag9 scaling9 alignment8 architecture8 attention8 causal-inference7 context-engineering7 experimentation7 infrastructure7 kv-cache7 latency7 llm-systems7 long-context7 mcp7 measurement7 orchestration7 reasoning7 reliability7 retrieval7 embeddings6 production6 reinforcement-learning6 security6 tool-use6 uncertainty6 anthropic5 counterfactual5 diffusion-models5
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N