Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “reliability”

Tagged “reliability”

7 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Training & Alignment 24 min

At Sixteen Thousand GPUs, Something Is Always Broken: Failures, Stragglers and Silent Data Corruption in Training Clusters

Over 54 days of Llama 3 405B pre-training, the job was interrupted 466 times, roughly once every three hours. At that failure rate the checkpoint interval barely matters; what decides how much of a sixteen-thousand-GPU cluster do…

gpu-fleet-and-capacity distributed-training reliability gpu ∑ ◫
Inference & Serving 24 min

Designing for a Collaborator That Is Sometimes Wrong: The Evidence Behind Human-AI Interaction Design

In a 2025 randomized trial, experienced developers using AI tools took 19% longer to finish their tasks while believing they had been 20% faster. Twenty-six years of human-AI interaction research explain the gap: an assistant's v…

interaction-design-for-ai product verification calibration ∑ ◫
Reasoning & Evaluation 24 min

Hidden Technical Debt, a Decade On: What Continuous Delivery for ML Actually Fixed

In 2015 a Google paper catalogued the ways machine learning systems rot, and a decade of MLOps tooling set out to pay that debt down. It paid down the debt that lives in pipelines and artefacts, and left the debt that lives in ju…

ci-cd-for-ml mlops evaluation-mlops deployment ∑ ◫
Platforms & Practice 24 min

The Leak in Every Training Set: Feature Stores, Point-in-Time Joins, and the Train-Serve Contract

A fraud model can score perfect recall offline and block nothing in production, because its training join looked a few hours into the future. Feature stores exist to enforce one contract: a training row may only see what the serv…

feature-stores mlops feature-engineering data ∑ ◫
Reasoning & Evaluation 24 min

When Did the World Change? Changepoint Detection From Page's CUSUM to Bayesian Online Inference

CUSUM has been provably optimal since 1986, yet on the first human-annotated changepoint benchmark a detector that never reports a change beat most of the field under default settings. Seventy years of changepoint theory, from Pa…

anomaly-and-changepoint changepoint-detection anomaly-detection statistics ∑ ◫
Reasoning & Evaluation 20 min

When the Judge Is Also a Player: LLM-as-Judge, Contamination, and Why Leaderboards Drift

A strong model grading other models looks like a free lunch for evaluation. It is not. Position, verbosity, and self-preference biases plus quietly leaked test sets mean a leaderboard number can move several points without any mo…

evaluation llm-as-judge benchmarks contamination ∑ ◫
Reasoning & Evaluation 14 min

Your eval pipeline is the moat, not your model choice

The model layer is commoditising and the answer flips every six months. The only durable advantage is the ability to A/B a model swap end-to-end in 48 hours and know whether it improved things for your users.

evaluation benchmarks llm-as-judge reliability
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N