Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “evaluation”

Tagged “evaluation”

31 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Safety, Security & Governance 24 min

A Decade of Adversarial Examples: Why Robustness Never Came Free

In 2014 a perturbation the size of one 8-bit colour step turned a 57.7 percent panda into a 99.3 percent gibbon. Twelve years, 300 million synthetic training images and more than 10^21 training FLOPs later, the best CIFAR-10 mode…

security safety evaluation benchmarks ∑ ◫
Platforms & Practice 24 min

A Model Version Is No Longer a File: Registries for Compound AI Systems

Model registries were built to version a trained artefact you own. An LLM application's behaviour comes from a hosted snapshot that retires on someone else's calendar, plus prompts, an index and tools you change weekly. The relea…

model-registry-and-versioning mlops llm-systems evaluation ∑ ◫
Safety, Security & Governance 24 min

Algorithmic Audits: What an Outside Examination of an AI System Can Actually Establish

Gender Shades measured a 34.4-point error gap on 1,270 faces and moved three vendors within seven months. New York City's mandatory bias audits produced 18 posted reports from 391 employers. The difference was not auditor skill b…

ai-assurance-and-audit evaluation governance compliance ∑ ◫
Reasoning & Evaluation 23 min

Clustering Has No Ground Truth: Impossibility, Validation, and What a Cluster Can Promise

In 2002 Jon Kleinberg proved that no clustering function can satisfy three properties almost everyone would ask for. Every algorithm is therefore a definition of what a cluster is, and every validation index is another definition…

clustering unsupervised-learning evaluation metrics ∑ ◫
Model Architecture 23 min

Context Rot: Why Bigger Context Windows Don't Mean Better Retrieval

A million-token window promises perfect recall of everything you feed it. Controlled tests on 18 frontier models show recall degrading steadily, unevenly, and well before the window fills, a pattern researchers now call context rot.

llm long-context rag context-engineering ∑ ◫
Reasoning & Evaluation 25 min

Error Bars for Evals: Why Most Benchmark Differences Are Noise

A 250-question benchmark carries a standard error of about three percentage points. Most of the model comparisons published on top of such benchmarks cannot distinguish the models they are comparing. Evaluations are experiments, …

evaluation benchmarks statistics mlops ∑ ◫
Reasoning & Evaluation 23 min

From Features to Circuits: What Attribution Graphs Explain, and the Fraction They Do Not

Swap the Texas features for British Columbia and Claude answers Victoria instead of Austin. That single intervention is the strongest evidence yet that a language model performs genuine multi-step reasoning inside one forward pas…

interpretability mech-interp safety alignment ∑ ◫
Reasoning & Evaluation 23 min

Guarantees Without Calibration: Conformal Prediction and the Limits of LLM Confidence

A language model's stated confidence is a number, not a probability. Conformal prediction offers the opposite trade: it promises nothing about any single answer and something exact about the long run, from any scorer, with one as…

uncertainty evaluation calibration safety ∑ ◫
Reasoning & Evaluation 24 min

Hidden Technical Debt, a Decade On: What Continuous Delivery for ML Actually Fixed

In 2015 a Google paper catalogued the ways machine learning systems rot, and a decade of MLOps tooling set out to pay that debt down. It paid down the debt that lives in pipelines and artefacts, and left the debt that lives in ju…

ci-cd-for-ml mlops evaluation-mlops deployment ∑ ◫
Training & Alignment 22 min

Language Modelling Is Compression: The Seventy-Year-Old Idea Underneath Every LLM

In 1951 Claude Shannon estimated the entropy of English by having people guess the next letter. In 2023 a 70-billion-parameter language model compressed a gigabyte of Wikipedia to 8.3% of its size, beating every compressor ever p…

information-theory compression entropy language-models ∑ ◫
Training & Alignment 27 min

Model Merging: Why Averaging Weights Works, and Where the Free Lunch Ends

Three 7B models that each scored under 30% on Japanese maths were averaged into one that scored 52%. No gradient was computed. Weight-space arithmetic is the cheapest capability gain in the field and the easiest one to fool yours…

model-merging fine-tuning task-vectors peft ∑ ◫
Training & Alignment 22 min

Prompts as Programs: What Changes When You Optimise the Prompt Instead of the Weights

A prompt optimiser that never touches a weight has been reported to beat GRPO by around six points using up to 35 times fewer rollouts. That result only makes sense once you stop treating the prompt as writing and start treating …

prompt-engineering llm optimization evaluation ∑ ◫
Page 1 of 3 · 31 posts Older →
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N