Skip to content
∑ Praveen T N AI & ML
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
AI & ML/ Writing/Tagged “safety”

Tagged “safety”

5 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving20 Agents & Orchestration11 Reasoning & Evaluation27 Safety, Security & Governance7 Platforms & Practice22
Training & Alignment 24 min

250 Documents: Why Data Poisoning Gets Easier as Models Get Bigger

The industry's defence against data poisoning was arithmetic: an attacker needs a percentage of the corpus, and a percentage of 260 billion tokens is unobtainable. In October 2025 the largest poisoning study ever run showed the r…

safety security data-poisoning backdoors ∑ ◫
Safety, Security & Governance 24 min

A Decade of Adversarial Examples: Why Robustness Never Came Free

In 2014 a perturbation the size of one 8-bit colour step turned a 57.7 percent panda into a 99.3 percent gibbon. Twelve years, 300 million synthetic training images and more than 10^21 training FLOPs later, the best CIFAR-10 mode…

security safety evaluation benchmarks ∑ ◫
Reasoning & Evaluation 23 min

From Features to Circuits: What Attribution Graphs Explain, and the Fraction They Do Not

Swap the Texas features for British Columbia and Claude answers Victoria instead of Austin. That single intervention is the strongest evidence yet that a language model performs genuine multi-step reasoning inside one forward pas…

interpretability mech-interp safety alignment ∑ ◫
Reasoning & Evaluation 23 min

Guarantees Without Calibration: Conformal Prediction and the Limits of LLM Confidence

A language model's stated confidence is a number, not a probability. Conformal prediction offers the opposite trade: it promises nothing about any single answer and something exact about the long run, from any scorer, with one as…

uncertainty evaluation calibration safety ∑ ◫
Training & Alignment 27 min

What the Model Remembers: Extraction, Memorisation, and the Price of a Privacy Guarantee

Two hundred dollars of API calls pulled more than ten thousand verbatim training examples out of ChatGPT. Memorisation is not a defect that better engineering removes; it scales log-linearly with everything the field is scaling. …

privacy safety security memorisation ∑ ◫
The library

1040 concepts, 11,261 flashcards and 128 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee Contact Privacy policy Terms of use

Written and maintained by Praveen T N.

© 2026 Praveen T N