Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “alignment”

Tagged “alignment”

8 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Training & Alignment 24 min

250 Documents: Why Data Poisoning Gets Easier as Models Get Bigger

The industry's defence against data poisoning was arithmetic: an attacker needs a percentage of the corpus, and a percentage of 260 billion tokens is unobtainable. In October 2025 the largest poisoning study ever run showed the r…

safety security data-poisoning backdoors ∑ ◫
Training & Alignment 20 min

Constitutional AI and RLAIF: Scaling Oversight Without Scaling Labels

Human preference labels are the most expensive ingredient in a modern aligned model. Constitutional AI replaced most of them with a written document and a model judging itself, and the idea quietly took over the alignment stack.

alignment rlhf rlaif constitutional-ai ∑ ◫
Reasoning & Evaluation 23 min

From Features to Circuits: What Attribution Graphs Explain, and the Fraction They Do Not

Swap the Texas features for British Columbia and Claude answers Victoria instead of Austin. That single intervention is the strongest evidence yet that a language model performs genuine multi-step reasoning inside one forward pas…

interpretability mech-interp safety alignment ∑ ◫
Training & Alignment 27 min

Model Merging: Why Averaging Weights Works, and Where the Free Lunch Ends

Three 7B models that each scored under 30% on Japanese maths were averaged into one that scored 52%. No gradient was computed. Weight-space arithmetic is the cheapest capability gain in the field and the easiest one to fool yours…

model-merging fine-tuning task-vectors peft ∑ ◫
Training & Alignment 21 min

RL from Verifiable Rewards: Training Models on Answers That Can Be Checked

Replace the reward model with a function that simply checks the answer, and a frontier reasoning model falls out of pure reinforcement learning. The catch is what 'checkable' quietly assumes, and what the model learns to exploit.

rlvr reinforcement-learning reasoning grpo ∑ ◫
Training & Alignment 10 min

RLHF vs DPO in production: what we learned shipping both

DPO is the right default for almost every preference-tuning project in 2026, but the cases where PPO still wins are sharper and more common than the simplicity pitch admits.

rlhf dpo alignment post-training
Training & Alignment 26 min

The Alignment Problem in Speech Recognition: CTC, RNN-T, and the Cost of Not Waiting

A ten-second utterance is a thousand acoustic frames. Its transcript is forty characters. Nobody wrote down which frames make which character, and the number of ways to line them up is astronomical. Every major speech architectur…

speech asr ctc rnn-t ∑ ◫
Training & Alignment 27 min

What the Model Remembers: Extraction, Memorisation, and the Price of a Privacy Guarantee

Two hundred dollars of API calls pulled more than ten thousand verbatim training examples out of ChatGPT. Memorisation is not a defect that better engineering removes; it scales log-linearly with everything the field is scaling. …

privacy safety security memorisation ∑ ◫
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N