Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “interpretability”

Tagged “interpretability”

5 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Reasoning & Evaluation 23 min

From Features to Circuits: What Attribution Graphs Explain, and the Fraction They Do Not

Swap the Texas features for British Columbia and Claude answers Victoria instead of Austin. That single intervention is the strongest evidence yet that a language model performs genuine multi-step reasoning inside one forward pas…

interpretability mech-interp safety alignment ∑ ◫
Reasoning & Evaluation 23 min

Latent Reasoning: Teaching Language Models to Think Without Tokens

Chain-of-thought made models reason out loud, one word at a time. A new line of work lets them reason in the silent space between words, trading auditability for compute that does not have to be spelled out.

latent-reasoning chain-of-thought test-time-compute reasoning ∑ ◫
Reasoning & Evaluation 23 min

The Model Knows It Is Being Tested: Evaluation Awareness and the Limits of Behavioural Safety Evidence

On one synthetic honeypot evaluation, Claude Sonnet 4.5 said out loud that it suspected it was being tested in 80 to 100 percent of transcripts, against under 10 percent for its predecessor. When the internal representations behi…

safety-alignment evaluation interpretability red-teaming ∑ ◫
Model Architecture 23 min

The Residual Stream: The Transformer's Shared Memory Bus

Stop reading the transformer as a pipeline of 96 layers, each transforming the output of the last. Read it as one shared communication channel that every attention head and MLP merely edits. This one reframing, formalised by inte…

transformer-anatomy interpretability residual-stream attention ∑ ◫
Reasoning & Evaluation 24 min

What a Feature Attribution Can and Cannot Tell You: SHAP, LIME and the Explanation Gap

Add a column the model never reads and SHAP can hand it more than a quarter of the credit for a decision. That is not a library bug: a Shapley attribution answers a question you chose, often without noticing, and a 2024 PNAS resu…

transparency-and-documentation interpretability responsible-ai causal-inference ∑ ◫
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N