Skip to content
∑ Praveen T N Learning Library
Concepts Flashcards Writing Editorial AI Feed Graph
Portfolio ↗
Overview Concepts Flashcards Writing Editorial AI Feed Graph Search Back to portfolio ↗
Library/ Writing/Tagged “jailbreaks”

Tagged “jailbreaks”

2 posts.

Clear
All Model Architecture19 Training & Alignment22 Inference & Serving18 Agents & Orchestration11 Reasoning & Evaluation26 Safety, Security & Governance7 Platforms & Practice20
Safety, Security & Governance 24 min

A Decade of Adversarial Examples: Why Robustness Never Came Free

In 2014 a perturbation the size of one 8-bit colour step turned a 57.7 percent panda into a 99.3 percent gibbon. Twelve years, 300 million synthetic training images and more than 10^21 training FLOPs later, the best CIFAR-10 mode…

security safety evaluation benchmarks ∑ ◫
Safety, Security & Governance 21 min

The Moderation Tax: How Guardrail Classifiers Trade Latency for Coverage

A guardrail is a classifier sandwich wrapped around your model, and every layer you add buys coverage with latency and false refusals. Here is how the layer actually works, what it costs, and where it breaks.

llm-safety guardrails content-moderation jailbreaks ∑ ◫
The library

1015 concepts, 11,105 flashcards and 123 long-form pieces on AI, NLP, deep learning, LLMs and agentic systems. Free, no sign-up, no paywall.

Sections Concepts Flashcards Writing AI Feed Knowledge Graph
Elsewhere Portfolio Architecture Practice RSS LinkedIn Buy me a coffee

Written and maintained by Praveen T N.

© 2026 Praveen T N