Safety & Alignment

Prompt injection, jailbreaks, Constitutional AI, reward hacking and mechanistic interpretability.

11concepts
54flashcards
93minutes of reading