Alignment & Post-Training

SFT, reward modelling, DPO/IPO/KTO/ORPO, model merging, and evaluating an aligned model.

21concepts
144flashcards
161minutes of reading

No beginner concepts in this track. Show all.