Alignment & Post-Training
SFT, reward modelling, DPO/IPO/KTO/ORPO, model merging, and evaluating an aligned model.
21concepts
144flashcards
161minutes of reading
No beginner concepts in this track. Show all.
SFT, reward modelling, DPO/IPO/KTO/ORPO, model merging, and evaluating an aligned model.
No beginner concepts in this track. Show all.