Training & Alignment
20 min
Constitutional AI and RLAIF: Scaling Oversight Without Scaling Labels
Human preference labels are the most expensive ingredient in a modern aligned model. Constitutional AI replaced most of them with a written document and a model judging itself, and the idea quietly took over the alignment stack.