Human Data & Annotation
Guideline design, inter-annotator agreement, preference collection, rater sourcing and label noise.
5concepts
60flashcards
35minutes of reading
- 01 Guidelines Are the Model Specification Why the annotation guideline determines what the model learns more than the architecture does, what a usable guideline contains, and the iteration loop that produces one.
- 02 Inter-Annotator Agreement and What It Bounds Why raw agreement overstates reliability, what Cohen's and Krippendorff's coefficients correct for, and why agreement is the ceiling on any model trained from those labels.
- 03 Label Noise and Learning Through It How random and systematic label noise differ in their effect on a model, why memorisation of noisy labels happens late in training, and the techniques that find mislabelled data cheaply.
- 04 Preference Data Collection Why pairwise comparison replaced absolute rating for alignment data, the biases that contaminate it, and the design choices that determine what a reward model actually learns.
- 05 Rater Sourcing, Quality and Welfare The tradeoffs between crowd, vendor, expert and internal annotation, the mechanisms that maintain quality, and the working conditions that shape both the data and the ethics.