Rater Sourcing, Quality and Welfare
The tradeoffs between crowd, vendor, expert and internal annotation, the mechanisms that maintain quality, and the working conditions that shape both the data and the ethics.
The people who label the data determine what the model learns, and how they are sourced, paid and managed determines both the data's quality and whether the project is defensible. These are usually treated as procurement questions and they are modelling questions.
The sourcing options
Crowd platforms give fast access to many workers at low cost, with high variance in quality and little context on the task. They suit high-volume tasks a non-expert can do with a clear guideline, and they require heavy quality machinery.
Managed vendors provide trained, supervised annotators with continuity, which means a team that learns the task and retains that learning. Costs more, produces better and more consistent data, and interposes a layer between the team and the annotators that can obscure problems.
Domain experts are necessary where the judgement requires expertise, medical, legal, or specialised technical annotation. Expensive, slow, and the only option that produces valid labels for those tasks. Their scarcity usually means expert effort should go to adjudication and guideline development rather than to bulk labelling.
Internal annotation by the team building the model gives the deepest understanding of the task and the fastest guideline iteration, and it does not scale and it carries the team's blind spots directly into the labels.
The common arrangement uses experts to write guidelines and adjudicate, a managed vendor for volume, and internal labelling of the first few hundred items to develop the guideline.
Quality mechanisms
Qualification tasks filter before production work begins, and they measure the ability to pass a test rather than sustained performance.
Gold questions seeded into the stream measure ongoing accuracy against known answers, at the cost of items that produce no data and the risk that annotators learn to recognise them.
Overlap on five to ten percent of items measures agreement continuously and is the most informative mechanism, since it detects drift and identifies which annotators diverge.
Feedback to annotators on their disagreements is what makes them better, and it is the mechanism most often skipped because it costs coordination.
Time monitoring catches rushing, and it also penalises careful work on hard items, so it should be a flag for review rather than a metric.
Welfare
Annotation work for AI systems includes exposure to violent, sexual and abusive content in safety and moderation labelling. Reporting on the conditions of this workforce has documented low pay, insufficient psychological support and unstable employment across several major projects.
This matters for two reasons that should not be conflated. It is an ethical question about the people doing the work, and it is separately a data quality question, because underpaid, rushed, unsupported annotators produce worse labels, and high turnover destroys the accumulated task understanding that makes a team good.
Concretely: pay above the platform floor, cap exposure to distressing content per session, provide genuine psychological support rather than a leaflet, employ stably enough that expertise accumulates, and know the conditions in the supply chain rather than accepting a vendor's assurance.
When it breaks
Cost per label is the wrong optimisation target. Cheap labels with high noise require more of them and cap model performance through the agreement bound. Cost per unit of usable signal is the quantity, and it frequently favours the more expensive option.
Vendor layers hide problems. A team that never sees an annotator's question does not learn that the guideline is unclear. A feedback channel from annotators back to the team is worth insisting on contractually.
Annotator populations skew the model. A workforce concentrated in particular countries, languages and income levels encodes those perspectives as the model's defaults, and the resulting behaviour is presented to users as neutral.
Turnover resets the learning. A team that has internalised a guideline is far more consistent than a new one, so employment stability is a direct input to data quality rather than only an ethical consideration.
12 flashcards for this concept
Click a card to reveal the answer.