intermediate
1 min answer
A design platform holds user-generated content, customer brand assets and internal data in one estate. How should sensitivity labelling work so it is actually applied?
Show the full answer Hide the answer
Why manual classification fails
It depends on the person creating the data knowing the taxonomy, caring, and choosing correctly — at the moment they are trying to do something else. Coverage ends up partial and skewed: the careful teams label, the rest do not, and unlabelled data is treated as unclassified rather than as sensitive.
What makes labelling work
- A small taxonomy. Three or four levels people can hold in their heads. A twelve-level scheme produces inconsistent application, which is worse than a coarse one applied consistently.
- Automated inference where possible — content scanning, source-based defaults, schema-based rules — with human confirmation rather than human origination.
- Inheritance and propagation. A derived dataset inherits the highest sensitivity of its inputs, automatically, through lineage. Without propagation, an aggregate of sensitive data is unlabelled and therefore unprotected.
- Labels that do something. If the label does not change access, retention, encryption or export behaviour, nobody maintains it. Consequence is what keeps classification current.
- Default to the more restrictive level for unlabelled data, which converts absence from a silent hole into visible friction someone will fix.
The specific hazard in this domain
Customer-owned content is not the platform's to classify freely. Brand assets uploaded by a customer carry the customer's obligations, not the platform's defaults — so the label needs to reflect whose data it is, which is a dimension separate from how sensitive the content is.