Model Provenance & Watermarking
Output watermarking, content credentials, fingerprinting weights and detecting extraction.
5concepts
58flashcards
36minutes of reading
- 01 Detecting Synthetic Media Without Watermarks Why passive detectors work in the lab and fail in deployment, the base rate problem that makes accusation dangerous, and what the evidence supports doing instead.
- 02 Model Fingerprinting and Weight Attribution How to prove a deployed model was derived from yours, the difference between backdoor-style and intrinsic fingerprints, and why fine-tuning is the adversary that matters.
- 03 Text Watermarking and the Detectability Tradeoff How a statistical signal is embedded in generated text by biasing the sampler, the detection test that makes it verifiable, and the reasons the scheme survives paraphrase poorly.