1. Synthetic Data advanced Multiple choice

    A live-streaming platform at Twitch scale load tests its chat-history service against two billion synthetic messages whose channel IDs were drawn uniformly at random. The test reports p99 of 40 ms at 5000 requests per second. Production at 3000 requests per second shows p99 of 800 ms with one shard at 95% CPU while the rest sit near 20%. Which explanation fits the evidence?

    3 min answer twitchsynthetic-datakey-skewhot-partition
  2. Synthetic Data intermediate Multiple choice

    A payments company must build a test dataset for its fraud-scoring service. Production holds 400M transactions with about 0.2% labelled fraudulent; legal will not permit production data in test environments. The team needs the dataset to exercise the model's decision boundary and the pipeline's handling of awkward records. Which approach fits the requirement as stated?

    2 min answer synthetic-datafraudtest-dataprivacy
  3. Synthetic Data advanced

    Review this: a bank generates synthetic test data by taking a production dump and replacing names, emails, account numbers and dates of birth with random values. The data volume matches production, the schema is identical, and the compliance team has signed it off as anonymised. What would you change?

    4 min answer synthetic dataanonymisationre-identificationreferential integrity
  4. Synthetic Data intermediate

    When is synthetic data sufficient and when does it mislead?

    2 min answer synthetic-datadistributionsskewperformance-testing
  5. Synthetic Data advanced

    Your synthetic test data generator produces valid records and the team says testing has got worse. What is likely wrong?

    1 min answer test-datagenerationquality