Synthetic Data
Generating data with the shape and edge cases of the real thing, and where it misleads.
4 to work through
-
intermediate Multiple choice
A payments company must build a test dataset for its fraud-scoring service. Production holds 400M transactions with about 0.2% labelled fraudulent; legal will not permit production data in test environments. The team needs the dataset to exercise the model's decision boundary and the pipeline's handling of awkward records. Which approach fits the requirement as stated?
2 min answer -
intermediate
When is synthetic data sufficient and when does it mislead?
2 min answer -
advanced Multiple choice
A live-streaming platform at Twitch scale load tests its chat-history service against two billion synthetic messages whose channel IDs were drawn uniformly at random. The test reports p99 of 40 ms at 5000 requests per second. Production at 3000 requests per second shows p99 of 800 ms with one shard at 95% CPU while the rest sit near 20%. Which explanation fits the evidence?
3 min answer -
advanced
Your synthetic test data generator produces valid records and the team says testing has got worse. What is likely wrong?
1 min answer
2 terms in this topic
Synthetic Data
Artificially generated data that preserves the statistical and structural properties of real data without containing anyone's actual records.
conceptSynthetic Data Fidelity
How closely generated data reproduces the shape, distribution and awkwardness of the real thing, which decides what the data can validly be used for.
Neighbouring topics
Testing & Quality Architecture
General material on designing a testing strategy as an architectural concern.
Test Architecture Strategy
Choosing what to verify where, given the failure modes that actually occur.
Test Pyramid Shapes
Pyramid, trophy and honeycomb, and the system properties that justify each shape.
Integration Test Boundaries
What sits inside a test's boundary, what is faked, and the confidence that follows.
Contract Testing at Scale
Keeping dozens of services compatible without an environment that runs all of them.
Consumer-Driven Contracts
Consumers declaring what they rely on, and providers verifying against those declarations.
Test Data Management
Realistic data without copying production personal data into a weaker environment.
Environment Parity
The differences between staging and production that decide which bugs survive to release.
Service Virtualisation
Standing in for a dependency you cannot call, and keeping the stand-in honest.
End-to-End Test Economics
Why broad end-to-end suites get slow, flaky and abandoned, and what to keep.
Non-Functional Test Strategy
Testing availability, latency, security and recovery rather than only behaviour.
Performance Test Design
Workload models, warm-up, think time, and the distribution the average hides.
Chaos as a Test
Fault injection with a hypothesis, a blast radius and an abort condition.
Security Testing in the Pipeline
SAST, DAST, dependency and secret scanning, and what to do with the findings.
Accessibility Testing
Automated checks, their ceiling, and the manual testing that has to sit above it.
Mutation Testing
Measuring whether tests would actually notice a defect, not just cover a line.
Flaky Test Management
Quarantine, detection, and the trust a suite loses once red stops meaning broken.
Testing in Production
Synthetic transactions, dark launches and shadow traffic, done deliberately and safely.
Quality Gates
Thresholds that block a release, who may override them, and how they decay.