Quiz
2707 questions of the kind that actually get asked — in interviews, in architecture review boards, and by the person who has to run the thing at 3 AM. Every answer states the trade-off rather than the slogan, and says when the obvious choice is the wrong one.
All areas2707
Architecture Fundamentals81
Distributed Systems101
Data Architecture90
Cloud Architecture87
Networking86
API & Integration Architecture78
Reliability & Resilience88
Observability81
Performance & Capacity Engineering90
Security Architecture95
Cost Architecture & FinOps92
Business Architecture93
Architecture Communication91
Enterprise Architecture91
Legacy Modernization92
AI-Era Architecture96
Software Architecture & Engineering84
Architecture Patterns84
Architecture Decision-Making91
The Architect's Meta-Skills92
Delivery & Release Engineering93
Platform Engineering & Developer Experience92
Testing & Quality Architecture90
Data Platform Architecture98
Streaming & Real-Time Data93
Data Governance & Semantics91
Frontend & Experience Architecture91
Edge, Mobile & IoT88
Regulatory & Data Protection Architecture90
Assurance, Audit & Model Risk98
51 questions in Data Platform Architecture.
-
File Formats & Compaction advanced
Review this configuration. A 60 TB events table is partitioned by ingest_date and written by a streaming job that commits every 60 seconds. A compaction job runs hourly over the last 24 hours with a 512 MB target and sorts each file by event_timestamp. Snapshot expiry runs monthly with 90-day retention. 85% of queries filter on customer_id over a 7-day range. What would you remove, what would you change and what would you leave alone?
3 min answer compactionsort keysnapshot expirystatistics -
Ingestion Patterns advanced Multiple choice
A connector extracts a SaaS object by paging with offset and a limit of 500 under a modified-at filter. Every nightly run loads exactly 10000 rows and finishes green. Reconciliation shows 61000 matching records at source. The API returns HTTP 200 with an empty page at offset 10000 and support confirms an undocumented deep-paging cap. Which change actually closes the gap?
3 min answer connectorspaginationsilent-failurewatermarks -
Ingestion Patterns advanced
A data-ingestion platform runs hundreds of connectors, each with different APIs, rate limits, failure modes and schemas. How should isolation, retries, scheduling, checkpointing, backfills and schema evolution be designed?
2 min answer airbytefivetranconnectorsisolation -
Ingestion Patterns advanced
A platform ingests high-volume telemetry from many sources. Which ingestion pattern properties matter most?
2 min answer ingestionbackpressureidempotencyschema -
Open Table Formats advanced
A team is choosing between Iceberg, Delta and Hudi for a new lakehouse. How would you approach the decision?
1 min answer data-platformtable-formatlock-in -
Open Table Formats advanced
An Iceberg lakehouse on object storage is healthy. Then the storage service starts serving reads at about four times its usual latency with no errors at all. Nothing is down. Walk through what happens over the next ten minutes and what stops it.
3 min answer icebergmetadataslow dependencyretry storm -
Open Table Formats advanced
An organisation is consolidating a warehouse and a lake onto open table formats. What guarantees do ACID-on-object-storage tables provide, where do small files and compaction bite, and how are streaming and batch writers coordinated?
3 min answer databricksdelta-lakeicebergcompaction -
Open Table Formats advanced Multiple choice
Which requirement most often forces an open table format over plain files in object storage?
2 min answer table-formatsdeletestransactionsschema-evolution -
Reverse ETL advanced
A warehouse table now feeds the CRM through reverse ETL. It breaks, and the analytics team and the CRM team each say it is the other's problem. How do you resolve it?
1 min answer data-platformreverse-etlownershipoperations -
Reverse ETL advanced
At 01:50 an upstream system renames a column and the nightly customer model's join silently yields 4,000 rows instead of 1.2 million. Every test on that model is a not-null check on columns that survived, so the model passes. At 02:30 reverse ETL syncs it to the CRM. By 09:15 sales reports that lifecycle stage and account owner are blank on more than a million accounts and overnight campaigns fired against the wrong segment. Nothing errored. What failed, and which design decision made it possible?
3 min answer reverse etlfull syncblast radiusvolume tests -
Slowly Changing Dimensions advanced
A customer dimension has always been Type 1 - attributes overwritten in place. Compliance now requires every order to report the segment and region the customer had at the time of the order, three years back. 140 reports read the dimension and 2.1 billion fact rows join it on the natural customer key. Sequence the conversion under live reporting.
3 min answer scdtype-2migrationsurrogate-keys -
Storage Layout & Partitioning advanced
A platform's analytical queries scan far more data than they need. Which storage layout decisions fix this?
2 min answer partitioningclusteringpruninglayout