1. Exception & Waiver Management advanced

    Your policy gate reads each waiver's expiry at evaluation time and blocks when it has passed. Sixty waivers granted against the same standard during one six-week release freeze all share an expiry date next Tuesday. Walk through what happens from Tuesday morning.

    3 min answer gitlabwaiversexpirypolicy-gates
  2. Human-in-the-Loop Design advanced

    A fraud system routes flagged transactions to human reviewers, who approve or reject within a target of 4 minutes each. A campaign triples flagged volume overnight. Reviewer headcount is unchanged. What happens to the quality of the human oversight, and what should the design do about it?

    2 min answer human-in-the-loopautomation-biascapacityqueueing
  3. Human-in-the-Loop Design advanced

    A high-consequence decision requires human oversight of an automated recommendation. What makes that oversight meaningful rather than nominal?

    2 min answer human-oversightautomation-biasworkflowaccountability
  4. Human-in-the-Loop Design advanced Multiple choice

    An anti-fraud system routes flagged transactions to human reviewers. Audit shows reviewers approved 98% of cases with a median handling time of 3 seconds. The control is documented as human oversight. What is the most accurate finding?

    3 min answer human oversightautomation biasthroughputcontrols
  5. Human-in-the-Loop Design advanced

    An approval step has a 99.8% approval rate. What does that tell you and what would you change?

    2 min answer human-in-the-looprubber-stampingescalationautomation-bias
  6. Model Evaluation & Red-Teaming advanced Multiple choice

    A generative assistant is about to be deployed to customer support, drafting replies that agents can edit before sending. Its retrieval corpus is the company's internal knowledge base. Which pre-deployment check most reduces the risk that actually matters here?

    2 min answer red-teamingevaluationragdeployment
  7. Model Evaluation & Red-Teaming advanced

    A team must demonstrate that a model behaves acceptably before deployment. What does evaluation need to cover, and what does red-teaming add?

    2 min answer scale-aievaluationred-teamsubgroups
  8. Model Evaluation & Red-Teaming advanced

    An interviewer says — you are launching a language-model feature to 2 million users. Design the evaluation gate that decides whether it ships, and tell me what it cannot tell you. Where do you take this?

    3 min answer llmevaluationred teamrelease gate
  9. Model Evaluation & Red-Teaming advanced

    Design pre-deployment evaluation for a generative AI system in a customer-facing context.

    1 min answer evaluationred-teamingadversarialsafety
  10. Model Risk Management advanced

    A deployed model performed well in validation and its business metric has declined over four months. Nothing has been deployed. What do you investigate?

    2 min answer aimonitoringdriftmodel-risk
  11. Model Risk Management advanced

    A recommendation model is retrained monthly. The challenger is promoted automatically when it beats the champion on an offline metric computed from the last 30 days of logged interactions. It passes every month, and live engagement has been flat for a year. What is happening?

    2 min answer champion challengerfeedback loopoffline evaluationmodel risk
  12. Model Risk Management advanced

    A vendor SaaS product embeds a model that scores customers, and its output drives an automated decision in your process. Your model governance framework covers models you build. What do you do?

    2 min answer model-riskthird-partygovernanceai