Evaluation & MLOps

Benchmarks, LLM-as-judge, red-teaming, model registries, drift detection and observability.

11concepts
57flashcards
91minutes of reading

No beginner concepts in this track. Show all.