Vision & Multimodal

ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.

12concepts
131flashcards
101minutes of reading