Vision & Multimodal

ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.

6concepts
32flashcards
51minutes of reading

No beginner concepts in this track. Show all.