Vision & Multimodal
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
6concepts
32flashcards
51minutes of reading
No beginner concepts in this track. Show all.
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
No beginner concepts in this track. Show all.