Vision & Multimodal
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
12concepts
131flashcards
101minutes of reading
No beginner concepts in this track. Show all.
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
No beginner concepts in this track. Show all.