Vision & Multimodal
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
—
0 known
0 to review
Loading deck…
Space to flipSpace flip · ← → move · K known · R review · S shuffle
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
Loading deck…
Space to flipSpace flip · ← → move · K known · R review · S shuffle