Flashcards
2,671 cards in 36 decks, written from the same source as the concepts. Space to flip, arrows to move, K and R to sort what you know from what you do not. Progress lives in your browser.
08
Multimodal & Applications
Beyond text — vision, speech, robotics and scientific discovery.
5decks
395cards
Vision & Multimodal
ViT, CLIP, diffusion, SAM, and the vision-language models that read images as tokens.
Speech Recognition
Spectrograms, CTC, RNN-T, Conformer, Whisper, streaming, diarisation and self-supervised audio.
Speech Synthesis
Acoustic models and vocoders, Tacotron, FastSpeech, HiFi-GAN, neural codecs and voice cloning.
Robotics & Embodied AI
Vision-language-action models, action tokenisation, diffusion policies and sim-to-real.
AI for Science
AlphaFold, protein language models, materials discovery, and the pitfalls of ML-for-science.