AI for Science
AlphaFold, protein language models, materials discovery, and the pitfalls of ML-for-science.
12concepts
137flashcards
99minutes of reading
- 01 Machine Learning for Materials Discovery How graph neural networks screen millions of candidate crystals for stability and how machine-learning interatomic potentials approximate DFT cheaply enough to simulate them, plus why an in-silico "stable" material is not yet a real one.
- 02 Pitfalls in ML-for-Science The failure modes, chiefly data leakage, that make machine-learning results in scientific papers look stronger than they replicate, and the reporting standards proposed to catch them.
- 03 Protein Language Models How masked-language-model pretraining over amino-acid sequences produces structure and function signal, and how ESMFold trades some accuracy for dropping the MSA search that AlphaFold2 depends on.
- 04 Self-Driving Labs and Autonomous Experimentation The A-Lab synthesised 41 of 58 target compounds in 17 days with no human intervention, and the dispute that followed — ending in a 2026 author correction — is the clearest available lesson in what an autonomous laboratory actually automates.
- 05 AI for Formal Mathematics How proof assistants turn mathematics into a verifiable reward signal, what AlphaGeometry and AlphaProof achieved at the IMO, and why autoformalisation remains the bottleneck.
- 06 AlphaFold2 and the Protein-Folding Problem How a deep-learning system read co-evolution signal out of aligned protein sequences to predict 3D structure at near-experimental accuracy, and what it still cannot do.
- 07 AlphaFold3 and Biomolecular Co-Folding How AlphaFold3 dropped the protein-only structure module for a diffusion head that denoises raw atoms, letting one model co-fold proteins with ligands, nucleic acids, ions, and modified residues, and how the open reimplementations caught up.
- 08 Generative Protein and Molecule Design Structure prediction reads nature's proteins; generative design writes new ones — how RFdiffusion denoises backbones into existence, why a separate network is needed to choose the sequence, and why the only benchmark that counts is a wet-lab success rate.
- 09 Machine-Learned Interatomic Potentials How equivariant graph networks reach near-quantum accuracy at a fraction of the cost, why symmetry is designed in rather than learned, and what breaks when a potential leaves its training chemistry.
- 10 Neural Operators and PDE Surrogates Why learning a mapping between function spaces is different from fitting a network to a grid, how the Fourier neural operator achieves resolution invariance, and what a surrogate cannot promise.
- 11 Neural Weather Prediction How graph and transformer models trained on reanalysis data overtook physics-based forecasting on most verification targets, what they still depend on, and where the learned approach genuinely fails.
- 12 Single-Cell Foundation Models Pretraining a transformer on tens of millions of single-cell transcriptomes produced scGPT and Geneformer, and then a zero-shot benchmark found both losing to highly variable gene selection — a case study in what "foundation model" does and does not transfer.