A complete, university-grade curriculum for deep learning — from a single neuron through transformers, pretraining, alignment, generative models and deployment. Every method is derived from first principles, implemented from scratch in NumPy, then reproduced with PyTorch or Hugging Face. All notebooks run in Google Colab with no local setup.
Four repositories, one curriculum. The first three are how a model learns; the fourth is what it is built from, and cuts across all three.
- Supervised learning — learning from labelled examples
- Unsupervised learning — finding structure with no labels
- Reinforcement learning — learning from reward
- Deep learning & transformers — the architecture the other three can each be built on (you are here)
Status: 20 of 20 lessons complete (36 notebooks)
- From First Principles: every algorithm derived from foundations, not quoted
- Dual Structure: theory (a) + practical (b) notebooks for each lesson
- Story-Driven: a real-world motivation before the mathematics
- Complete Implementations: from-scratch NumPy, then production libraries
- Google Colab Compatible: runs in the browser, no local setup
- Machine-Verified: every notebook executes top-to-bottom on CPU in under 10 minutes,
checked by
scripts/verify_notebook.shbefore its task closed
See CURRICULUM_PLAN.md for the full table, including data used per lesson.
- Lesson 0: Introduction to Deep Learning
0a_intro_deep_learning_theory.ipynb— what stacking non-linear layers buys, representation learning, and the computational graph0b_intro_deep_learning_practical.ipynb— tensors, autograd,nn.Module, and the training-loop skeleton, on MNIST
- Lesson 1: From Linear Models to Neurons
1a_linear_to_neurons_theory.ipynb— the perceptron and logistic regression as a one-layer net; loss surfaces; batch vs mini-batch vs SGD1b_linear_to_neurons_practical.ipynb— the same model in PyTorch, withDataLoaderand an overfit-one-batch sanity check
- Lesson 2: Multilayer Perceptrons and Backpropagation
2a_mlp_backprop_theory.ipynb— the chain rule to backprop for an L-layer MLP, with manual gradients and gradient checking2b_mlp_backprop_practical.ipynb— a PyTorch MLP, autograd checked against the manual gradients
- Lesson 3: Training Dynamics
3a_training_dynamics_theory.ipynb— Xavier/He initialisation, vanishing/exploding gradients, BatchNorm/LayerNorm, and momentum/RMSProp/Adam, all derived3b_training_dynamics_practical.ipynb— an initialisation × normalisation × optimiser ablation with learning-rate schedules, on CIFAR-10
- Lesson 4: Regularisation and Generalisation
4a_regularisation_theory.ipynb— bias-variance in deep nets; weight decay as L2; dropout as an implicit ensemble; early stopping; augmentation4b_regularisation_practical.ipynb— augmentation pipelines and dropout/weight-decay ablations, with learning curves
- Lesson 5: Convolutional Networks
5a_convolutional_networks_theory.ipynb— convolution as a linear operator, parameter sharing, receptive fields and pooling, with backprop through conv from scratch5b_convolutional_networks_practical.ipynb— a CNN on CIFAR-10 compared against a dense baseline, with feature-map visualisation
- Lesson 6: Modern Architectures and Transfer Learning
6a_modern_architectures_theory.ipynb— why residual connections ease optimisation, norm placement, and the major architecture families6b_modern_architectures_practical.ipynb— fine-tuning a pretrained torchvision ResNet, feature extraction vs full fine-tune
- Lesson 7: Sequence Models
7a_sequence_models_theory.ipynb— the RNN forward pass and backpropagation through time, vanishing gradients, and LSTM/GRU gates, with a character RNN from scratch7b_sequence_models_practical.ipynb— an LSTM language model with teacher forcing and temperature sampling
- Lesson 8: Embeddings and Tokenisation
8a_embeddings_tokenisation_theory.ipynb— the distributional hypothesis and skip-gram with negative sampling, derived8b_embeddings_tokenisation_practical.ipynb— byte-pair encoding from scratch vs thetokenizerslibrary, and a pretrained embedding space visualised
- Lesson 9: Attention
9a_attention_theory.ipynb— the alignment problem; additive and scaled dot-product attention derived as a differentiable soft lookup9b_attention_practical.ipynb— seq2seq with attention on a synthetic task, with attention heatmaps and masking
- Lesson 10: The Transformer
10a_the_transformer_theory.ipynb— multi-head self-attention, positional encodings, residual+layer-norm blocks, and causal masking10b_the_transformer_practical.ipynb— a decoder-only Transformer trained and compared against the Lesson 7 LSTM, with head/layer ablation
- Lesson 11: Language Model Pretraining
11a_lm_pretraining_theory.ipynb— the autoregressive objective, cross-entropy and perplexity, scaling intuition, and a mini-GPT's forward pass11b_lm_pretraining_practical.ipynb— a GPT-style model trained briefly and compared against Hugging Face GPT-2 inference
- Lesson 12: Fine-tuning and Adaptation
12a_finetuning_adaptation_theory.ipynb— full fine-tuning vs feature extraction; LoRA's low-rank update derived; catastrophic forgetting12b_finetuning_adaptation_practical.ipynb— Hugging Face fine-tuning of a small encoder, and LoRA via PEFT
- Lesson 13: Alignment: RLHF and Preference Optimisation
13a_alignment_rlhf_theory.ipynb— preference modelling and the Bradley-Terry reward model, derived; the RLHF loop; DPO derived13b_alignment_rlhf_practical.ipynb— a reward model trained on a tiny preference set, DPO on a small language model, and a reproduced reward-hacking failure
- Lesson 14: Generative Models
14a_generative_models_theory.ipynb— autoencoders; the VAE ELBO and reparameterisation trick derived; the diffusion forward/reverse process derived14b_generative_models_practical.ipynb— a VAE and a minimal diffusion model, trained and sampled from, on MNIST
- Lesson 15: Efficient and Scalable Deep Learning
15a_efficient_scalable_theory.ipynb— compute/memory accounting; mixed precision; int8 quantisation derived; distillation15b_efficient_scalable_practical.ipynb— automatic mixed precision, dynamic quantisation and distillation, measured against a full-precision baseline
- Lesson X1: Debugging Deep Networks
X1_debugging_deep_networks.ipynb— loss-curve taxonomy, gradient norms, dead ReLUs, overfit-one-batch, and NaN hunting
- Lesson X2: Evaluation and Benchmarking
X2_evaluation_benchmarking.ipynb— splits and leakage, metrics beyond accuracy, calibration, and statistical comparison of runs
- Lesson X3: Deployment and Safety
X3_deployment_safety.ipynb— TorchScript/ONNX export, latency and throughput, monitoring and drift, robustness, and responsible use
- Lesson X4: Research Frontiers
X4_research_frontiers.ipynb— scaling laws, mixture of experts, state-space models, multimodal models, and interpretability
scripts/setup_env.sh # build .venv from requirements.txt
scripts/verify_notebook.sh notebooks/0a_intro_deep_learning_theory.ipynbThe curriculum is written by autonomous Claude Code sessions under the
consciousness plugin. CONSCIOUSNESS/
carries the directive, stories, features, tasks and every review verdict — the full
provenance of what was built, by which seat, against which acceptance criteria.
See LICENSE.md.