From Machine Learning to Large Language Models

Ten complete modules that take an engineer who knows calculus, linear algebra and basic probability from a first loss function to training, aligning and serving a language model.

AI · Deep learning · LLMs10 modules · 100–150 hoursEN / 中文

How the series works

All ten modules are available in English and Simplified Chinese. Each is a self-contained unit with five study sessions and about ten hours of scheduled activities. Allow 10–15 hours per module, depending on time spent on derivations, lab reruns and review. A session mixes reading with doing.

A study plan at the top of every module lists its sessions with their timings. Tick them off as you go; your progress is kept in your browser.

The modules

Module 01
Machine learning foundations
What learning from data is, written out as mathematics you can compute: a model, a loss derived from a noise model, an optimiser whose behaviour an eigenvalue predicts, and the habit that separates a model that works from one that only looks as if it does, which is evaluating it honestly.
10–15 h · 5 labs · 15 exercises
Module 02
Neural networks and backpropagation
How a stack of linear maps and nonlinearities learns its own features, how reverse-mode differentiation delivers the whole gradient for at most two more forward passes, and why training is mostly a matter of keeping numbers in range: initialisation, optimisers, normalisation, regularisation and numerical stability, each with the check that shows when it has failed.
10–15 h · 5 labs · 15 exercises
Module 03
Convolutional networks
The convolutional network built from the operation up: a convolution computed by hand and in code, the cost of every layer counted, the residual connection that made depth trainable, and the same machinery used to detect objects, to segment images and volumes into measured surfaces, and to check what a classifier is actually looking at.
10–15 h · 6 labs · 15 exercises
Module 04
Recurrent networks and sequences
Recurrent networks built from the equations up: the state, backpropagation through time and why its gradients vanish, the LSTM and GRU that fixed it, honest forecasting and monitoring of engineering sensor streams, the encoder–decoder whose bottleneck produced attention, and why the transformer replaced recurrence before linear recurrences and state-space models brought it back.
10–15 h · 5 labs · 15 exercises
Module 05
Other networks worth knowing
Autoencoders and VAEs, GANs, diffusion models, graph networks, physics-informed networks and neural operators, contrastive learning and mixture of experts: what each one optimises, its equations derived, a worked example in numbers, a lab that runs on a laptop, and how each one fails.
10–15 h · 5 labs · 15 exercises
Module 06
The transformer
This module takes attention apart until you can compute it by hand, differentiate it, tile it the way FlashAttention does and rotate it with RoPE; then it assembles the modern decoder-only block, counts its parameters and FLOPs against published models, and trains a small GPT on a laptop CPU whose loss curve and attention heads you can read.
10–15 h · 6 labs · 15 exercises
Module 07
Large language models
What a large language model computes and how to reason about it with numbers: the next-token objective and its units, tokens, the scaling laws, prompting and decoding, the context window, the ways it fails, how to judge a claim about one, and what it costs to run. A hypothetical bilingual model of about 9.5B parameters, adapted to draft and check safety-case arguments, runs through the module and into Modules 08 to 10.
10–15 h · 6 labs · 15 exercises
Module 08
LLM pretraining
Everything between a pile of text and a base model: the compute budget, the data pipeline, the architecture and optimiser settings that hold at scale, the memory and parallelism arithmetic that decides what fits where, what breaks in a long run, and the two cheaper relatives an organisation actually runs, mid-training and continued pretraining. You work out on paper how a 9.5B-parameter base like the case study’s is made and what continuing its pretraining would cost, then build a deduplication pass and pretrain, destabilise and adapt small GPTs on your own laptop.
10–15 h · 5 labs · 15 exercises
Module 09
LLM post-training
Turns a base model into an assistant: supervised fine-tuning and its mechanics, low-rank adapters, reward models and RLHF, DPO derived in full, and reinforcement learning with verifiable rewards, together with the evaluation that decides whether any of it worked.
10–15 h · 6 labs · 15 exercises
Module 10
Inference and serving
Derive latency and capacity from first principles: prefill, decode, the roofline model, KV cache, continuous batching and paging, quantisation and speculative decoding. Test the mechanisms in runnable CPU labs, then account for the case study’s service objectives, cost and reliability.
10–15 h · 6 labs · 15 exercises

Learning path

Read the modules in order. Module 06 needs 02 and 04; Modules 08 and 09 need 06 and 07. Module 10 builds on 06–09, including formats, the running case and the merged post-trained model.

FOUNDATIONS ARCHITECTURES LARGE LANGUAGE MODELS Module 01 ML foundations Module 02 Neural networks Module 03 Convolutional networks Module 05 Other networks Module 04 Recurrent networks Module 06 The transformer Module 07 Large language models Module 08 Pretraining Module 09 Post-training Module 10 Inference and serving Arrows point from a module to the modules that build on it. Dashed: readable any time after Module 07.
Figure 0.1

The ten modules and what each depends on. Foundations (01–02) lead to the architectures (03–06); the transformer (06) leads to large language models (07), which branch into pretraining (08), post-training (09) and inference (10).

Suggested schedules

Two rules the series keeps

Every number comes from somewhere. When a module says a method reaches an accuracy, costs a number of FLOPs or needs a number of gigabytes, it says how that number was obtained, so you can redo it.

Every method can fail, and the text says how. A technique explained without its failure mode is advertising. Every module has a section on what goes wrong, because that is the section an engineer needs at three in the morning.

Who it is for

Engineers and postgraduate students who are comfortable with calculus, linear algebra and basic probability but have not worked in machine learning. The aim is not a survey. It is to leave you able to read a modern model paper, understand what a training run is doing to a model, and make engineering decisions about models with judgement rather than by recipe.

Setting up

Python 3.11 or later and a virtual environment are enough for every lab. A GPU is optional; Google Colab is a free alternative to a local install.

python -m venv .venv
source .venv/bin/activate          # on Windows: .venv\Scripts\activate
pip install torch numpy scipy scikit-learn matplotlib
pip install transformers datasets tokenizers tiktoken huggingface_hub pandas pyarrow   # Modules 07 to 10

Modules 07 to 10 download small open models from Hugging Face, the largest about one gigabyte; each lab states its download size.

Notation

Vectors are bold lower case, \mathbf{x}; matrices are bold upper case, \mathbf{W}; scalars are italic, y. A dataset is \mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^{N}. Model parameters are collected in \theta. A loss is \mathcal{L}(\theta) and its gradient is \nabla_\theta \mathcal{L}. Expectation over data is \mathbb{E}_{(\mathbf{x},y)\sim\mathcal{D}}. Logarithms are natural unless written \log_2. Tensor shapes are written (B, T, d) for batch, sequence length and width. Code is Python with PyTorch, kept short enough to type.

Every module ends with its key terms in English and Chinese, for readers who read papers in English and discuss them in Chinese.

Further reading for the series