Ten complete modules that take an engineer who knows calculus, linear algebra and basic probability from a first loss function to training, aligning and serving a language model.
All ten modules are available in English and Simplified Chinese. Each is a self-contained unit with five study sessions and about ten hours of scheduled activities. Allow 10–15 hours per module, depending on time spent on derivations, lab reruns and review. A session mixes reading with doing.
A study plan at the top of every module lists its sessions with their timings. Tick them off as you go; your progress is kept in your browser.
Read the modules in order. Module 06 needs 02 and 04; Modules 08 and 09 need 06 and 07. Module 10 builds on 06–09, including formats, the running case and the merged post-trained model.
The ten modules and what each depends on. Foundations (01–02) lead to the architectures (03–06); the transformer (06) leads to large language models (07), which branch into pretraining (08), post-training (09) and inference (10).
Every number comes from somewhere. When a module says a method reaches an accuracy, costs a number of FLOPs or needs a number of gigabytes, it says how that number was obtained, so you can redo it.
Every method can fail, and the text says how. A technique explained without its failure mode is advertising. Every module has a section on what goes wrong, because that is the section an engineer needs at three in the morning.
Engineers and postgraduate students who are comfortable with calculus, linear algebra and basic probability but have not worked in machine learning. The aim is not a survey. It is to leave you able to read a modern model paper, understand what a training run is doing to a model, and make engineering decisions about models with judgement rather than by recipe.
Python 3.11 or later and a virtual environment are enough for every lab. A GPU is optional; Google Colab is a free alternative to a local install.
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
pip install torch numpy scipy scikit-learn matplotlib
pip install transformers datasets tokenizers tiktoken huggingface_hub pandas pyarrow # Modules 07 to 10
Modules 07 to 10 download small open models from Hugging Face, the largest about one gigabyte; each lab states its download size.
Vectors are bold lower case, \mathbf{x}; matrices are bold upper case, \mathbf{W}; scalars are italic, y. A dataset is \mathcal{D} = \{(\mathbf{x}_i, y_i)\}_{i=1}^{N}. Model parameters are collected in \theta. A loss is \mathcal{L}(\theta) and its gradient is \nabla_\theta \mathcal{L}. Expectation over data is \mathbb{E}_{(\mathbf{x},y)\sim\mathcal{D}}. Logarithms are natural unless written \log_2. Tensor shapes are written (B, T, d) for batch, sequence length and width. Code is Python with PyTorch, kept short enough to type.
Every module ends with its key terms in English and Chinese, for readers who read papers in English and discuss them in Chinese.