LLMs from First Principles
For engineers and technical managers: from basic transformers to a working nano-scale GPT.
What you will be able to do
- Explain the basic structure of the transformer, compared with a feed-forward network.
- Build character-level language models, from RNNs and GRUs to the causal (autoregressive) transformer.
- Tokenise text with byte pair encoding.
- Pretrain a small GPT-style language model.
- Fine-tune it to follow instructions.
- Apply reinforcement learning from rewards, and explain how reward hacking arises.
- Build a minimal software-engineering (SWE) agent: a loop in which a model calls tools such as running commands.
- Discuss the safety risks of letting an agent run commands.
Every part is built small enough to run on a laptop.
Prerequisites
Python, and some introduction to deep learning, for example Introduction to Deep Learning or an online course.
Outline
Duration: 2 days, on-site or remote.
- Set models: the basic structure of the transformer, compared with a feed-forward network
- Character-level language models: from RNNs and GRUs to the causal (autoregressive) transformer
- Tokenisation: byte pair encoding
- A full GPT: pretraining a small language model
- Fine-tuning to follow instructions
- Reinforcement learning from rewards, including reward hacking
- Building a SWE agent
Sample materials
- Set models: feed-forward network, Deep Sets and Set Transformer
- Recurrent neural networks
- Gated recurrent units
- Character-level transformer
- Byte pair encoding tokenizer
- Pretraining a small GPT
- Supervised fine-tuning
- Reinforcement learning from verifiable rewards
- A minimal software-engineering agent
More in yoavram/nanochat.
Related workshops
Contact
Tell me about your team and what you would like them to be able to do. A few lines are enough to start; I will reply to set up a scoping call.
Email Yoav yoav@yoavram.com