Part 2 gave us attention with RoPE. Part 1 gave us a 4,096-token vocabulary. This post stacks them into a real GPT, trains it, and watches what happens.
A 5.8M-parameter GPT trains on 1M tokens of Sherlock Holmes in 59 seconds on a laptop RTX 3070. The same run takes 48 minutes on that laptop’s 8-core CPU. That is 48x. Both devices reach nearly the same loss because they see the same batches. The harder lesson comes later: with only 1M tokens, the model starts memorizing the book within about 1,250 steps, and the validation loss is the only number that tells you when.
Train a Tiny GPT at Home series: Part 1: Tokenizers · Part 2: Attention · Part 3: Training (you are here) · Part 4: Inference
Full example: Clone the working files at github.com/KingPin/sumguy-examples/…/part-3-training
Run everything from llm/tiny-gpt/ after the setup in the series README. The code imports Part 1’s tokenizer and Part 2’s attention, so train that tokenizer first. The series README installs the CPU build of torch, so the --device cuda runs need a CUDA build of torch instead, or the Docker command from the Part 3 README.
cd part-3-trainingpython model.py # parameter count, under 1 spython train.py --device cuda --steps 1500 --out gpu1500 # about 1 minutepython train.py --device cpu --threads 8 --steps 1500 --out cpu1500 # about 1 hourpython train.py --device cuda --steps 5000 --out gpu # the overfitting runAll numbers below come from one laptop: an NVIDIA RTX 3070 Laptop GPU (8 GB) and an Intel Core i7-11800H (8 cores, 16 threads), running PyTorch 2.14.1 in the pytorch/pytorch:2.14.1-cuda13.2-cudnn9-runtime Docker image. CPU and GPU share one machine, so the comparison is fair. Part 2’s CPU numbers came from a different machine (an AMD Ryzen 7 255), so do not compare CPU numbers across posts. That image lacks the regex package the tokenizer needs. The README shows pip install --break-system-packages regex==2026.9.29 inside the throwaway container, because PEP 668 blocks a plain pip install there.
The Model: Six Blocks and a Budget
A GPT is a short recipe. Token ids go through an embedding table, then through N identical blocks, then through a final norm and a linear layer that scores every vocabulary entry. Each block does two jobs. Attention lets tokens look at earlier tokens. An MLP lets each token think on its own. Both jobs are wrapped in a residual add.
class Block(nn.Module): def __init__(self, cfg: Config): super().__init__() d = cfg.d_model self.norm1, self.norm2 = nn.LayerNorm(d), nn.LayerNorm(d) self.attn = RoPEAttention(cfg) self.mlp = nn.Sequential( nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d), nn.Dropout(cfg.dropout) )
def forward(self, x): x = x + self.attn(self.norm1(x)) # mix information across tokens return x + self.mlp(self.norm2(x)) # process each token on its ownThe norm sits before each sublayer (pre-norm), and the block adds its output back onto its input. That residual path is why a stack of six blocks still trains: gradients have a direct road back to the embedding.
Running python model.py prints where the parameters live:
embedding (shared with output head) 1,048,576one block 788,736 attention 262,144 MLP 525,5686 blocks 4,732,416total 5,781,504Two things stand out. The MLP holds two thirds of every block (525,568 of 788,736), because it expands 256 dimensions to 1,024 and back. Attention gets the fame, the MLP gets the weights. Also, the embedding is only counted once. The output head reuses the embedding table:
self.head.weight = self.emb.weight # tied: one table reads tokens in and scores them outThat tie saves another 1,048,576 parameters, and it has a side effect we will use at the end: the table that reads tokens in is the same one that scores them out, so it gets pressure from both directions. The total is 5,781,504, well under the 30M promise from Part 1, with d_model 256, 4 heads of 64 dimensions, and a vocab of 4,096.
A Sanity Check Before Any Training
An untrained model should know nothing. Nothing means a uniform guess over 4,096 tokens, and the loss of a uniform guess is ln(4096) = 8.318. The script asserts that:
untrained loss 8.379, ln(4096) = 8.318untrained model guesses uniformly: OKIf this number is 20 or 2, the init is broken and no amount of training will rescue it. Two init details get it right. Weights start at a standard deviation of 0.02. The layers that write into the residual stream (attn.out and the MLP’s second linear) start smaller, at 0.02 divided by the square root of 2 times the layer count, so the sum of twelve additions stays about the same size as depth grows.
The Device Gotcha
RoPEAttention subclasses Part 2’s MultiHeadAttention and reuses its rope(). That function builds its angle tables with torch.arange, which defaults to the CPU. On a GPU that crashes with a device mismatch. The fix is a context manager, plus a cast back, because rope returns fp32 even when the rest of the model runs in bf16:
with torch.device(q.device): q, k = rope(q).to(v.dtype), rope(k).to(v.dtype)The Training Loop
The corpus is 1,030,773 tokens. The last 10% is held out and never trained on: 927,695 tokens train, 103,078 val. Training sees random 256-token windows. The target for every position is the next token, so y is x shifted by one:
x = torch.stack([d[s : s + cfg.context] for s in starts])y = torch.stack([d[s + 1 : s + 1 + cfg.context] for s in starts]) # next tokenA batch is 32 windows of 256 tokens, 8,192 tokens per step. One pass over the training set is about 113 steps. The loss is cross-entropy between the predicted distribution and the real next token, at all 256 positions at once.
The optimizer settings:
- AdamW with learning rate 1e-3 and betas (0.9, 0.95).
- Weight decay 0.1 on matrices only. Norms and biases are 1D, so they are excluded.
- Warmup then cosine. The learning rate ramps up over
min(200, steps // 10)steps (150 for a 1,500-step run), then decays along a cosine to 10% of its peak. - Gradient clipping at 1.0.
- Dropout 0.1.
- bf16 autocast on the GPU, fp32 on the CPU.
The schedule code is short enough to read in one go:
def lr_at(step): # linear warmup, then cosine decay down to 10% if step < warmup: return args.lr * (step + 1) / warmup t = (step - warmup) / max(1, args.steps - warmup) return args.lr * (0.1 + 0.9 * 0.5 * (1 + math.cos(math.pi * t)))To make the CPU and GPU runs comparable, model init and batch sampling both happen on the CPU generator with seed 1337, then move to the device. Same init, same batches, so the curves should match step for step.
Timing needs one trick. GPU calls are asynchronous: Python returns before the kernel finishes. The loop calls torch.cuda.synchronize() before starting and before stopping each step’s timer, and it excludes evaluation time. Without that, the GPU would look impossibly fast.
CPU vs GPU: Same Run, 48x Apart
Run A is the GPU for 1,500 steps in bf16. Run B is the CPU with 8 threads in fp32, same 1,500 steps.
| Step | GPU train | GPU val | CPU train | CPU val |
|---|---|---|---|---|
| 250 | 4.399 | 4.599 | 4.393 | 4.604 |
| 500 | 3.843 | 4.141 | 3.835 | 4.121 |
| 750 | 3.537 | 3.929 | 3.533 | 3.925 |
| 1000 | 3.323 | 3.842 | 3.312 | 3.831 |
| 1250 | 3.163 | 3.791 | 3.154 | 3.782 |
| 1500 | 3.078 | 3.773 | 3.070 | 3.762 |
The curves never drift more than 0.02 apart. Now the clock:
GPU 206,699 tok/s training time 59s peak memory allocated 1,008 MiBCPU 4,293 tok/s training time 2862sThat is 2,862 / 59, about 48x. A second identical GPU run reproduced every loss to three decimals, so the seed handling works. The CPU ended slightly lower (3.762 vs 3.773). That gap mixes two causes: bf16 rounding, and different dropout masks, because the CPU and CUDA draw random numbers differently even from the same seed. Either way it is noise-level.
Forty-eight minutes is tolerable for a one-off. It is not tolerable for the fifth hyperparameter tweak. That gap is the case for owning a GPU, and it matches the tradeoffs in CUDA vs ROCm vs CPU.
Which CPU Setting Is Fastest
If you must train on a CPU, the thread count matters. bench.py cpu sweeps batch size in fp32, and python bench.py cpu 4, 8 and 16 set the thread count (tokens per second):
threads batch 8 batch 16 batch 32 batch 64 4 3,261 2,857 2,951 3,192 8 4,125 4,077 3,838 3,629 16 3,570 3,590 3,087 2,791Sixteen threads (the hyperthreads) is slower than 8 physical cores, at every batch size. PyTorch already defaults to the physical core count on this machine (8), so do not raise --threads to the hyperthread count. Smaller batches also did a bit better at 8 threads. The box runs other services in the background, so treat the CPU numbers as having some noise.
What Fills 8 GB
The model itself is tiny. 5.8M parameters in fp32 is 23 MB. What fills VRAM is activations, and they scale with batch size. bench.py cuda measures both precisions. The peak column is torch.cuda.max_memory_allocated(), the memory PyTorch hands out to tensors:
precision batch tokens/s ms/step peak MiB fp32 8 90,696 22.6 403 fp32 16 93,226 43.9 722 fp32 32 100,406 81.6 1,361 fp32 64 104,009 157.5 2,638 fp32 128 106,198 308.6 5,194 fp32 256 OOM bf16 8 138,810 14.8 333 bf16 16 201,637 20.3 551 bf16 32 209,909 39.0 1,008 bf16 64 219,457 74.7 1,921 bf16 128 229,713 142.6 3,749 bf16 256 OOMThree readings:
- Bf16 roughly doubles throughput. About 100k tokens/s in fp32 against about 210k in bf16 at batch 32.
- Bf16 also cuts memory: 1,361 to 1,008 MiB at batch 32 (about 26%), and 5,194 to 3,749 MiB at batch 128 (about 28%).
- Bigger batches stop paying. Going from 32 to 128 adds under 10% speed and almost quadruples VRAM. Batch 256 runs out of memory at both precisions.
nvidia-smi shows more than that column. During a batch-32 bf16 training run it reported 1,291 MiB for the process: 1,008 MiB of tensors, 80 MiB of allocator cache on top, and the CUDA context. Budget about 1.5 GB for this model. An 8 GB card is plenty, and a 2 GB card would fit the training batch size. If you want to predict where a bigger model lands, GPU Memory Math walks through the arithmetic.
Overfitting: The Run That Went Too Long
The 1,500-step run ends at val 3.773. More steps must be better, right? Run C trains for 5,000 steps (44 passes over the data) on the GPU. The log below is abridged:
step 0 train 8.347 val 8.338step 250 train 4.464 val 4.689step 500 train 3.862 val 4.165step 750 train 3.549 val 3.959step 1000 train 3.311 val 3.882step 1250 train 3.095 val 3.860 <- best valstep 1500 train 2.891 val 3.862step 2000 train 2.530 val 3.943step 2500 train 2.208 val 4.064step 3000 train 1.906 val 4.212step 4000 train 1.449 val 4.435step 5000 train 1.237 val 4.57141.0M tokens seen (44.2 passes over the training set), training time 201sTraining loss keeps falling to 1.237. Validation loss bottoms out at step 1,250 (3.860) and then climbs to 4.571. The model has stopped learning how Holmes stories work and started memorizing this particular text. A model that has memorized its training set looks great on the training loss and gets worse on everything it has not seen.
If you only watched the train column, you would call this run a triumph. The val column is the honest one.
train.py saves a checkpoint only when val improves, so gpu.pt holds step 1,250, not step 5,000. Even so, this run’s best point (3.860) loses to the 1,500-step run (3.773). The reason is the schedule. At step 1,250 the long run’s learning rate is still 9.0e-4, five times the short run’s 1.7e-4. The long run fits the training text faster (train 3.095 vs 3.163) and generalizes worse (val 3.860 vs 3.791). The short run’s low learning rate at the end of its schedule earns the last bit of val loss.
The lesson: with small data, longer is not better. Size the schedule to the data. The fixes for overfitting are more data, more dropout, a smaller model, or stopping early. Here, the cheapest fix was a shorter schedule. Early stopping alone only got back to 3.860.
What It Learned
Samples printed during training, prompt “Holmes”, temperature 0.8, 40 tokens, whitespace collapsed. First the untrained model:
Holmes habitionanc obl dangerous children disc enough NoreverR towardsolmesorter�iterHere pers son said alternish occup|apedld mindwhereager reasoning lessiuszz needlp bearingribleittizeodyHalfway through the 1,500-step GPU run (step 750):
Holmes, and I waited there all I had to do so. I was absent about that the window of a thruce and a man who met me alping-night and IAt the end of that run (step 1,500):
Holmes, who has waited for all his time to get towards the window. It is clear that the window of the window, and his whole reasoning is not too much for the next.And the CPU run at step 1,500:
Holmes,” said the Inspector, “it would be alone to the conclusion that Horner occurred to the police, and that is where I was the greatest concentration of the problem whichA longer sample, from python sample.py gpu1500.pt "Holmes" --seed 1 (80 tokens):
Holmes rose to the door and stood by hisconfirmed eyes, sailly shortly at the door, with a wooden line oflamps which seemed to be rushing down the way within.
“On the first of the last train-street, Brother White,” murmured Holmes.“And now, though you can use your help.”
“ThenPunctuation, dialogue quotes, Holmes-isms like “murmured Holmes”, and sentence rhythm are all there. Meaning is not. The model loops on “the window of the window” and invents words like “sailly” by gluing BPE pieces together. That is the right result for 5.8M parameters and 1M tokens. sample.py is deliberately naive: it reruns the whole context for every new token. Part 4 fixes that.
The Embedding Check, Rerun
Part 1 ended with a promise. It printed cosine similarities for ” Holmes”, ” Watson”, and ” telegram” on an untrained table and got noise. python embeddings_after.py gpu1500.pt reruns the same check on the trained checkpoint. One note: Part 1 used its own N(0,1) table, while this post’s untrained column is this model’s own init (std 0.02, seed 1337), so the untrained numbers differ from Part 1’s.
pair untrained trained Holmes / Watson -0.006 +0.355 Holmes / telegram -0.045 +0.014 Watson / telegram +0.124 +0.201Holmes and Watson jumped from noise to +0.355. Holmes and telegram stayed near zero, which is correct: a detective and a piece of mail are different kinds of thing. The CPU checkpoint gives similar but not identical numbers (Holmes/Watson +0.329).
The nearest neighbours show how the model organizes the vocabulary:
nearest to ' Holmes': ' Lestrade' 0.62, ' Sherlock' 0.59, ' McMurdo' 0.52, ' Gregson' 0.50, ' Milverton' 0.49, ' Phelps' 0.44nearest to ' Watson': ' sir' 0.58, ' madam' 0.55, ' Doctor' 0.53, ' Mortimer' 0.51, ' gentlemen' 0.46, ' Lestrade' 0.46nearest to ' telegram': ' letter' 0.60, ' note' 0.57, ' narrative' 0.56, ' message' 0.55, ' photograph' 0.54, ' wire' 0.51nearest to ' London': ' England' 0.50, ' America' 0.50, ' town' 0.49, ' Europe' 0.44, ' South' 0.43, ' country' 0.41nearest to ' said': ' cried' 0.63, ' remarked' 0.61, ' observed' 0.55, ' says' 0.54, ' explained' 0.52, ' answered' 0.51Neighbours group by role. ” Holmes” sits with other named characters. ” Watson” sits with forms of address (” sir”, ” madam”, ” Doctor”), probably because characters address Watson the way they say “sir,”. ” telegram” sits with other documents, ” London” with places, ” said” with speech verbs. Nobody told the model any of this. Next-token prediction did.
What’s Next
We have a trained model and a sampler that is slow on purpose. Part 4 covers inference: a KV cache so each new token stops recomputing the whole context, and sampling strategies, with measured tokens per second. sample.py here reruns the full context at every token, which gives Part 4 a baseline to beat.
Common Questions
Can I train a small GPT on a CPU?
Yes. The 5.8M-parameter model trained for 1,500 steps on an 8-core CPU in 2,862 seconds (48 minutes) and reached val loss 3.762. The same run took 59 seconds on an RTX 3070 Laptop GPU. A CPU works for one-off runs, but it makes experimenting painfully slow.
How much VRAM do I need to train a small language model?
For this 5.8M-parameter model, about 1.5 GB. Training at batch 32 in bf16 used 1,291 MiB in nvidia-smi, of which PyTorch tensors took 1,008 MiB. The weights are only 23 MB in fp32. Activations fill the memory, and batch 256 ran out of memory on 8 GB at both precisions.
Why does validation loss go up while training loss goes down?
Overfitting. The model starts memorizing its training text instead of learning patterns that carry over to unseen text. In the 5,000-step run, training loss fell to 1.237 while validation loss rose from 3.860 at step 1,250 to 4.571. More data, more dropout, a smaller model, or early stopping fix it.
Should I use bf16 or fp32 for training on an RTX 30-series GPU?
Use bf16. On the RTX 3070 Laptop GPU, bf16 ran about 210,000 tokens per second against about 100,000 for fp32 at batch 32, and cut peak VRAM from 1,361 to 1,008 MiB. The final loss was 3.773 in bf16 against 3.762 in fp32, which is noise-level here.