Past the basics of attention and into the architecture: what heads are, why order needs injecting, and how residual connections let depth train at all.
3 units, about 3 hours. Free.
Compute attention weights by hand and explain why several small heads beat one large one.
2Demonstrate that attention is order-blind and fix it with positional encodings.
3Show numerically why deep stacks need residual connections and what layer normalisation stabilises.