The Attention Atlas: Mechanisms That Power Modern AI
A Visual, Mathematical, and Practical Guide
Learn major attention families from first principles: Q/K/V, Transformer blocks, positional methods, sparse MoBA routing, DeltaNet and Kimi Delta Attention, Flash execution, Kimi hybrid case studies, and a real tiny-Transformer capstone.
23 chapters— in publication order.
Input · Chapter 00 · 4 sections · 98 min
Why did machines need attention, and what does a whole Transformer do with it?
The Origin of Attention
- 0.1The Origin of Attention25m
- 0.2From Seq2Seq Bottlenecks to Attention18m
- 0.3The Whole Forward Pass, Matrix by Matrix30m
- 0.4The Same Network, Neuron by Neuron25m
Block · Chapter 01 · 1 section · 45 min
How does one word decide which other words matter, and how much?
Scaled Dot-Product Attention
Block · Chapter 02 · 1 section · 40 min
How can one word look for several different things at the same time?
Multi-Head Attention
Start chapter 2Block · Chapter 03 · 1 section · 40 min
How does a model learn to write one word at a time without peeking at the words that come later?
Causal (Masked) Self-Attention
Block · Chapter 04 · 1 section · 35 min
How can one text look things up in another, such as a translation reading its source?
Cross-Attention
Start chapter 4Block · Chapter 05 · 1 section · 50 min
What happens to the output of attention inside a full layer of the model?
Transformer Blocks: Putting Attention in Context
Start chapter 5Memory · Chapter 06 · 1 section · 40 min
Can all heads share one set of keys and values, so the model remembers less?
Multi-Query Attention (MQA)
Memory · Chapter 07 · 1 section · 40 min
How many sets of keys and values does a model really need?
Grouped-Query Attention (GQA)
Position · Chapter 08 · 2 sections · 75 min
Why does attention ignore word order on its own, and how can it learn distance?
Relative Position Bias Attention
Position · Chapter 09 · 1 section · 40 min
How can turning queries and keys by an angle tell attention how far apart two words are?
RoPE — Rotary Position Embedding
Position · Chapter 10 · 1 section · 40 min
How can a model trained on short texts keep working on longer ones?
ALiBi — Attention with Linear Biases
Long context · Chapter 11 · 2 sections · 80 min
Can a token gather context without comparing itself with every other token?
Linear Attention
Long context · Chapter 12 · 1 section · 40 min
What if each token reads only a window of recent tokens?
Sliding Window Attention
Long context · Chapter 13 · 2 sections · 80 min
How can a few well-chosen links reach across a long text?
Sparse Attention — BigBird
New ideas · Chapter 14 · 1 section · 40 min
How can a GPU compute exactly the same attention without writing down the whole table of scores?
Flash Attention
New ideas · Chapter 15 · 1 section · 40 min
How can attention give less weight to words that do not help?
Differential Attention
New ideas · Chapter 16 · 2 sections · 85 min
How can a model store fewer numbers than the keys and values themselves?
Multi-Head Latent Attention (MLA)
Compare · Chapter 17 · 1 section · 45 min
What does every method in this book do to the same word?
Mechanism Comparison
Start chapter 17Compare · Chapter 18 · 1 section · 40 min
How does one formula fill in hidden words, write text and look at pictures?
Attention in BERT, GPT, and Vision Transformers
Start chapter 18Compare · Chapter 19 · 1 section · 40 min
What can an attention map honestly tell you, and what can it not?
Visualizing and Interpreting Attention
Start chapter 19Output · Chapter 20 · 1 section · 45 min
How do you check that your own attention code is right?
Implementation Lab: Attention from Scratch
Start chapter 20Output · Chapter 21 · 1 section · 50 min
What limits remain when the text grows from 5 tokens to a million?
Limitations and Efficient-Attention Roadmap
Start chapter 21Output · Chapter 22 · 1 section · 60 min
Can you build a small Transformer yourself and look inside it?
Capstone: Build and Inspect a Tiny Transformer
Start chapter 22Finish by building a tiny Transformer.
The capstone ties together token IDs, embeddings, positional information, attention weights, a mini Transformer block, shape checks, and interpretation questions.
Capstone: Build and Inspect a Tiny Transformer
An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block
Open chapter30 sections. Begin with one.
Chapter 0 — The Origin of Attention — is where every reader starts.