The Attention Atlas: Mechanisms That Power Modern AI
A Visual, Mathematical, and Practical Guide
Master attention from first principles through modern Transformer practice: seq2seq motivation, Q/K/V intuition, 15 core mechanisms, BERT/GPT/ViT case studies, visualization, implementation labs, efficiency trade-offs, and a final tiny-Transformer capstone.
Foundation— Origin, seq2seq bridge, shared example, and core attention.
The Origin of Attention
From selective focus to neural sequence models — how resource allocation became a useful way to think about attention in modern AI
Scaled Dot-Product Attention
Vaswani et al. 2017 — the foundational attention mechanism
Attention Forms— Multi-head, causal, and cross-attention.
Multi-Head Attention
Vaswani et al. 2017 — H independent attention heads in parallel
Causal (Masked) Self-Attention
Radford et al. 2018 — autoregressive masking for language generation
Cross-Attention
Vaswani et al. 2017 — the bridge between encoder and decoder
Transformer Architecture— Residuals, normalization, feed-forward layers, and masks.
Transformer Blocks: Putting Attention in Context
How attention fits inside real encoder and decoder layers with residual connections, layer normalization, feed-forward networks, and masks
KV-Cache Optimization— MQA, GQA for efficient inference.
Multi-Query Attention (MQA)
Shazeer 2019 — shared K/V across all heads for fast inference
Grouped-Query Attention (GQA)
Ainslie et al. 2023 — the sweet spot between MHA and MQA
Positional Encoding— Relative bias, RoPE, ALiBi.
Relative Position Bias Attention
Shaw et al. 2018 / T5 2020 — distance-aware attention scoring
RoPE — Rotary Position Embedding
Su et al. 2021 — position via rotation of Q and K vectors
ALiBi — Attention with Linear Biases
Press et al. 2022 — linear distance penalty, no positional embeddings
Efficient Attention— Linear, Sliding Window, Sparse.
Linear Attention
Katharopoulos et al. 2020 — O(Nd²) instead of O(N²d)
Sliding Window Attention
Beltagy et al. 2020 — local window for O(N) complexity
Sparse Attention — BigBird
Zaheer et al. 2020 — local + global + random for long documents
Modern Innovations— Flash, Differential, MLA.
Flash Attention
Dao et al. 2022 — IO-aware tiling of the same attention operator
Differential Attention
Microsoft Research 2024 — noise cancellation via dual softmax
Multi-Head Latent Attention (MLA)
DeepSeek-V2 2024 — compressed KV-cache via learned bottleneck
Comparison— All 15 mechanisms side-by-side.
Mechanism Comparison
Fourteen mechanism variants plus one Flash execution reference side-by-side
Model Families and Interpretation— BERT, GPT, ViT, and responsible attention analysis.
Attention in BERT, GPT, and Vision Transformers
How encoder-only, decoder-only, and vision Transformer models use the same attention equation differently
Visualizing and Interpreting Attention
How to read attention heatmaps, inspect heads responsibly, and avoid treating attention weights as complete explanations
Implementation Labs— From-scratch code and decision guidance.
Implementation Lab: Attention from Scratch
Build scaled dot-product attention, multi-head attention, and a mini Transformer block with shape checks and debugging guidance
Limitations and Efficient-Attention Roadmap
Why standard attention becomes expensive, how efficient variants trade accuracy, memory, speed, and implementation complexity
Capstone— Build and inspect a tiny Transformer.
Capstone: Build and Inspect a Tiny Transformer
An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block
Finish by building a tiny Transformer.
The capstone ties together token IDs, embeddings, positional information, attention weights, a mini Transformer block, shape checks, and interpretation questions.
Capstone: Build and Inspect a Tiny Transformer
An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block
Open chapter25 sections. Begin with one.
Chapter 0 — The Origin of Attention — is where every reader starts.