← All books
Book · Intermediate · 13+ hours

The Attention Atlas: Mechanisms That Power Modern AI

A Visual, Mathematical, and Practical Guide

Master attention from first principles through modern Transformer practice: seq2seq motivation, Q/K/V intuition, 15 core mechanisms, BERT/GPT/ViT case studies, visualization, implementation labs, efficiency trade-offs, and a final tiny-Transformer capstone.

23Chapters
25Sections
14hReading
11Parts
Part I·2 chapters · 3 sections

FoundationOrigin, seq2seq bridge, shared example, and core attention.

The Origin of Attention

From selective focus to neural sequence models — how resource allocation became a useful way to think about attention in modern AI

2 sections58 min read
Start chapter
  1. 01The Origin of Attention30m
  2. 02From Seq2Seq Bottlenecks to Attention28m

Scaled Dot-Product Attention

Vaswani et al. 2017 — the foundational attention mechanism

1 section60 min read
Start chapter
  1. 01Scaled Dot-Product Attention60m
Part II·3 chapters · 3 sections

Attention FormsMulti-head, causal, and cross-attention.

Multi-Head Attention

Vaswani et al. 2017 — H independent attention heads in parallel

1 section35 min read
Start chapter
  1. 01Multi-Head Attention35m

Causal (Masked) Self-Attention

Radford et al. 2018 — autoregressive masking for language generation

1 section40 min read
Start chapter
  1. 01Causal (Masked) Self-Attention40m

Cross-Attention

Vaswani et al. 2017 — the bridge between encoder and decoder

1 section35 min read
Start chapter
  1. 01Cross-Attention35m
Part III·1 chapter · 1 sections

Transformer ArchitectureResiduals, normalization, feed-forward layers, and masks.

Transformer Blocks: Putting Attention in Context

How attention fits inside real encoder and decoder layers with residual connections, layer normalization, feed-forward networks, and masks

1 section25 min read
Start chapter
  1. 01Transformer Blocks: Putting Attention in Context25m
Part IV·2 chapters · 2 sections

KV-Cache OptimizationMQA, GQA for efficient inference.

Multi-Query Attention (MQA)

Shazeer 2019 — shared K/V across all heads for fast inference

1 section30 min read
Start chapter
  1. 01Multi-Query Attention (MQA)30m

Grouped-Query Attention (GQA)

Ainslie et al. 2023 — the sweet spot between MHA and MQA

1 section30 min read
Start chapter
  1. 01Grouped-Query Attention (GQA)30m
Part V·3 chapters · 4 sections

Positional EncodingRelative bias, RoPE, ALiBi.

Relative Position Bias Attention

Shaw et al. 2018 / T5 2020 — distance-aware attention scoring

2 sections53 min read
Start chapter
  1. 01Why Attention Needs Position18m
  2. 02Relative Position Bias Attention35m

RoPE — Rotary Position Embedding

Su et al. 2021 — position via rotation of Q and K vectors

1 section30 min read
Start chapter
  1. 01RoPE — Rotary Position Embedding30m

ALiBi — Attention with Linear Biases

Press et al. 2022 — linear distance penalty, no positional embeddings

1 section30 min read
Start chapter
  1. 01ALiBi — Attention with Linear Biases30m
Part VI·3 chapters · 3 sections

Efficient AttentionLinear, Sliding Window, Sparse.

Linear Attention

Katharopoulos et al. 2020 — O(Nd²) instead of O(N²d)

1 section32 min read
Start chapter
  1. 01Linear Attention32m

Sliding Window Attention

Beltagy et al. 2020 — local window for O(N) complexity

1 section28 min read
Start chapter
  1. 01Sliding Window Attention28m

Sparse Attention — BigBird

Zaheer et al. 2020 — local + global + random for long documents

1 section30 min read
Start chapter
  1. 01Sparse Attention — BigBird30m
Part VII·3 chapters · 3 sections

Modern InnovationsFlash, Differential, MLA.

Flash Attention

Dao et al. 2022 — IO-aware tiling of the same attention operator

1 section45 min read
Start chapter
  1. 01Flash Attention45m

Differential Attention

Microsoft Research 2024 — noise cancellation via dual softmax

1 section30 min read
Start chapter
  1. 01Differential Attention30m

Multi-Head Latent Attention (MLA)

DeepSeek-V2 2024 — compressed KV-cache via learned bottleneck

1 section35 min read
Start chapter
  1. 01Multi-Head Latent Attention (MLA)35m
Part VIII·1 chapter · 1 sections

ComparisonAll 15 mechanisms side-by-side.

Mechanism Comparison

Fourteen mechanism variants plus one Flash execution reference side-by-side

1 section45 min read
Start chapter
  1. 01Fourteen Mechanisms and One Execution Reference45m
Part IX·2 chapters · 2 sections

Model Families and InterpretationBERT, GPT, ViT, and responsible attention analysis.

Attention in BERT, GPT, and Vision Transformers

How encoder-only, decoder-only, and vision Transformer models use the same attention equation differently

1 section25 min read
Start chapter
  1. 01Attention in BERT, GPT, and Vision Transformers25m

Visualizing and Interpreting Attention

How to read attention heatmaps, inspect heads responsibly, and avoid treating attention weights as complete explanations

1 section22 min read
Start chapter
  1. 01Visualizing and Interpreting Attention22m
Part X·2 chapters · 2 sections

Implementation LabsFrom-scratch code and decision guidance.

Implementation Lab: Attention from Scratch

Build scaled dot-product attention, multi-head attention, and a mini Transformer block with shape checks and debugging guidance

1 section35 min read
Start chapter
  1. 01Implementation Lab: Attention from Scratch35m

Limitations and Efficient-Attention Roadmap

Why standard attention becomes expensive, how efficient variants trade accuracy, memory, speed, and implementation complexity

1 section25 min read
Start chapter
  1. 01Limitations and Efficient-Attention Roadmap25m
Part XI·1 chapter · 1 sections

CapstoneBuild and inspect a tiny Transformer.

Capstone: Build and Inspect a Tiny Transformer

An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block

1 section35 min read
Start chapter
  1. 01Capstone: Build and Inspect a Tiny Transformer35m
The capstone

Finish by building a tiny Transformer.

The capstone ties together token IDs, embeddings, positional information, attention weights, a mini Transformer block, shape checks, and interpretation questions.

Chapter 22·1 sections

Capstone: Build and Inspect a Tiny Transformer

An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block

Open chapter

25 sections. Begin with one.

Chapter 0 — The Origin of Attention — is where every reader starts.