Skip to content
← All books
Book · Intermediate · 15+ hours

The Attention Atlas: Mechanisms That Power Modern AI

A Visual, Mathematical, and Practical Guide

Learn major attention families from first principles: Q/K/V, Transformer blocks, positional methods, sparse MoBA routing, DeltaNet and Kimi Delta Attention, Flash execution, Kimi hybrid case studies, and a real tiny-Transformer capstone.

23Chapters
30Sections
20hReading
8Parts
Curriculum

23 chapters— in publication order.

The capstone

Finish by building a tiny Transformer.

The capstone ties together token IDs, embeddings, positional information, attention weights, a mini Transformer block, shape checks, and interpretation questions.

Chapter 22·1 sections

Capstone: Build and Inspect a Tiny Transformer

An end-to-end beginner project that tokenizes a tiny sequence, runs attention, inspects weights, and passes data through a mini Transformer block

Open chapter

30 sections. Begin with one.

Chapter 0 — The Origin of Attention — is where every reader starts.