Skip to content

17 technical books on transformers, deep learning, probability and the systems they run on. Every idea arrives three ways at once: the equation, the code, and a diagram you can touch.

That headline was written one word at a time, each picked from a probability distribution. Behind it is a loss surface: click anywhere to drop a ball and watch gradient descent find a valley.

17 volumes. Pull one off the shelf.

355 chapters and 1834 lessons, from the maths underneath to training models at scale.

Data Structures & Algorithms

Master data structures and algorithms from fundamentals to advanced topics. Interactive visualizations, multi-language implementations (C, Python, C++, TypeScript), and real-world applications with interview preparation.

8 chapters · 41 lessons

Open the book
See the whole catalogue

Look inside a real transformer.

This is the model the books build with you: a small GPT with every matrix and every neuron on screen. Step through the forward pass, switch to the neuron view, change the sentence. Nothing here is a video.

Open the lesson it comes from

One idea, three ways at once.

This is how every lesson is built. Pick a piece of the attention formula and watch the same piece light up in the code and in the picture.

The equation

softmax ⁣(QK⊤dk)V\mathrm{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V

The code

# one attention headscores = Q @ K.Tscores = scores / math.sqrt(d_k)weights = softmax(scores, dim=-1)out = weights @ V

The picture

÷√dkrows sum to 1outQK⊤

Every query is compared with every key: one score per pair of tokens.

Start with one chapter tonight.

Open the free sample chapter. No card needed.