Premium Content

Tokenisation and Vocabularies

The full tokenisation pipeline from raw text to token IDs, and the design decisions that affect model quality.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Why Subword Tokenisation?
  • 2. Byte-Pair Encoding from Scratch
  • 3. SentencePiece and Unigram LM Tokenisation
  • 4. Vocabulary Size Trade-offs
  • 5. Special Tokens and Chat Templates
  • 6. Tokenisation Pitfalls
Subscribe to Unlock

Already have an account? Sign in