Premium Content

Inference Optimisation

Prefill/decode split, KV cache management, speculative decoding, quantisation, and expert load balancing.

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Two Very Different Problems: Prefill vs Decode
  • 2. KV Cache Management and PagedAttention
  • 3. Speculative Decoding with MTP
  • 4. Post-Training Quantisation
  • 5. Expert Load Balancing at Inference
Subscribe to Unlock

Already have an account? Sign in