Premium Content

Multi-Head Latent Attention (MLA)

DeepSeek-V2 2024 — compressed KV-cache via learned bottleneck

This chapter requires a subscription to access.

What you'll unlock:

  • 1. Multi-Head Latent Attention (MLA)
Subscribe to Unlock

Already have an account? Sign in