Research

Watch · narrated whiteboard episodesL2

The Math Behind Modern LLM Inference

The algorithms that make a large language model answer on demand — attention and the KV cache, speculative decoding, quantization error, and vector-index recall — derived properly, with the complexity and the point where each one breaks. Grounded in the primary papers.

Murali Chillakuru·4 episodes
  1. 11 min Episode 1Attention and the KV Cache: The Arithmetic of a Forward PassA whiteboard derivation of why serving a language model is a memory problem — from the attention equation to the exact bytes the KV cache holds.
  2. 12 min Episode 2Speculative Decoding: The Accept-Reject Math of Faster TokensA whiteboard walkthrough of how a cheap draft model and a rejection-sampling rule cut latency while leaving the output distribution untouched.
  3. 10 min Episode 3Quantization Error: How Low Can the Bits Go Before Quality BreaksA whiteboard walkthrough of the arithmetic of rounding a model's weights to a coarse grid, why each bit buys about six decibels, and where the error finally wins.
  4. 12 min Episode 4Vector Indexes: The Recall-vs-Latency Trade-off of HNSW and IVFA whiteboard walkthrough of why approximate nearest-neighbour search is a single dial between how many true neighbours you find and how long you wait, and how IVF and HNSW set that dial.