Research

Watch · narrated walkthroughs

The Math Behind Modern LLM Inference

The algorithms that make a large language model answer on demand — attention and the KV cache, speculative decoding, quantization error, and vector-index recall — derived properly, with the complexity and the point where each one breaks. Grounded in the primary papers.

Murali Chillakuru·4 episodes
  1. 1
  2. 2
  3. 3
  4. 4