LLM Inference Math1Attention and the KV Cache: The Arithmetic of a Forward PassWatch2Speculative Decoding: The Accept-Reject Math of Faster TokensWatch3Quantization Error: How Low Can the Bits Go Before Quality BreaksWatch4Vector Indexes: The Recall-vs-Latency Trade-off of HNSW and IVFWatch