Runbooks
Each runbook is a self-contained sequence of chapters. Start at chapter 00; every page links forward, back, and to a cheat sheet, interview questions and an FAQ.
Tokenisation, the forward pass, KV cache, attention variants, quantisation, paging, multi-GPU serving, serving engines, the decode loop, capacity planning and production SLOs.
Start here → 17 chaptersChunking, parsing hard content, identity and deletes, access control, embedding models, index structures, tuning, quantisation, sharding, filtering, hybrid retrieval and evaluation.
Start here →