Runbooks

Runbooks

Two runbooks for engineers who run AI systems in production.

Each runbook is a self-contained sequence of chapters. Start at chapter 00; every page links forward, back, and to a cheat sheet, interview questions and an FAQ.

16 chapters

LLM Inference Runbook

Tokenisation, the forward pass, KV cache, attention variants, quantisation, paging, multi-GPU serving, serving engines, the decode loop, capacity planning and production SLOs.

Start here →
17 chapters

RAG Runbook

Chunking, parsing hard content, identity and deletes, access control, embedding models, index structures, tuning, quantisation, sharding, filtering, hybrid retrieval and evaluation.

Start here →