{"product_id":"deep-dive-into-sglang-volume-yin-yang-9798182198233","title":"Deep Dive into SGLang, Volume I: Runtime, Scheduling, Memory, and Decoding","description":"\u003cp\u003eDeep Dive into SGLang, Volume I explains SGLang's runtime request path as a set of algorithms, data structures, and systems tradeoffs. \u003c\/p\u003e\u003cp\u003e\u003c\/p\u003eThis volume focuses on the core inference-serving loop: tokenization, admission, prefill and decode scheduling, prefix lookup, KV-cache allocation, forward execution, sampling, streaming, and cleanup. Instead of treating SGLang as a catalog of commands, it builds a working model of the runtime state that makes modern LLM serving possible. \u003cp\u003e\u003c\/p\u003eVolume I covers: \u003cul\u003e\n\u003cli\u003eThe inference serving problem and request lifecycle\u003c\/li\u003e\n\u003cli\u003eTransformer inference cost models\u003c\/li\u003e\n\u003cli\u003eSGLang's runtime as a distributed state machine\u003c\/li\u003e\n\u003cli\u003eContinuous batching and chunked prefill\u003c\/li\u003e\n\u003cli\u003eKV-cache memory management\u003c\/li\u003e\n\u003cli\u003eRadixAttention and prefix reuse\u003c\/li\u003e\n\u003cli\u003eHierarchical caching\u003c\/li\u003e\n\u003cli\u003eAttention backends and forward batches\u003c\/li\u003e\n\u003cli\u003eModel architecture registration\u003c\/li\u003e\n\u003cli\u003eSampling and logits processing\u003c\/li\u003e\n\u003cli\u003eStructured outputs, reasoning parsers, and tool-call parsing\u003c\/li\u003e\n\u003cli\u003eSpeculative decoding\u003c\/li\u003e\n\u003cli\u003eDiffusion language models and blockwise decoding\u003c\/li\u003e\n\u003c\/ul\u003e\u003cbr\u003eEach chapter explains the systems problem before implementation details, then ties technical claims back to SGLang source files, tests, or benchmark artifacts. The book is written for readers who already know basic transformer and GPU-systems vocabulary and want a deeper, source-grounded understanding of how a production LLM serving runtime is organized. \u003cp\u003e\u003c\/p\u003eThis is Volume I of Deep Dive into SGLang. Volume II continues into quantization, distributed serving, kernels, hardware backends, benchmarking, correctness, multimodal serving, diffusion inference, and technical extension. \u003cp\u003e\u003c\/p\u003eIndependent explanatory guide. Not affiliated with or endorsed by the SGLang project.\u003cbr\u003e\u003cbr\u003e\u003cb\u003eAuthor:\u003c\/b\u003e Yin Yang\u003cbr\u003e\u003cb\u003eISBN-13:\u003c\/b\u003e 9798182198233\u003cbr\u003e\u003cb\u003ePublisher:\u003c\/b\u003e Independently Published\u003cbr\u003e\u003cb\u003eLanguage:\u003c\/b\u003e English\u003cbr\u003e\u003cb\u003ePublished:\u003c\/b\u003e 06\/21\/2026\u003cbr\u003e\u003cb\u003ePages:\u003c\/b\u003e 544\u003cbr\u003e\u003cb\u003eFormat:\u003c\/b\u003e Paperback\u003cbr\u003e\u003cb\u003eWeight:\u003c\/b\u003e 2.74lbs\u003cbr\u003e\u003cb\u003eSize:\u003c\/b\u003e 11.00h x 8.50w x 1.10d","brand":"Yin Yang","offers":[{"title":"Paperback","offer_id":49148599369983,"sku":"9798182198233","price":29.99,"currency_code":"USD","in_stock":true}],"url":"https:\/\/www.whiterainbookhouse.com\/products\/deep-dive-into-sglang-volume-yin-yang-9798182198233","provider":"WR Book House","version":"1.0","type":"link"}