{"product_id":"building-and-customizing-inference-engines-ranadhir-ghosh-9798185385777","title":"Building and Customizing Inference Engines for LLMs: From First Principles to Production - A Complete Guide to Designing High-Performance, Efficient,","description":"Open the hood of modern AI.\u003cp\u003eEveryone is learning to \u003ci\u003euse\u003c\/i\u003e large language models. Almost no one understands how they actually \u003ci\u003erun\u003c\/i\u003e. This book closes that gap - carrying you, without a single hand-wave, from \u003cb\u003e\"what is a token?\"\u003c\/b\u003e to a working, batched, quantized, multi-GPU inference server you built yourself.\u003c\/p\u003e\u003cp\u003eAn LLM's weights hold the intelligence, but they do nothing on their own. The \u003cb\u003einference engine\u003c\/b\u003e is the software that decides how those billions of numbers are moved, cached, batched, and multiplied against a stream of concurrent users. Get it wrong and a state-of-the-art model crawls and falls over at ten users. Get it right and the \u003ci\u003esame weights on the same GPU\u003c\/i\u003e serve dozens - at a fraction of the latency and cost.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eWhat you will learn: \u003c\/b\u003e\u003c\/p\u003e\u003cul\u003e\n\u003cli\u003eWhy decoding is \u003cb\u003ememory-bandwidth-bound\u003c\/b\u003e - the one law that explains PagedAttention, quantization, and FlashAttention as inevitable consequences, not tricks.\u003c\/li\u003e\n\u003cli\u003eThe full anatomy of an engine: tokenizer, scheduler \u0026amp; continuous batching, memory manager, KV cache \u0026amp; PagedAttention, quantization (INT8\/INT4\/FP8), sampler, and Mixture-of-Experts routing.\u003c\/li\u003e\n\u003cli\u003eThe hot path: fused GPU kernels and FlashAttention, writing your own kernels in Triton\/CUDA\/Rust, speculative decoding, and multi-GPU tensor\/pipeline\/expert parallelism.\u003c\/li\u003e\n\u003cli\u003eReal engines dissected: \u003cb\u003evLLM, TensorRT-LLM, llama.cpp, MLC-LLM, and SGLang\u003c\/b\u003e - and how to deploy Llama, Qwen, and DeepSeek on each.\u003c\/li\u003e\n\u003cli\u003eProduction skills: benchmarking, profiling, cost\/SLO capacity planning, and eighteen-plus real-world war stories.\u003c\/li\u003e\n\u003c\/ul\u003e\u003cp\u003e\u003cb\u003eBuilt for three altitudes.\u003c\/b\u003e Beginners get the concept and the mental model; intermediate engineers get the algorithms, the math, and runnable Python; advanced engineers get kernel-level detail, C++\/Rust considerations, failure modes, and the frontier.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eWho it's for: \u003c\/b\u003e LLM \u0026amp; inference engineers, AI platform engineers and architects, forward-deployed engineers, researchers, and technology leaders who must reason about latency, throughput, memory, and cost - not just call an API.\u003c\/p\u003e\u003cp\u003eYou will build \u003cb\u003eDwarfStar\u003c\/b\u003e, a minimal but real inference engine, growing it stage by stage into a batched, quantized, C++\/CUDA-accelerated server. Close this book and you will open the source of vLLM or TensorRT-LLM and feel at home.\u003c\/p\u003e\u003cbr\u003e\u003cbr\u003e\u003cb\u003eAuthor:\u003c\/b\u003e Ranadhir Ghosh\u003cbr\u003e\u003cb\u003eISBN-13:\u003c\/b\u003e 9798185385777\u003cbr\u003e\u003cb\u003ePublisher:\u003c\/b\u003e Independently Published\u003cbr\u003e\u003cb\u003eLanguage:\u003c\/b\u003e English\u003cbr\u003e\u003cb\u003ePublished:\u003c\/b\u003e 07\/04\/2026\u003cbr\u003e\u003cb\u003ePages:\u003c\/b\u003e 434\u003cbr\u003e\u003cb\u003eFormat:\u003c\/b\u003e Paperback\u003cbr\u003e\u003cb\u003eWeight:\u003c\/b\u003e 1.65lbs\u003cbr\u003e\u003cb\u003eSize:\u003c\/b\u003e 10.00h x 7.00w x 0.88d","brand":"Ranadhir Ghosh","offers":[{"title":"Paperback","offer_id":49174189113599,"sku":"9798185385777","price":25.0,"currency_code":"USD","in_stock":true}],"url":"https:\/\/www.whiterainbookhouse.com\/products\/building-and-customizing-inference-engines-ranadhir-ghosh-9798185385777","provider":"WR Book House","version":"1.0","type":"link"}