{"product_id":"advanced-gpu-assembly-programming-third-gareth-thomas-9798184529363","title":"Advanced GPU Assembly Programming Third Edition: A Technical Reference for NVIDIA and AMD Architectures","description":"\u003cp\u003e\u003cb\u003eAdvanced GPU Assembly Programming\u003c\/b\u003e\u003c\/p\u003e\u003cp\u003e\u003cb\u003eMost GPU performance problems are not source-code problems. They are machine-code problems.\u003c\/b\u003e\u003c\/p\u003e\u003cp\u003eA kernel can look clean in CUDA or HIP and still lose the war at the hardware level.\u003c\/p\u003e\u003cp\u003eThe compiler may choose an instruction sequence you did not expect. A branch may split a warp or wavefront into masked paths. A load pattern may explode into extra memory transactions. A tensor pipeline may sit underfed while the code looks \"mathematically right.\" Occupancy may look healthy while register pressure, wait states, barriers, cache behavior, or issue slots quietly cap throughput.\u003c\/p\u003e\u003cp\u003eThat is where this book begins.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eAdvanced GPU Assembly Programming\u003c\/b\u003e is for advanced CUDA, HIP, AI-systems, HPC, and compiler engineers who need to read GPU machine code, understand NVIDIA and AMD execution behavior, and push kernels closer to the hardware performance ceiling.\u003c\/p\u003e\u003cp\u003eThis is not an introductory CUDA book.\u003c\/p\u003e\u003cp\u003eIt is not a beginner HIP guide.\u003c\/p\u003e\u003cp\u003eIt is not another surface-level explanation of \"parallel programming on GPUs.\"\u003c\/p\u003e\u003cp\u003eThis is a low-level technical reference for engineers who already understand kernels and now need to understand what those kernels become after compilation.\u003c\/p\u003e\u003cp\u003eIf you are optimizing AI inference, LLM kernels, GEMM, attention, scientific workloads, compiler output, CUDA-to-HIP portability, or architecture-specific performance, the question is no longer: \u003c\/p\u003e\u003cp\u003e\"Does the kernel run?\"\u003c\/p\u003e\u003cp\u003eThe question is: \u003c\/p\u003e\u003cp\u003e\u003cb\u003eWhat is the machine actually doing, and how close is it to the real limit?\u003c\/b\u003e\u003c\/p\u003e\u003cp\u003eInside, you will learn how to reason about: \u003c\/p\u003e\u003cul\u003e\n\u003cli\u003e\n\u003cb\u003eSIMT execution: \u003c\/b\u003e warps, wavefronts, active masks, divergence, reconvergence, predication, and independent thread scheduling\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eMachine code: \u003c\/b\u003e PTX, SASS, AMD ISA, instruction encoding, disassembly, source correlation, and compiler idioms\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eExecution resources: \u003c\/b\u003e SMs, CUs, schedulers, issue slots, scoreboards, barriers, wait states, register files, and occupancy limits\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eMemory behavior: \u003c\/b\u003e coalescing, global memory, shared memory, LDS, cache policy, HBM bandwidth, alignment, sectors, and bank conflicts\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eTensor and matrix pipelines: \u003c\/b\u003e tensor cores, MMA, MFMA, TMEM, FP8, FP6, FP4, block scaling, operand staging, and accumulator flow\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eAsynchronous execution: \u003c\/b\u003e cp.async, tensor-memory movement, producer\/consumer roles, barriers, staged pipelines, and latency hiding\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003ePerformance evidence: \u003c\/b\u003e Nsight Compute, Nsight Systems, rocprof, Radeon GPU Profiler, Omniperf, roofline analysis, and microbenchmarking\u003c\/li\u003e\n\u003cli\u003e\n\u003cb\u003eReal workloads: \u003c\/b\u003e high-performance GEMM, attention, reductions, scans, sparse computation, atomics, scatter\/gather, and irregular kernels\u003c\/li\u003e\n\u003c\/ul\u003e\u003cp\u003eThe value of this book is not that it tells you GPUs are fast.\u003c\/p\u003e\u003cp\u003eYou already know that.\u003c\/p\u003e\u003cp\u003eThe value is that it gives you the machinery to diagnose why a kernel is not fast enough.\u003c\/p\u003e\u003cp\u003eWhy did the compiler emit that instruction sequence?\u003c\/p\u003e\u003cp\u003eWhy did this memory access pattern create extra traffic?\u003c\/p\u003e\u003cp\u003eWhy are tensor units idle?\u003c\/p\u003e\u003cp\u003eWhy did a theoretically good tiling strategy lose throughput?\u003c\/p\u003e\u003cp\u003eWhy did NVIDIA and AMD behave differently?\u003c\/p\u003e\u003cp\u003eWhy did a change that looked harmless at source level move the bottleneck somewhere else?\u003c\/p\u003e\u003cp\u003eThis book helps you connect source code, compiler decisions, disassembly, profiler counters, memory transactions, lane masks, and architectural constraints into one coherent performance model.\u003c\/p\u003e\u003cp\u003e\u003cb\u003eAdvanced GPU Assembly Programming\u003c\/b\u003e was written for the engineer who wants the layer beneath CUDA, HIP, Triton, compiler output, and vendor libraries.\u003c\/p\u003e\u003cbr\u003e\u003cbr\u003e\u003cb\u003eAuthor:\u003c\/b\u003e Gareth Thomas\u003cbr\u003e\u003cb\u003eISBN-13:\u003c\/b\u003e 9798184529363\u003cbr\u003e\u003cb\u003ePublisher:\u003c\/b\u003e Independently Published\u003cbr\u003e\u003cb\u003eLanguage:\u003c\/b\u003e English\u003cbr\u003e\u003cb\u003ePublished:\u003c\/b\u003e 07\/08\/2026\u003cbr\u003e\u003cb\u003ePages:\u003c\/b\u003e 386\u003cbr\u003e\u003cb\u003eFormat:\u003c\/b\u003e Paperback\u003cbr\u003e\u003cb\u003eWeight:\u003c\/b\u003e 1.97lbs\u003cbr\u003e\u003cb\u003eSize:\u003c\/b\u003e 11.00h x 8.50w x 0.80d","brand":"Gareth Thomas","offers":[{"title":"Paperback","offer_id":48997886853375,"sku":"9798184529363","price":33.97,"currency_code":"USD","in_stock":true}],"url":"https:\/\/www.whiterainbookhouse.com\/products\/advanced-gpu-assembly-programming-third-gareth-thomas-9798184529363","provider":"WR Book House","version":"1.0","type":"link"}