GPU Bit Tricks: LLM Kernel Arithmetic Optimizations (Generative AI LLM Programming)
Buy Now, Pay Later
- – 4-month term
- – No impact on credit to apply
- – Instant approval decision
- – Secure and straightforward checkout
Ready to go? Add this product to your cart and select a plan during checkout.
Payment plans are offered through our trusted finance partners Klarna, Affirm, Afterpay, Zip, Apple Pay, and Google Pay. No-credit-needed leasing options through Acima may also be available at checkout.
Learn more about financing & leasing here.
This item is eligible for return within 30 days of receipt
To qualify for a full refund, items must be returned in their original, unused condition. If an item is returned in a used, damaged, or materially different state, you may be granted a partial refund.
To initiate a return, please visit our Returns Center.
View our full returns policy here.
Description
GPU Bit Tricks: LLM Kernel Arithmetic Optimizations: For LLM engineers writing device kernels in CUDA C++, this book covers all sorts of low-level bit tricks, arithmetic optimizations, and approximations to make things faster. Highlights: - Integer bitwise magic tricks - Floating-point bit manipulation - Approximations and other magic - Mathematical manipulations - Blackwell hardware exploitation Table of Contents: Part I: LLM Kernel Arithmetic 1. LLM Arithmetic Overview 2. Blackwell GPU Optimizations 3. Rubin and Feynman GPU Optimizations 4. Grace CPU Optimizations 5. Low-Bit Quantization 6. Block- Scaled Quantization (BSQ) 7. Bitshift Quantization 8. Zero-Multiplication Models Part II: Integer Optimizations 9. Bitwise Operator Overview 10. Bitwise Intrinsics 11. Bit Fiddling Magic 12. Bit Data Structures Part III: Floating- Point Optimizations 13. Floating-Point Bit Formats 14. Floating-Point Bit Magic 15. Fixed-Point Numbers 16. Block Floating Point (BFP)Part IV: Arithmetic Optimizations 17. Arithmetic Optimizations 18. Approximate Arithmetic 19. Math Tricks 20. BF16x9 Emulation 21. FP64 Emulation 22. Vector Magnitude Approximations 23. Integer Overflow, Saturation and ClampingPart V: Branchless Arithmetic 24. Branchless Coding Techniques 25. Branchless Pitfalls 26. Branchless Conditional TestsPart VI: Advanced Optimizations 27. SWAR Techniques 28. Polynomial Approximation of Softmax 29. Polynomial Approximation in Tensor Cores 30. Advanced Number SystemsAppendix A: Long List of Low Latency Techniques Appendix B. 500 LLM Inference Optimization Techniques Appendix C. 200 CUDA C++ Optimization Techniques Read more
Accessibility : Learn more
Publication date : June 12, 2026
Language : English
File size : 1.4 MB
Screen Reader : Supported
Enhanced typesetting : Enabled
X-Ray : Not Enabled
Word Wise : Not Enabled
Print length : 474 pages
Frequently asked questions
To initiate a return, please visit our Returns Center.
View our full returns policy here.
- Klarna Financing
- Affirm Pay in 4
- Affirm Financing
- Afterpay Financing
- Zip Pay in 4
- Financing through Apple Pay
- Financing through Google Pay
Learn more about financing & leasing here.