Further reading
The papers and posts worth your time
The primary sources. This field moves fast enough that the papers are often more current than anything written about them, and most of these are readable in an evening.
Architecture
The papers the field is built on, from the original transformer through diffusion and multimodal encoders.
- “Attention is All You Need”by Ashish Vaswani et al. (Neural Information Processing Systems, 2017)
- “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”by Jacob Devlin et al. (North American Chapter of the Association for Computational Linguistics, 2019)
- “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models”by Junnan Li et al. (International Conference on Machine Learning, 2023)
- Deep Learningby Ian Goodfellow, Yoshua Bengio, and Aaron Courville (The MIT Press, 2016)
- Deep Learning with Python (2nd Edition)by François Chollet (Manning, 2021)
- “Denoising Diffusion Probabilistic Models”by Jonathan Ho et al. (ArXiv abs/2006.11239, 2020)
- “DiT: Scalable Diffusion Models with Transformers”by William Peebles and Saining Xie (International Conference on Computer Vision (ICCV), 2023)
- “FlashAttention: Fast and Memory-Efficient Exact Attention with IO Awareness”by Tri Dao et al. (ArXiv abs/2205.14135, 2022)
- “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”by Tri Dao (ArXiv abs/2307.08691, 2023)
- “FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision”by Jay Shah et al. (ArXiv abs/2407.08608, 2024)
- Flash-Attention-4by Tri Dao (Dao AI Research Lab, 2025)
- “Imagen Video: High Definition Video Generation with Diffusion Models”by Jonathan Ho et al. (ArXiv abs/2210.02303, 2022)
- “Language Models Are Few-Shot Learners”by Tom Brown et al. (ArXiv abs/2005.14165)
- “Learning Transferable Visual Models from Natural Language Supervision”by Alec Radford et al. (International Conference on Machine Learning, 2021)
- “Longformer: The Long-Document Transformer”by Iz Beltagy et al. (ArXiv abs/2004.05150, 2020)
- “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”by Albert Gu and Tri Dao (ArXiv abs/2312.00752, 2023)
- “Matryoshka Representation Learning”by Aditya Kusupati et al. (Neural Information Processing Systems, 2022)
- “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer”by Noam Shazeer et al. (ArXiv abs/1701.06538, 2017)
- “Reformer: The Efficient Transformer”by Nikita Kitaev et al. (ArXiv abs/2001.04451, 2020)
- “Robust Speech Recognition via Large-Scale Weak Supervision”by Alec Radford et al. (International Conference on Machine Learning, 2022)
- “RoFormer: Enhanced Transformer with Rotary Position Embedding”by Jianlin Su et al. (ArXiv abs/2104.09864, 2021)
- “SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis”by Dustin Podell et al. (ArXiv abs/2307.01952, 2023)
- “Segment Anything”by Alexander Kirillov et al. (2023 IEEE/CVF International Conference on Computer Vision (ICCV), 2023)
- “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks”by Nils Reimers and Iryna Gurevych (ArXiv abs/1908.10084, 2019)
- “The Llama 3 Herd of Models”by Aaron Grattafiori et al. (ArXiv 2407.21783, 2024)
- “Video Diffusion Models”by Jonathan Ho et al. (ArXiv abs/2204.03458, 2022)
- “Visual Instruction Tuning”by Haotian Liu et al. (ArXiv abs/2304.08485, 2023)
Developer Tools
Engines, frameworks, and the libraries you will actually import.
- BitsAndBytes by bitsandbytes-foundation
- ComfyUI by comfyanonymous
- “CUDA by Example: An Introduction to General Purpose GPU Programming”by Jason Sanders and Edward Kandrot (NVIDIA developer, 2025)
- “CUDA C++ Programming Guide Release 13.0”by NVIDIA (2025)
- “CUDA cuBLAS Release 13.0”by NVIDIA (2025)
- CUTLASS by NVIDIA
- DeepGEMM by DeepSeek-ai
- Hugging Face Diffusers by Hugging Face
- LMCacheby LMCache Project
- NVIDIA Dynamo Documentationby NVIDIA (2025)
- NVIDIA Nsight Systems by NVIDIA
- NVIDIA Triton Inference Server by NVIDIA
- ONNX Runtime by Microsoft
- “PyTorch Performance Tuning Guide”by Szymon Migacz (PyTorch Foundation, 2020)
- “PyTorch Profiler”by Shivam Raikundalia (PyTorch Foundation, 2021)
- SGLang Projectby LMSYS Org
- “TensorRT Documentation”by NVIDIA (2025)
- TensorRT-LLMby NVIDIA
- Transformersby Hugging Face
- vLLM Project by The Linux Foundation
Frontier Open Models
Model cards and technical reports for the open weights worth serving.
GPU Infrastructure
Hardware documentation, architecture whitepapers, and interconnect.
- Designing Data-Intensive Applicationsby Martin Kleppmann (O’Reilly Media, 2017)
- Grace Hopper / Grace Blackwell Systemsby NVIDIA
- GPU Glossaryby Frye et al. (Modal, 2025)
- InfiniBandby NVIDIA
- Kubernetes Documentationby The Kubernetes Authors (The Linux Foundation, 2025)
- NVIDIA Blackwell Architecture Technical Brief: Built for the Age of AI Reasoningby NVIDIA (2025)
- NVIDIA H100 Tensor Core GPU Architecture: Exceptional Performance, Scalability and Security for the Data Centerby NVIDIA (2023)
- “NVIDIA Tesla: A Unified Graphics and Computing Architecture”by E. Lindholm et al. (IEEE Micro, March–April 2008)
- NVLink / NVSwitchby NVIDIA
- Programming Massively Parallel Processors: A Hands-on Approachby Wen-mei Hwu, David Kirk, Izzat El Hajj (Morgan Kaufmann, 2022)
- SemiAnalysisby Dylan Patel (SemiAnalysis, 2025)
- Site Reliability Engineering: How Google Runs Production Servicesedited by Betsy Beyer et al. (O’Reilly Media, 2017)
Inference Optimization Research
Where the techniques in Chapter 5 come from: FlashAttention, speculation, quantization, paged attention.
- “Adversarial Diffusion Distillation”by Axel Sauer et al. (European Conference on Computer Vision, 2023)
- “Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation”by Ofir Press, Noah Smith, and Mike Lewis (ArXiv abs/2108.12409, 2021)
- “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration”by Song Han (MIT, 2024)
- Cache-DITby Vipshop
- “CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion”by Jiayi Yao et al. (Proceedings of the Twentieth European Conference on Computer Systems, 2024)
- “Adding Conditional Control to Text-to-Image Diffusion Models”by Lvmin Zhang et al. (International Conference on Computer Vision, 2023)
- “Beyond the Buzz: A Pragmatic Take on Inference Disaggregation”by Tiyasa Mitra et al. (ArXiv abs/2506.05508, 2025)
- “Break the Sequential Dependency of LLM Inference Using Lookahead Decoding”by Yichao Fu et al. (ArXiv abs/2402.02057, 2024)
- “EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty”by Yuhui Li et al. (ArXiv abs/2401.15077, 2024)
- “EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees”by Yuhui Li et al. (Conference on Empirical Methods in Natural Language Processing, 2024)
- “EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test”by Yuhui Li et al. (ArXiv abs/2503.01840, 2025)
- “FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving”by Ye et al. (ArXiv abs/2501.01005, 2025)
- “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers”by Elias Frantar (ArXiv abs/2210.17323, 2022)
- “High-Resolution Image Synthesis with Latent Diffusion Models”by Robin Rombach et al. (Conference on Computer Vision and Pattern Recognition (CVPR), 2021)
- “Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference”by Simian Luo (ArXiv abs/2310.04378, 2023)
- “LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale”by Tim Dettmers et al. (ArXiv abs/2208.07339, 2022)
- “Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads”by Tianle Cai et al. (ArXiv abs/2401.10774, 2024)
- “Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”by Mohammad Shoeybi et al. (ArXiv abs/1909.08053, 2019)
- “Efficient Memory Management for Large Language Model Serving with PagedAttention”by Woosuk Kwon et al. (Proceedings of the 29th Symposium on Operating Systems Principles, 2023)
- “Fast Inference from Transformers via Speculative Decoding”by Yaniv Leviathan et al. (International Conference on Machine Learning, 2022)
- “Ring Attention with Blockwise Transformers for Near-Infinite Context”by Hao Liu et al. (ArXiv abs/2310.01889, 2023)
- “SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration”by Jintao Zhang et al. (ArXiv abs/2410.02367, 2024)
- Sequence/Context Parallelismby Megatron-LM for NVIDIA
- SmoothQuantby Song Han (MIT)
- “SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot”by Elias Frantar and Dan Alistarh, (ArXiv abs/2301.00774, 2023)
- “SpecVLM: Fast Speculative Decoding in Vision-Language Models”Haiduo Huang et al. (ArXiv abs/2509.11815, 2025)
- “TeaCache: Timestep Embedding Aware Cache”by Feng Liu et al. (Alibaba TongYi Vision Intelligence Lab, ArXiv abs/2411.19108, 2025)
Intelligence Evaluation
Benchmarks, their construction, and their well-documented limits.
- ARC AGI Prizeby Greg Kamradt (2025)
- Evals for AI Engineers: Systematically Measuring and Improving AI Applicationsby Shreya Shankar and Hamel Husain (O’Reilly Media, forthcoming 2026)
- Grade School Math: Training Verifiers to Solve Math Word Problemsby Karl Cobbe and Vineet Kosaraju (ArXiv abs/2110.14168, 2021)
- “How to Fine-Tune Qwen3 to GPT-4o Level Performance”by Greg Schoeninger (Fine-Tune Fridays, Oxen AI, 2025)
- “Humanity’s Last Exam”by Long Phan et al. (Center for AI Safety and Scale AI, ArXiv abs/2501.14249, 2025)
- “HumanEval: Evaluating Large Language Models Trained on Code”by Mark Chen et al. (OpenAI, 2021)
- “MMLU: Measuring Massive Multitask Language Understanding”by Dan Hendrycks et al. (Proceedings of the International Conference on Learning Representations (ICLR), 2021)
- “MTEB: Massive Text Embedding Benchmark”by Niklas Muennighoff et al. (Conference of the European Chapter of the Association for Computational Linguistics, 2022)
- “SWE-Bench: Can Language Models Resolve Real-World Github Issues?”by Carlos Jimenez et al. (Proceedings of the International Conference on Learning Representations (ICLR), 2024)
Reproduced from Appendix B of Inference Engineering. Links go to the primary sources; anything worth understanding properly is better read there than summarized here.