AI Systems Performance Engineering
A 1,062-page O’Reilly field guide and active open source lab for GPU architecture, CUDA, PyTorch, distributed training, and modern inference serving.
1.8K+ stars · 250+ forks · 2,700+ authored commits
Hi, I’m
AI systems performance engineer · Product leader · Founder · Advisor · 3× O’Reilly author
I build and explain high-performance AI systems. My work spans GPU kernels, distributed training, high-throughput inference, and the product systems that move AI into production.
Building open source tools for GPU performance and agent evaluation, while advising teams that are taking AI systems into production.
Recent work
The current work connects low-level performance engineering with the product discipline needed to ship useful AI.
A 1,062-page O’Reilly field guide and active open source lab for GPU architecture, CUDA, PyTorch, distributed training, and modern inference serving.
1.8K+ stars · 250+ forks · 2,700+ authored commits
An evidence-first workbench for evaluating agent prompts, tools, transcripts, and traces without losing the receipts behind each conclusion.
Prompts · Traces · Evals · Evidence
A set of agent workflows for profiling and improving GPU inference, including benchmarking, quantization, speculative decoding, and performance reports.
32 focused workflows
Open source impact
I publish the code, benchmark results, and review trail behind the work. These are three recent examples, including two upstream contributions.
Flattened the MLA KV-cache gather and added four-way instruction-level parallelism. Mirage merged the patch with bit-identical output.
Read the merged PRPublished agent workflows and an MCP server for profiling, benchmarking, quantization, speculative decoding, and performance reporting.
Explore the toolkitProposed a broadcast-bias baddbmm decomposition that also measured 1.31× faster at the surrounding mixture-of-experts layer.
Review the PyTorch PRThe flagship AI Systems Performance Engineering repository has earned 1.8K+ stars and 250+ forks, with 2,700+ commits authored across the book and lab.
See all public work on GitHubFeatured appearances
Recent podcast conversations about GPU performance, software and hardware codesign, and how agentic coding is changing engineering.
SuperDataScience · Episode 973
Jon Krohn and I go from memory bandwidth and GPU profiling to agentic coding and the systems judgment AI engineers still need.
49K+ YouTube views
MLOps Community · Episode 363
A technical conversation about PyTorch, CUDA, GPU architecture, mechanical sympathy, and the cost of production inference.
85-minute technical deep dive
Books
From production data science to generative AI and the performance of the full system beneath it.
Optimize training and inference across GPU hardware, CUDA kernels, compilers, PyTorch, networking, storage, and multinode systems.
Build context-aware multimodal applications with foundation models, fine-tuning, reinforcement learning, RAG, and production deployment patterns.
Implement end-to-end machine learning pipelines with data engineering, model training, tuning, and production deployment on AWS.
Course
A practical course built with DeepLearning.AI and AWS for people who want to understand the full generative AI lifecycle.
The course covers transformers, model selection, scaling laws, fine-tuning, evaluation, reinforcement learning, inference, and deployment. It includes 47 lessons and three graded assignments.
More talks
Selected talks and community sessions from the last two years.
AI Performance Engineering · Rob Ferguson
I hosted a practical session on getting more technical and business value from AI startup accelerator programs.
Watch or listenAI Performance Engineering
A shared recap of inference engines, disaggregated prefill and decode, and the systems work behind faster models.
Watch or listenAI Performance Engineering
Practical techniques for adapting inference systems as request patterns, memory pressure, and latency targets change.
Watch or listenO’Reilly AI Superstream
Performance patterns for agentic workloads using DeepSeek, NVIDIA Dynamo, vLLM, CUDA, and PyTorch.
Watch or listenCommunity
I cohost monthly technical sessions with Antje Barth for engineers who build the systems behind modern AI. Our global network reaches more than 100,000 people worldwide. The flagship Meetup group has hosted 375 past events, and the YouTube channel has 193 videos with more than 780,000 views.
Speaking, workshops, and advisory
I’m open to selected podcast conversations, technical talks, performance workshops, and work with founders building AI infrastructure.