Hi, I’m

Chris Fregly

AI systems performance engineer · Product leader · Founder · Advisor · 3× O’Reilly author

I build and explain high-performance AI systems. My work spans GPU kernels, distributed training, high-throughput inference, and the product systems that move AI into production.

Now

Building open source tools for GPU performance and agent evaluation, while advising teams that are taking AI systems into production.

Chris Fregly
1,062pages in the 2025 book
1.8K+GitHub stars
440K+course enrollments
100K+community members worldwide
780K+YouTube views

Recent work

Systems depth, packaged for builders

The current work connects low-level performance engineering with the product discipline needed to ship useful AI.

Open source · 2026

Agent Harness Optimization

An evidence-first workbench for evaluating agent prompts, tools, transcripts, and traces without losing the receipts behind each conclusion.

Prompts · Traces · Evals · Evidence

Open source · 2026

GPU Performance Tuning

A set of agent workflows for profiling and improving GPU inference, including benchmarking, quantization, speculative decoding, and performance reports.

32 focused workflows

Open source impact

Measured work, with the review trail attached

I publish the code, benchmark results, and review trail behind the work. These are three recent examples, including two upstream contributions.

Merged upstream · Mirage · June 2026

2.3× faster

Cut a GB300 KV-cache gather from 21.3 μs to 9.2 μs

Flattened the MLA KV-cache gather and added four-way instruction-level parallelism. Mirage merged the patch with bit-identical output.

Read the merged PR

Open source release · June 2026

32 workflows

Made GPU performance tuning repeatable

Published agent workflows and an MCP server for profiling, benchmarking, quantization, speculative decoding, and performance reporting.

Explore the toolkit

Open PyTorch proposal · June 2026

1.45× faster

Benchmarked a faster Inductor path on GB300

Proposed a broadcast-bias baddbmm decomposition that also measured 1.31× faster at the surrounding mixture-of-experts layer.

Review the PyTorch PR

The flagship AI Systems Performance Engineering repository has earned 1.8K+ stars and 250+ forks, with 2,700+ commits authored across the book and lab.

See all public work on GitHub

Featured appearances

Conversations on the systems beneath AI

Recent podcast conversations about GPU performance, software and hardware codesign, and how agentic coding is changing engineering.

Chris Fregly speaking with Demetrios Brinkmann on the MLOps Community podcastVideo still · MLOps Community

Podcast · February 24, 2026

MLOps Community · Episode 363

Software and Hardware Codesign

A technical conversation about PyTorch, CUDA, GPU architecture, mechanical sympathy, and the cost of production inference.

85-minute technical deep dive

Books

Three O’Reilly books across the AI stack

From production data science to generative AI and the performance of the full system beneath it.

Cover of Generative AI on AWS

O’Reilly · 2023 · New translations in 2024 and 2025

Generative AI on AWS

Build context-aware multimodal applications with foundation models, fine-tuning, reinforcement learning, RAG, and production deployment patterns.

Cover of Data Science on AWS

O’Reilly · 2021

Data Science on AWS

Implement end-to-end machine learning pipelines with data engineering, model training, tuning, and production deployment on AWS.

Course

Generative AI with Large Language Models

A practical course built with DeepLearning.AI and AWS for people who want to understand the full generative AI lifecycle.

Co-instructor · Intermediate

Learn the full generative AI lifecycle

The course covers transformers, model selection, scaling laws, fine-tuning, evaluation, reinforcement learning, inference, and deployment. It includes 47 lessons and three graded assignments.

  • 440K+enrollments
  • 4.8learner rating
  • 24languages

More talks

Technical sessions from kernels to clusters

Selected talks and community sessions from the last two years.

Hosted session · June 2026

AI Performance Engineering · Rob Ferguson

Hacking AI Accelerators

I hosted a practical session on getting more technical and business value from AI startup accelerator programs.

Watch or listen

Community recap · March 2026

AI Performance Engineering

Performance Highlights from GTC 2026

A shared recap of inference engines, disaggregated prefill and decode, and the systems work behind faster models.

Watch or listen

Talk · July 2025

AI Performance Engineering

Dynamic Inference Tuning with CUDA and vLLM

Practical techniques for adapting inference systems as request patterns, memory pressure, and latency targets change.

Watch or listen

Talk · June 2025

O’Reilly AI Superstream

High-Performance Agentic AI Inference Systems

Performance patterns for agentic workloads using DeepSeek, NVIDIA Dynamo, vLLM, CUDA, and PyTorch.

Watch or listen

Community

AI performance is a team sport

I cohost monthly technical sessions with Antje Barth for engineers who build the systems behind modern AI. Our global network reaches more than 100,000 people worldwide. The flagship Meetup group has hosted 375 past events, and the YouTube channel has 193 videos with more than 780,000 views.

100K+worldwide members
375past events
193videos
780K+views

Speaking, workshops, and advisory

Bring me the hard AI systems questions

I’m open to selected podcast conversations, technical talks, performance workshops, and work with founders building AI infrastructure.