Nvidia has AI agents build faster CUDA kernels
September 29, 2026

Nvidia's open KDA project has coding agents research, write, test, and optimize CUDA kernels. Published results show large gains, but only on tightly bounded workloads.
What this is about
Nvidia Research has released Kernel Design Agents, or KDA, an open workflow in which coding agents develop specialized CUDA kernels. These small GPU programs are the engine room of many AI systems: they help determine how quickly models train, generate images, or produce answers.
The project is not a general coding assistant. A task needs a measurable definition, representative inputs, a correct reference, and explicit target hardware. The agent studies existing approaches, generates candidates, checks correctness, and measures speed. Only a validated candidate is promoted. Code and documentation are available under open licenses.
What Kernel Design Agents actually does
KDA treats optimization as a repeatable experiment. First, a contract defines the computation, the input shapes it must handle, and the command that validates it. The agent then designs variants, compiles them, compares outputs with the reference, and analyzes GPU profiles.
The public wishlist currently supports only Nvidia B200 and B300 GPUs. Developers can submit reproducible tasks as pull requests and others can vote on them. Selected results listed by the team include 1.39 times the performance of a human state of the art in an MLSys contest, a 6.1-times speedup for a balanced k-means kernel, and improvements merged into SGLang, TensorRT-LLM, or FlashInfer.
These figures are not universal model speedups. Some concern a single kernel, while others cover a weighted group of operations. A faster kernel only makes the whole model substantially faster when that operation accounts for a meaningful share of total runtime.
Why it matters
Strong CUDA optimization requires knowledge of memory access, thread layout, numerical formats, and specific GPU architectures. Specialists are scarce. A verifiable agent process can try more variants while documenting results so that a human can understand and reproduce them.
This matters especially for developers of open models. A model can be freely available yet expensive to run because important operations are poorly adapted to new hardware. When optimized kernels are merged upstream into common runtimes, many projects benefit without employing their own GPU specialist. The human role also shifts: less manual writing of every variant, more precise tests, boundaries, and release criteria.
In plain language
A CUDA kernel is like a precise work plan for a large kitchen. The recipe stays the same, but the order changes: who chops, which stove is used, and how unnecessary walking is avoided. KDA lets an agent try many work plans. The dish can only be served after it passes both the taste test and the stopwatch.
A practical example
A team runs an image model that spends 40 milliseconds per request in one matrix operation. It gives the agent 20 common input shapes, a verified reference implementation, and an error tolerance. KDA produces twelve candidates; seven are correct and three are faster. The best reduces the operation to 10 milliseconds.
If the full model previously took 200 milliseconds, it does not automatically fall to 50 milliseconds. Only 30 milliseconds are saved, reducing total time in this example to 170 milliseconds. An end-to-end test is still needed to determine whether compilation overhead, memory use, and rare inputs preserve the gain.
Scope and limits
- The public wishlist currently supports only B200 and B300, and results do not transfer automatically to other GPUs.
- Benchmark gains apply to defined shapes and workloads. Untested inputs can be slower or incorrect.
- Agent-generated low-level code still needs reference tests, numerical checks, profiling, and human approval.
KDA therefore does not show the replacement of GPU engineers. It is a structured amplifier. Its most important feature is not code generation alone but the combination of a reproducible task, automatic measurement, and disclosed limits.
SEO & GEO keywords
Nvidia, Kernel Design Agents, KDA, CUDA, GPU optimization, B200, B300, coding agents, FlashInfer, SGLang, TensorRT-LLM, open source
π‘ In plain English
KDA lets coding agents optimize small but important GPU programs. Every proposal must match a reference result and is measured on real hardware. Gains can be large, but they apply only to the tested tasks and GPUs.
Key Takeaways
- βKDA automates research, implementation, validation, and profiling of CUDA kernels.
- βThe public wishlist currently supports Nvidia B200 and B300.
- βNvidia publishes code, benchmarks, and reproduction steps for selected results.
- βA large kernel gain does not automatically become an equally large end-to-end gain.
- βHuman approval and strong reference tests remain necessary.
FAQ
What is a CUDA kernel?
A CUDA kernel is a small program that performs a clearly defined computation in parallel on an Nvidia GPU.
Can anyone submit a task?
Yes, through the public wishlist. A task needs reproducible inputs, a reference, and a supported target GPU.
Does KDA replace GPU specialists?
No. People must define tasks, tests, and promotion criteria and review the results.