Choose your first GPU kernel language by what 2026 hiring demands.

As GPU kernel engineering and LLM inference evolve, aspiring engineers in 2026 face a pivotal choice: dive deep into legacy C++ CuTe and CUTLASS, or embrace the emerging CuTeDSL in Python.

3 min readMachine Learning

The job postings are telling you one thing, and NVIDIA is telling you another. For someone starting fresh in 2026, the honest answer is that the legacy C++ CuTe and CUTLASS template path is no longer the best first investment of your time. The new stack, built around CuTeDSL, Triton, and Mojo, is where the momentum is, and the production evidence is already visible in FlashAttention-4, FlashInfer, and SGLang's collaboration roadmap with NVIDIA. That doesn't mean C++ is dead. It means the deep template metaprogramming skills that used to be a barrier to entry are becoming a maintenance skill, not a creative one.

Here's what that means for you practically. If you want to contribute to the next generation of inference engines, learning CuTeDSL first gets you to the same performance level as CUTLASS without the template gymnastics. You get JIT compilation, faster iteration, and direct TorchInductor integration, which means you can test ideas in hours instead of days. That's not a toy path. That's the path NVIDIA is recommending for new kernels, and the projects that matter are following. Triton then becomes your second language, not because it replaces CuTeDSL, but because it lets you move between research and production without rewriting everything. Mojo, or Rust for serving, rounds out the stack for the systems side, where you need control over memory and concurrency without drowning in C++ boilerplate.

Now, the job postings are not lying. They still list C++17 and CUTLASS as hard requirements because hiring is reactive, not forward-looking. But here's the reality: a candidate who understands kernel design principles, can read C++ to understand legacy code, and can write new kernels in CuTeDSL or Triton is more valuable than someone who can only write template-heavy CUTLASS code. The former can adapt. The latter is locked into a style that NVIDIA is actively moving away from. So keep light C++ for reading and debugging, but do not make it your primary skill. The people who get hired and ship will be the ones who can move fast, and the new stack is built for speed.

If you're mapping out your learning order, start with CuTeDSL to understand the fundamentals of tiling, memory coalescing, and shared memory without the syntax fighting you. Then move to Triton to see how those concepts scale across different hardware targets. Add Mojo or Rust for the serving layer once you can write a working kernel. Pick a concrete project, like a simplified FlashAttention variant, and implement it in each language. That exercise will teach you more than a year of reading CUTLASS templates. The shift is real, and the people who embrace it now will be the ones defining what the next generation of kernel engineering looks like.

From Machine Learning

For people just starting out in GPU kernel engineering or LLM inference (FlashAttention / FlashInfer / SGLang / vLLM style work), most job postings still list “C++17, CuTe, CUTLASS” as hard requirements.

At the same time NVIDIA has been pushing CuTeDSL (the Python DSL in CUTLASS 4.x) hard since late 2025 as the new recommended path for new kernels — same performance, no template metaprogramming, JIT, much faster iteration, and direct TorchInductor integration.

Read the original at Machine Learning