ML Compiler Lab
GPU kernels and model compilation
I built CuTe kernels for a two-layer residual MLP and compare persistent execution with conventional kernels and CUDA Graph replay. The tensor-core megakernel performs both projections in one launch and keeps the hidden state in shared memory.
CuTe DSL · CUDA Graphs · PyTorch · IREE / MLIR · C++
The graph and its execution
H = ReLU(X @ W1 + b1)
Y = ReLU(H @ W2 + b2 + X) Own a complete row tile
Each thread block computes sixteen independent rows. It stages operands, runs the first projection, stores the rounded hidden state in shared memory, then runs the second projection. Block barriers protect each producer/consumer phase; there is no global spinning barrier.
Check the graph before selecting a schedule
A separate FP32 path exports the PyTorch graph and verifies its operators, shapes and row dependencies. An original C++ pass in IREE lowers that contract into a persistent or conventional schedule. Cross-row reductions are rejected by the row-owned frontend.
Inspect the compiler output
The CNN baseline retains FX, MLIR, LLVM IR and PTX. CuTe kernels retain generated GPU code for inspection. Attention and an explicitly decomposed recurrent cell also run through the stock compiler.
Measured comparisons
The assessment shapes below were separate from the initial development cases. The comparisons include a megakernel that stores the intermediate in global memory, which helps distinguish memory reuse from keeping the same operator schedule in one launch.
Rows 49, channels 48, hidden width 96. Input, weights and hidden state are FP16; accumulation, bias and output are FP32.
| Execution | Median CUDA-event interval | Median host completion |
|---|---|---|
| Conventional CuTe + graph replay | 15.36 µs | 23.79 µs |
| Shared-memory megakernel + graph replay | 12.10 µs | 20.23 µs |
| Global-memory intermediate + graph replay | 14.34 µs | 22.55 µs |
| cuBLAS/PyTorch + graph replay | 18.43 µs | 27.17 µs |
Recorded on one RTX 3090. Device-resident input and output; 1,000 observations per variant across five rounds in one process. CUDA-event intervals can include dispatch idle gaps. These are small synthetic workloads, not wearable-device or whole-LLM latency.
The smaller graphs can benefit from persistence. Larger shapes can favor vendor kernels or a different tile schedule. The implementation retains those losses and supplies a conventional path rather than assuming one kernel is best for every shape.
Recorded aggregate data ↗Failures the checks caught
An early tile-tuning script captured an empty CUDA Graph on the wrong stream. Another version left partial-channel MMA extents unpadded and produced nonfinite results. Those attempts were rejected. Replay now has to overwrite a poisoned output, and the fixed kernels are checked for memory, shared-memory race and synchronization errors.
These tests establish arithmetic and execution behavior with original random weights. The supplied quality evaluator can compare task accuracy and data slices when a user provides a model and a licensed, representative evaluation set. This study makes no perception-accuracy or demographic-fairness claim.