← All projects

ML Compiler Lab

GPU kernels and model compilation

I built CuTe kernels for a two-layer residual MLP and compare persistent execution with conventional kernels and CUDA Graph replay. The tensor-core megakernel performs both projections in one launch and keeps the hidden state in shared memory.

CuTe DSL · CUDA Graphs · PyTorch · IREE / MLIR · C++

The graph and its execution

H = ReLU(X @ W1 + b1)
Y = ReLU(H @ W2 + b2 + X)

Own a complete row tile

Each thread block computes sixteen independent rows. It stages operands, runs the first projection, stores the rounded hidden state in shared memory, then runs the second projection. Block barriers protect each producer/consumer phase; there is no global spinning barrier.

Check the graph before selecting a schedule

A separate FP32 path exports the PyTorch graph and verifies its operators, shapes and row dependencies. An original C++ pass in IREE lowers that contract into a persistent or conventional schedule. Cross-row reductions are rejected by the row-owned frontend.

Inspect the compiler output

The CNN baseline retains FX, MLIR, LLVM IR and PTX. CuTe kernels retain generated GPU code for inspection. Attention and an explicitly decomposed recurrent cell also run through the stock compiler.

Measured comparisons

The assessment shapes below were separate from the initial development cases. The comparisons include a megakernel that stores the intermediate in global memory, which helps distinguish memory reuse from keeping the same operator schedule in one launch.

Rows 513, channels 96, hidden width 192. Input, weights and hidden state are FP16; accumulation, bias and output are FP32.

Execution Median CUDA-event interval Median host completion
Conventional CuTe + graph replay 26.62 µs 34.73 µs
Shared-memory megakernel + graph replay 35.78 µs 43.86 µs
Global-memory intermediate + graph replay 41.98 µs 50.50 µs
cuBLAS/PyTorch + graph replay 21.50 µs 30.32 µs

Recorded on one RTX 3090. Device-resident input and output; 1,000 observations per variant across five rounds in one process. CUDA-event intervals can include dispatch idle gaps. These are small synthetic workloads, not wearable-device or whole-LLM latency.

Recorded assessment comparisons for conventional, persistent and cuBLAS graph replay

The smaller graphs can benefit from persistence. Larger shapes can favor vendor kernels or a different tile schedule. The implementation retains those losses and supplies a conventional path rather than assuming one kernel is best for every shape.

Recorded aggregate data ↗

Failures the checks caught

An early tile-tuning script captured an empty CUDA Graph on the wrong stream. Another version left partial-channel MMA extents unpadded and produced nonfinite results. Those attempts were rejected. Replay now has to overwrite a poisoned output, and the fixed kernels are checked for memory, shared-memory race and synchronization errors.

These tests establish arithmetic and execution behavior with original random weights. The supplied quality evaluator can compare task accuracy and data slices when a user provides a model and a licensed, representative evaluation set. This study makes no perception-accuracy or demographic-fairness claim.