← All projects

Tile Accelerator

A compiler and simulator for tiled neural-network workloads

I built a small accelerator instruction set and a compiler that makes its memory transfers and dependencies explicit. It supports a residual MLP and a depthwise/pointwise CNN block exported from PyTorch. An independent C++ engine executes the binary so the generated program can be checked against the source model.

Python · C++ · Binary ISA · IREE VM bridge

From a graph to an executable program

LOAD2D  input tile → scratch A
LOAD2D  weight tile → scratch B
MATMUL  A × B → accumulator C
EPILOGUE bias / residual / ReLU
STORE2D accumulator → output

Check the working set

A, B, the accumulator, bias and residual have to fit the target's scratch budget. Tile boundaries and partial reduction tiles are explicit. Unsupported operators and shapes fail compilation with a diagnostic.

Encode and execute

The versioned binary includes commands, constants and an input/output contract. The C++ interpreter checks bounds, initialization, dependencies and finite results. An IREE VM native-module bridge can run the same binary across a host runtime boundary.

Test precision and request state

FP16 operand mode uses nearest/ties-to-even rounding and retains FP32 accumulation, bias and output. Tests cover subnormals, overflow, repeated requests, malformed binaries and source-model agreement. Overflow fails before an activation can hide it.

Study DMA overlap and architecture choices

Double buffering uses a second pair of operand slots. It can prefetch the next reduction tile while the matrix engine works, with dependencies that prevent a buffer from being overwritten before its consumer finishes.

Schedule Estimated cycles Estimated latency
Serial tile transfers 20,384 101.92 µs
Double-buffered transfers 17,237 86.19 µs

This example uses a 17 × 65 × 79 MLP, 32 KiB scratch, a 200 MHz modeled clock and the declared DMA/compute rates. Both schedules pass numerical checks. The times are analytical estimates, not measurements from a physical accelerator.

Architecture-model latency and energy estimates with explicit scratch and compute assumptions

The study also varies scratch capacity, matrix throughput, tile width, bandwidth and component energy. Its area quantity is an illustrative proxy. Calibration against a technology library or real hardware would be needed for chip power and area claims.

Recorded study data ↗

What the project demonstrates

The software connects a supported model graph to an instruction stream, binary format, runtime boundary and numerical test. The functional interpreter and the cost model remain separate, so a correct answer does not imply an accurate timing prediction.

Evaluation and failure handling ↗