Tile Accelerator
A compiler and simulator for tiled neural-network workloads
I built a small accelerator instruction set and a compiler that makes its memory transfers and dependencies explicit. It supports a residual MLP and a depthwise/pointwise CNN block exported from PyTorch. An independent C++ engine executes the binary so the generated program can be checked against the source model.
Python · C++ · Binary ISA · IREE VM bridge
From a graph to an executable program
LOAD2D input tile → scratch A
LOAD2D weight tile → scratch B
MATMUL A × B → accumulator C
EPILOGUE bias / residual / ReLU
STORE2D accumulator → output Check the working set
A, B, the accumulator, bias and residual have to fit the target's scratch budget. Tile boundaries and partial reduction tiles are explicit. Unsupported operators and shapes fail compilation with a diagnostic.
Encode and execute
The versioned binary includes commands, constants and an input/output contract. The C++ interpreter checks bounds, initialization, dependencies and finite results. An IREE VM native-module bridge can run the same binary across a host runtime boundary.
Test precision and request state
FP16 operand mode uses nearest/ties-to-even rounding and retains FP32 accumulation, bias and output. Tests cover subnormals, overflow, repeated requests, malformed binaries and source-model agreement. Overflow fails before an activation can hide it.
Study DMA overlap and architecture choices
Double buffering uses a second pair of operand slots. It can prefetch the next reduction tile while the matrix engine works, with dependencies that prevent a buffer from being overwritten before its consumer finishes.
| Schedule | Estimated cycles | Estimated latency |
|---|---|---|
| Serial tile transfers | 20,384 | 101.92 µs |
| Double-buffered transfers | 17,237 | 86.19 µs |
This example uses a 17 × 65 × 79 MLP, 32 KiB scratch, a 200 MHz modeled clock and the declared DMA/compute rates. Both schedules pass numerical checks. The times are analytical estimates, not measurements from a physical accelerator.
The study also varies scratch capacity, matrix throughput, tile width, bandwidth and component energy. Its area quantity is an illustrative proxy. Calibration against a technology library or real hardware would be needed for chip power and area claims.
Recorded study data ↗What the project demonstrates
The software connects a supported model graph to an instruction stream, binary format, runtime boundary and numerical test. The functional interpreter and the cost model remain separate, so a correct answer does not imply an accurate timing prediction.
Evaluation and failure handling ↗