Interactive simulation of the r/chipdesign OpenTransformer 4x4 compute core. Explore how tile reuse enables 16x16 matrix multiplication and why result storage write-controllers determine whether RTL fits on an FPGA.
4×4 Processing Element ArrayTile: 1/16 | Step: 0/7
Reusing 16 MAC processing elements across 16 smaller tiles to compute the full 16×16 matrix product (256 MAC operations per tile).
Gowin GW1NR-9 (Tang Nano 9K) SynthesisFIT_OK
Toggling result write architecture between parallel unrolled registers and 1-per-clock sequential writes over 16 clocks.
LUT Utilization54% (4,676 / 8,640)
0%100% Target Capacity200%
Flip-Flop / Register Usage28% (1,814 / 6,480)
Sequential Write Strategy: Writing 1 result per clock over 16 clocks uses dedicated BRAM/multiplexed registers, lowering LUT logic to 54% and fitting cleanly within Tang Nano 9K limits.
Python ↔ FPGA UART Testbench256/256 MATCH
Simulates sending 16×16 matrices over 115200 baud UART, reading hardware output, and checking golden matrix arithmetic.