LLM Inference Accelerator

Jan. 2025 - Apr. 2025 · Waterloo, ON

  • SystemVerilog
  • FPGA
  • Vivado
  • RTL Design

Project Overview

Implemented the IBERT LLM model for inference on a PYNQ FPGA, with a focus on hardware compute units and throughput.

  • Wrote and optimized RTL implementations of systolic arrays, matrix multipliers, softmax, and layer normalization for synthesis on FPGA.
  • Developed a suite of throughput-optimized, pipelined compute units in SystemVerilog for the PYNQ FPGA.