Experimental Machine Learning Hardware Accelerators in Scala (SpinalHDL)
PyTorch-like developer ergonomics • Automated FPGA compilation • Validated physical FPGA inference
SpinalML is an experimental hardware Machine Learning acceleration library designed for FPGA synthesis and simulation, written in Scala using SpinalHDL. It bridges the gap between high-level deep learning concepts and silicon RTL, enabling developers and researchers to describe neural networks with a PyTorch-like API while generating functional, pipelined, cycle-accurate hardware.
Note
Student Research & Experimental Project: SpinalML is an exploratory student project under active development. While end-to-end inference has been physically verified on hardware, primitives, quantization pipelines, and toolchain features are continuously evolving and should not be considered production-grade.
FPGA-based AI acceleration is typically confined to expensive enterprise platforms or locked behind opaque vendor tools. SpinalML was born out of an experimental vision: exploring neural network execution (with the long-term ambition of evaluating quantized Small Language Models) on resource-constrained edge FPGAs.
- Transparent RTL: Unlike traditional C-based HLS tools that can produce opaque Verilog, SpinalML generates deterministic, cycle-accurate RTL with full architectural visibility through SpinalHDL.
- Resource-Conscious Design: Custom mixed-precision (INT4, W4A8, FP8), temporal resource sharing, and double-buffered DMA streaming designed to fit models into devices with limited logic (down to 20K LUTs).
- PyTorch-like Ergonomics with Hardware Control: Define networks with a clean, declarative API (
Sequential,Conv2D,Linear,Attention) while retaining direct visibility over physical registers, FIFOs, and DSP slices.
flowchart LR
subgraph Build ["spinalml build"]
direction LR
A["Scala Model"] --> B["SpinalHDL\n(Verilog RTL)"]
B --> C["Yosys\n(Synthesis)"]
C --> D["nextpnr\n(Place & Route)"]
D --> E["Bitstream\n(top.fs)"]
end
subgraph Flash ["spinalml flash"]
direction LR
E --> F["openFPGALoader"]
F --> G["FPGA Silicon\n(Tang Primer 20K)"]
end
Any neural network model defined with SpinalML can be simulated and verified without writing boilerplate testbenches:
# Automatically scaffold, simulate, and verify any model under Verilator
spinalml test examples/Mnist/Model.scala- Automated SoC Simulation: Scaffolds an AXI4 memory harness and AXI-Lite control plane under Verilator 5.
- Bit-Exact Software Oracles: Validates hardware outputs against
ModelReplicawith zero numerical deviation (deviation = 0.000). - Formal Verification: Streaming flows, handshakes, and operations across the codebase are formally proven against protocol violations and deadlocks via SymbiYosys and SMT solvers (
cvc4) (spinalml test-all-formal).
SpinalML provides standalone, single-file executables for Windows, Linux (x64 & ARM64), and macOS with zero Python dependencies required.
Note
Disk Space Requirement:
Please ensure you have 2 to 3 GB of free disk space. Running spinalml setup automatically provisions the full open-source FPGA toolchain (OSS CAD Suite, Mill, Verilator, Yosys, nextpnr, openFPGALoader) into ~/.spinalml_tools.
-
Download the precompiled binary for your operating system from the latest GitHub Releases.
-
Add the binary to your PATH to invoke
spinalmlfrom any directory:- Linux / macOS:
chmod +x spinalml # Move to system PATH (e.g. /usr/local/bin) or export the folder: sudo mv spinalml /usr/local/bin/spinalml # or: export PATH="$PATH:/path/to/binary/folder"
- Windows (PowerShell):
# Add binary location to current session PATH (or add permanently in System Settings): $env:Path += ";$PWD"
- Linux / macOS:
-
Initialize the toolchain:
spinalml setup
-
Verify, compile & simulate a sample model directly (no git clone required): Download the self-contained sample component
Universal1DDemo.scala:# Download standalone sample model curl -O https://raw.githubusercontent.com/Juste-Leo2/spinalML/main/tests/universal/Universal1DDemo.scala # Run bit-exact cycle-accurate C++ simulation with Verilator spinalml test Universal1DDemo.scala # Turnkey FPGA build for Sipeed Tang Primer 20K (Gowin GW2A-18) spinalml build Universal1DDemo.scala --board tang-primer-20k # Flash bitstream to SRAM via openFPGALoader spinalml flash --board tang-primer-20k
If you are developing SpinalML or running Python Cocotb co-simulations, clone the repository and follow:
Guide: Setting up SpinalML from Source & Running All Tests
(Includes the essential Python 3.12 global interpreter requirement for Cocotb VPI on Ubuntu / WSL 2, and uv bootstrap instructions).
To see what you can build and run on physical silicon with SpinalML, check out the examples/ directory:
- MNIST Hardware Accelerator on Tang Primer 20K: A complete end-to-end demonstration featuring an ultra-compact mixed-precision (W4A8) CNN accelerator running in real time on the Gowin GW2A-18, paired with an interactive Gradio web drawing canvas and live UART telemetry.
SpinalML abstracts away tedious hardware handshakes. You can construct neural networks with high-level declarative components:
Tensor[T]: Multi-dimensional tensor abstraction embedding physical hardwareStreaminterfaces (valid/readyhandshakes).- Automated Pipelining: Layers handle FIFO buffering, dimensional broadcasting, and backpressure automatically.
- Sequential API: Define your model cleanly using a PyTorch-like layer list.
import spinal.core._
import spinal.lib.bus.amba4.axi.Axi4Config
import spinalML.nn._
import spinalML.dtypes._
case class TinyMLP(
override val axiConfig: Axi4Config = Axi4Config(addressWidth = 32, dataWidth = 64, idWidth = 4)
) extends Accelerator(
dataType = I8(), // 8-bit integer quantization
inputShape = Seq(1, 16), // 1D input vector of length 16 (1x16)
modelSpec = Seq(
Linear(inFeatures = 16, outFeatures = 32),
ReLU(),
Linear(inFeatures = 32, outFeatures = 4)
),
axiConfig = axiConfig
)No boilerplate App object or manual Verilog runner needed: the SpinalML CLI automatically elaborates your class, wraps it in the SoC bus, and compiles it directly:
python cli/main.py build TinyMLP.scala --board tang-primer-20kFor advanced experiments and architectural exploration, SpinalML provides modular hardware building blocks and quantization experiments:
-
Transformers & Attention Mechanisms:
-
ClassicalAttention(embedDim, numHeads): Multi-Head Attention blocks with batched streaming matrix multiplications (matmul) and hardware-accelerated Softmax.
-
-
Mixed Precision & Weight-Only Quantization (wXaY):
- Run activations in floating-point (
FP8,BF16) with compact integer weights (I4,I8) usingcustomWeightTypeand compile-timeweightScales. - On-the-fly precision adaptation via
Requantize(shift, targetType)andCast(targetType).
- Run activations in floating-point (
-
2D Computer Vision:
-
Conv2D,MaxPool2D, andAvgPool2Dbacked by BRAM-based line buffers (Mem+readSync) with zero CPU overhead. - Normalization layers:
LayerNormand inference-foldedBatchNorm.
-
-
SoC & Memory Infrastructure:
- Autonomous AXI4 DMA memory engines (
DMAReader,DMAReader2D) withStreamDoubleBufferto overlap memory transfers with compute. - Hardware CSR bridge with UART packet framing for direct host-to-FPGA streaming inference.
- Autonomous AXI4 DMA memory engines (
-
Hardware LUT & Area Optimization Knobs:
-
lanes&Repack(newLanes): Bus streaming width (lanes) is automatically inferred per layer, but can be manually adapted viaRepackto scale parallelism up or down to save routing LUTs. -
temporalAccumulator Streaming: Matrix-multiplication layers accumulate full$M \times N$ output tables in registers by default (temporal = 0). Settingtemporal > 0inAccelerator(..., temporal = N)drains completed rows on-the-fly, reclaiming thousands of LUTs and flip-flops on area-constrained FPGAs.
-
Explore detailed guides in the docs/ directory:
- Getting Started Guide: Core concepts, tensors, streams, and your first hardware layers.
- High-Level Tutorial: Complete guide to the PyTorch-like API, quantization, and DAG topologies.
- Operations API Reference: Detailed specifications of all supported hardware operations and layers.
- CLI Reference: Commands for compilation, simulation, synthesis, formal verification, and flashing.
- Project Structure & Architecture: Repository layout, 3-layer hardware sandwich architecture, and verification framework.
- Project Roadmap: High-level architectural roadmap, ONNX ingestion, and edge SLM goals.
- Full Technical Roadmap: Complete phase-by-phase implementation plan and detailed checklist.
- FPGA Board Roadmap: Hardware compatibility matrix and community board support guidelines.
- UART Bridge & Protocol: Specifications for physical UART host communication and CSR bridges.
- Application Examples: Silicon-validated hardware projects, including the Tang Primer 20K MNIST Accelerator.
- Supported FPGA Boards: Hardware compatibility status.
Physical in-circuit synthesis, place-and-route, bitstream generation, and real-time UART inference are validated on hardware for:
| Vendor | Board | FPGA Device | Toolchain / Programmer | Target Slug | Hardware Status |
|---|---|---|---|---|---|
| Gowin | Sipeed Tang Primer 20K | GW2A-LV18PG256C8/I7 | Yosys + nextpnr-himbaechel / openFPGALoader | tang-primer-20k |
✅ Hardware Validated |
Note
SpinalML generates generic, vendor-agnostic RTL. For the complete matrix of planned FPGA targets (including AMD/Xilinx Artix-7/Zynq, Lattice ECP5/iCE40, and Intel Cyclone V) or to contribute a board definition, see the FPGA Board Roadmap.
Contributions are welcome! Whether you are adding support for a new FPGA board target, implementing hardware neural network layers, or improving toolchain automation, check out CONTRIBUTING.md to get started.
- SpinalHDL: For providing an unmatched, expressive hardware description language that makes digital design concise, modular, and robust.
- The Yosys & Open-Source FPGA Community: For building phenomenal open-source synthesis, PnR, and bitstream tools (
yosys,nextpnr,openFPGALoader).
SpinalML is an individual student project. Designing a full-stack hardware ML acceleration library from scratch—spanning Scala/SpinalHDL RTL primitives, custom DMA streaming engines, automated CLI build pipelines, verification testbenches, and physical board deployment—would have been practically impossible for a single student to accomplish without the leverage of modern Generative AI.
- AI Code Generation: The vast majority of the code across this repository was generated with the assistance of LLMs:
- Gemini Flash & Pro (via Antigravity)
- DeepSeek Flash & Pro (via OpenCode)
- GLM 5.3 Flash (via OpenCode)
- Human Role & Quality Oversight: The author's role focused on system architecture, engineering decisions, and quality control. Countless hours were dedicated to testing, debugging hardware timing and synthesis anomalies, designing verification testbenches, validating physical silicon inference over UART, and continuously reviewing generated code to ensure architectural integrity, rigor, and clean design patterns throughout the project.
If you use SpinalML in your research or hardware projects, please cite:
@software{adamo2026spinalml,
author = {Adamo, Léonard},
title = {{spinalML: Hardware Machine Learning Accelerators with SpinalHDL}},
year = {2026},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/Juste-Leo2/spinalML}},
note = {Student project, Université de Montpellier, France}
}This project is licensed under the MIT License - see the LICENSE file for details.
Copyright (c) 2026 Léonard Adamo.