MARATTO

article

Neural Network Quantization for FPGA: A Unified Taxonomy and Survey of Methods, Platforms, and Deployment

Abstract

Neural network quantization is a key enabler for deploying deep learning models on field-programmable gate arrays (FPGAs), as it significantly reduces memory footprint, computational complexity, and energy consumption while preserving competitive inference accuracy. By exploiting customized datapaths and fine-grained parallelism, FPGAs are particularly well suited for low-precision arithmetic. Nevertheless, the rapid expansion of quantization techniques, FPGA platforms, and deployment frameworks has given rise to a complex design space in which selecting suitable solutions under accuracy, performance, and resource constraints remains challenging. This paper introduces a unified taxonomy and presents a focused survey of neural network quantization for FPGA deployment. Existing approaches are organized along key design dimensions, including quantization stage, numerical precision, granularity, hardware awareness, and deployment abstraction level. Primary techniques, such as post-training quantization and quantizationaware training, are reviewed alongside advanced strategies encompassing mixed-precision, extreme low-bit, and hardwareaware quantization. In addition, representative FPGA platforms and deployment frameworks are analyzed, and major application domains benefiting from quantized FPGA-based inference are summarized. Finally, open challenges are discussed and future research directions toward more automated, portable, and efficient FPGA-based deep learning systems are outlined.

Research topics

  • Embedded Systems Design Techniques
  • Advanced Neural Network Applications
  • VLSI and FPGA Design Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iraset68627.2026.11538436

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.