Design of Quantized Deep Neural Network Hardware Inference Accelerator Using Systolic Architecture

Dary Mochamad Rifqie, Yasser Abd. Djawad, Faizal Arya Samman, Ansari Saleh Ahmar, M. Miftach Fakhri
https://doi.org/10.35877/454RI.asci2689

Abstract

This paper presents a hardware inference accelerator architecture of quantized deep neural networks (DNN). The proposed accelerator implements all computation in a quantize version of DNN including linear transformations like matrix multiplications, nonlinear activation functions such as ReLU, quantization and dequantization operation. The hardware accelerator of quantized DNN consists of matrix multiplication core which is implemented in systolic array architecture, and the QDR core for computing the operation of quantization, dequantization, and ReLU. This proposed hardware architecture is implemented in Verilog Hardware Description Language (HDL) code using modelsim. To validate, we simulated the quantized DNN using Python programming language and compared the results with our proposed hardware accelerator. The result of this comparison shows a very slight difference, confirming the validity of our quantized DNN hardware accelerator.

Keywords

Downloads

Download data is not yet available.

References (18)

  1. Adiono, T., Meliolla, G., Setiawan, E., & Harimurti, S. (2019). Design of Neural Network Architecture using Systolic Array Implemented in Verilog Code. ISESD 2018 - International Symposium on Electronics and Smart Devices: Smart Devices for Big Data Analytic and Machine Learning, 126, 1–4. https://doi.org/10.1109/ISESD.2018.8605478
  2. Adiono, T., Putra, A., Sutisna, N., Syafalni, I., & Mulyawan, R. (2021). Low Latency YOLOv3-Tiny Accelerator for Low-Cost FPGA Using General Matrix Multiplication Principle. IEEE Access, 9, 141890–141913. https://doi.org/10.1109/ACCESS.2021.3120629
  3. Amin, M., & Adiono, T. (2019). Area Optimized CNN Architecture Using Folding Approach. 2019 International Conference on Electrical Engineering and Informatics (ICEEI), 206–209. https://doi.org/10.1109/ICEEI47359.2019.8988879
  4. Baba, A. (2024). Neural networks from biological to artificial and vice versa. Biosystems, 235, 105110. https://doi.org/https://doi.org/10.1016/j.biosystems.2023.105110
  5. Bishop, C. M., & Bishop, H. (2024). Foundations and Concepts Deep Learning.
  6. Courbariaux, M., David, J. P., & Bengio, Y. (2015). Training deep neural networks with low precision multiplications. 3rd International Conference on Learning Representations, ICLR 2015 - Workshop Track Proceedings, Section 5, 1–10.
  7. El Omary, S., Lahrache, S., & El Ouazzani, R. (2024). Attention mechanism-based model for cardiomegaly recognition in chest X-Ray images. IAES International Journal of Artificial Intelligence (IJ-AI), 13(1), 1005. https://doi.org/10.11591/ijai.v13.i1.pp1005-1013
  8. Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., & Keutzer, K. (2021). A Survey of Quantization Methods for Efficient Neural Network Inference. CoRR, abs/2103.1. https://arxiv.org/abs/2103.13630
  9. Holly, S., Wendt, A., & Lechner, M. (2020). Profiling Energy Consumption of Deep Neural Networks on NVIDIA Jetson Nano. 2020 11th International Green and Sustainable Computing Workshops, IGSC 2020. https://doi.org/10.1109/IGSC51522.2020.9290876
  10. Jones, D. H., Powell, A., Bouganis, C. S., & Cheung, P. Y. K. (2010). GPU versus FPGA for high productivity computing. Proceedings - 2010 International Conference on Field Programmable Logic and Applications, FPL 2010, March 2014, 119–124. https://doi.org/10.1109/FPL.2010.32

How to Cite

Rifqie, D. M., Djawad, Y. A., Samman, F. A., Ahmar, A. S., & Fakhri, M. M. (2024). Design of Quantized Deep Neural Network Hardware Inference Accelerator Using Systolic Architecture. Journal of Applied Science, Engineering, Technology, and Education, 6(1), 27–33. https://doi.org/10.35877/454RI.asci2689