Loading...

FullPack: full vector utilization for sub-byte quantized matrix-vector multiplication on general purpose CPUs

Katebi, H ; Sharif University of Technology | 2024

28 Viewed
  1. Type of Document: Article
  2. DOI: 10.1109/LCA.2024.3370402
  3. Publisher: 2024
  4. Abstract:
  5. Sub-byte quantization on popular vector ISAs suffers from heavy waste of vector as well as memory bandwidth. The latest methods pack a number of quantized data in one vector, but have to pad them with empty bits to avoid overflow to neighbours. We remove even these empty bits and provide full utilization of the vector andmemory bandwidth by our data-layout/compute codesign scheme. We implemented FullPack on TFLite for Vector- Matrix multiplication and showed up to 6.7X speedup, 2.75X on average on single layers, which translated to 1.56-2.11X end-to-end speedup on DeepSpeech. © 2002-2011 IEEE
  6. Keywords:
  7. Deep learning ; Hardware acceleration ; Bandwidth ; Computer graphics ; Deep learning ; Image coding ; Matrix algebra ; Program processors ; Quantization (signal) ; Vector quantization ; Vectors ; Full vectors ; General purpose CPUs ; Hardware acceleration ; Load modeling ; Memory bandwidths ; Memory-management ; Quantization (signal) ; Register ; Vector-matrix multiplications ; Graphics processing unit
  8. Source: IEEE Computer Architecture Letters ; Volume 23, Issue 2 , 2024 , Pages 142-145 ; 15566056 (ISSN)
  9. URL: https://ieeexplore.ieee.org/document/10449368