Focused quantization for sparse CNNs

Zhao, Y; Gao, X; Bates, D; Mullins, R; Xu, CZ

Focused quantization for sparse CNNs

Accepted version

Peer-reviewed

Repository URI

https://www.repository.cam.ac.uk/handle/1810/304831

Repository DOI

https://doi.org/10.17863/CAM.51914

Files

Accepted version (1.4 MB)

Type

Conference Object

Authors

Zhao, Y

Gao, X

Bates, Daniel

https://orcid.org/0000-0002-9526-3342

Mullins, Robert

https://orcid.org/0000-0002-8393-2748

Xu, CZ

Abstract

Deep convolutional neural networks (CNNs) are powerful tools for a wide range of vision tasks, but the enormous amount of memory and compute resources required by CNNs pose a challenge in deploying them on constrained devices. Existing compression techniques, while excelling at reducing model sizes, struggle to be computationally friendly. In this paper, we attend to the statistical properties of sparse CNNs and present focused quantization, a novel quantization strategy based on power-of-two values, which exploits the weight distributions after fine-grained pruning. The proposed method dynamically discovers the most effective numerical representation for weights in layers with varying sparsities, significantly reducing model sizes. Multiplications in quantized CNNs are replaced with much cheaper bit-shift operations for efficient inference. Coupled with lossless encoding, we built a compression pipeline that provides CNNs with high compression ratios (CR), low computation cost and minimal loss in accuracy. In ResNet-50, we achieved a 18.08x CR with only 0.24% loss in top-5 accuracy, outperforming existing compression methods. We fully compressed a ResNet-18 and found that it is not only higher in CR and top-5 accuracy, but also more hardware efficient as it requires fewer logic gates to implement when compared to other state-of-the-art quantization methods assuming the same throughput.

Keywords

cs.LG, cs.LG, stat.ML

Journal Title

Advances in Neural Information Processing Systems

Conference Name

Advances in Neural Information Processing Systems

Journal ISSN

1049-5258

Volume Title

32

Publisher DOI

https://doi.org/10.17863/CAM.51914

Rights

Sponsorship

This work is supported in part by the National Key R&D Program of China (No. 2018YFB1004804), the National Natural Science Foundation of China (No. 61806192). We thank EPSRC for providing Yiren Zhao his doctoral scholarship.

Collections

Cambridge University Research Outputs