Generative Low-bitwidth Data Free Quantization
Published in arXiv (Cornell University) • Mar 7, 2020
NobleIDNI4P31W11R01S76
Authors:,,
Shoukai Xu
Haokun Li
Bohan Zhuang
Abstract
Neural network quantization is an effective way to compress deep models and improve their execution latency and energy efficiency, so that they can be deployed on mobile or embedded devices. Existing quantization methods require original data for calibration or fine-tuning to get better performance....
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!