Integrating FITcompress to PQuant
Authors/Creators
Description
This work presents the integration of FITCompress into the PQuant framework, adapting the algorithm to support fixed-point quantization and various pruning methods for practical hardware deployment. FITCompress was modified to work with PQuant’s Q(m,n) fixed-point representation by developing a principled approach for distributing bits between integer and fractional parts based on dynamic range analysis. To address the lack of explicit activation quantization in the original FITCompress paper, a calibration-based method was implemented that determines optimal bit allocations for activation units, pooling layers and model inputs. Our integration supports joint optimization of quantization and pruning for methods that use a global sparsity target, such as PDP and Wanda, while also offering quantization-only optimization for other pruning approaches. Experimental results on ResNet-20 with CIFAR-10 show that FITCompress achieves substantial compression with only minimal accuracy loss. With PDP and Wanda, the compressed models maintained accuracies within about one percentage point of the uncompressed baseline, even when reduced to just $6\%$ (PDP) and $4\%$ (Wanda) of the original model's size. FITCompress was also analyzed with other pruning methods that do not rely on a global sparsity target, demonstrating that the approach remains flexible even without direct pruning optimization. This integration bridges the gap between FITCompress’s theoretical framework and hardware-oriented quantization, enabling efficient deployment on resource-constrained devices.
Files
CERN_report_final.pdf
Files
(1.6 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:b78398191dde06b8a39045d880ca5f9c
|
1.6 MB | Preview Download |