AI Model Compression Techniques
DOI:
https://doi.org/10.5281/zenodo.19614793Keywords:
model compression; quantisation; pruning; knowledge distillation; low-rank factorisation; GPTQ; LoRA; efficient inference; edge AI; parameter-efficient fine-tuningAbstract
The rapid growth of deep learning model sizes -- from millions to hundreds of billions of parameters -- has created a severe deployment gap between what frontier research produces and what edge devices, resource-constrained servers, and latency-sensitive applications can run. Model compression addresses this gap through four primary techniques: pruning (removing redundant weights or structures), quantisation (reducing numerical precision), knowledge distillation (training smaller student models to replicate larger teacher behaviour), and low-rank factorisation (decomposing weight matrices into compact representations). This study provides the most comprehensive systematic comparison to date, evaluating sixteen compression configurations across five base architectures (BERT-Large, LLaMA-2-13B, ResNet-152, ViT-L/16, and Whisper-Large) on eight task domains. Compression configurations span 1-bit to 8-bit quantisation (GPTQ, AWQ, LLM.int8(), SmoothQuant), unstructured and structured pruning at 50-90% sparsity, task-specific and task-agnostic distillation, and LoRA/QLoRA parameter-efficient fine-tuning. Key findings: 4-bit GPTQ quantisation preserves 97.4-98.8% of full-precision task performance across LLM tasks with 3.8x memory reduction. Structured N:M sparsity at 2:4 achieves 1.72x NVIDIA GPU inference speedup with 94.2-96.8% accuracy retention. Task-specific distillation achieves 82-88% teacher performance at 10-15x parameter reduction. A compression recipe selection framework balancing accuracy retention, memory reduction, and inference speedup is proposedDownloads
Published
2026-08-19
Issue
Section
Articles
How to Cite
AI Model Compression Techniques. (2026). Bio-QI Journal, 3(2), 81-89. https://doi.org/10.5281/zenodo.19614793

