Quantization
Quantization is the process of mapping high-precision floating-point numbers (e.g., FP32 or FP16) to lower-precision formats (e.g., INT8, INT4, or even 1.58-bit).
Purpose
- Memory Reduction: Fits larger models into smaller VRAM (e.g., consumer GPUs).
- Speed: Improves inference throughput by utilizing specialized low-precision hardware instructions.
Common Formats
- GPTQ: Post-training quantization for weight-only.
- AWQ: Activation-aware Weight Quantization.
- TurboQuant: A specialized kernel for fast quantized inference.
- BitNet: 1-bit LLMs that utilize binary or ternary weights.
#TODO: Deep dive into the trade-offs between precision and perplexity.
