Quick Contact

✉
fahimkhan20148@gmail.com
📱
+971 507 286 133
Back to Notes
March 30, 2026
llmquantizationoptimization

Quantization

Quantization is the process of mapping high-precision floating-point numbers (e.g., FP32 or FP16) to lower-precision formats (e.g., INT8, INT4, or even 1.58-bit).

Purpose

  • Memory Reduction: Fits larger models into smaller VRAM (e.g., consumer GPUs).
  • Speed: Improves inference throughput by utilizing specialized low-precision hardware instructions.

Common Formats

  • GPTQ: Post-training quantization for weight-only.
  • AWQ: Activation-aware Weight Quantization.
  • TurboQuant: A specialized kernel for fast quantized inference.
  • BitNet: 1-bit LLMs that utilize binary or ternary weights.

#TODO: Deep dive into the trade-offs between precision and perplexity.