TurboQuant
TurboQuant is an optimized inference kernel designed for extremely fast quantized LLM inference on GPUs.
Key Features
- Optimized Kernels: Hand-tuned CUDA kernels for specific low-precision operations.
- Minimal Overhead: Reduces the compute overhead of de-quantizing weights on the fly.
#TODO: Compare TurboQuant with other kernels like BitsAndBytes or AutoGPTQ.
