Large Language Models (LLMs) MOC
A central index for exploring concepts, architectures, and strategies related to LLM.
Core Concepts
- Tokenization: Character, word, and sub-word (BPE) techniques.
- Encoder / decoder: Feature extraction (bidirectional) vs autoregressive generation (causal).
- Needle-in-a-Haystack: Benchmarking long-context retrieval and biases.
Retrieval-Augmented Generation (RAG)
- RAG Overview: Modalities (Text, Hybrid, Full Multimodal).
- RAG Strategies MOC: (Candidate MOC)
- Re-ranking: Two-stage retrieval for precision.
- Agentic RAG: Autonomous tool selection.
- Knowledge Graphs (GraphRAG): Relational retrieval.
Performance & Optimization
- Quantization: Model compression (GPTQ, AWQ, 1.58-bit).
- TurboQuant: Fast quantized kernels.
- Token Caching: Reusing KV caches for speed.
- Speculative Decoding: Draft model generation with target validation.
- Serving Engines:
Multimodal Models
- MLLM: Multimodal LLMs capable of processing images, audio, and video.
- HYMMRag: Hybrid Multimodal RAG.
- FMMRag: Full Multimodal RAG.
