AI Acronyms
Common abbreviations and short-hands in artificial intelligence.
14 terms
AI moves fast enough that the acronyms often outrun the explanations — RLHF, DPO, LoRA, RoPE, KTO show up in papers and release notes long before anyone writes a plain-English definition. This category exists specifically to close that gap: each entry expands the acronym, explains what the technique actually does, and links to the fuller explainer where one exists.
Several of these are alternative or newer approaches to the same underlying problem — RLHF, DPO, and KTO are all ways to align a model to human preferences, just with different training mechanics and trade-offs. Reading a few of these together tends to be more useful than reading any one in isolation, since the acronym alone rarely tells you why a team would pick one over another.
- BM25 Retrieval
A probabilistic term-matching keyword search algorithm widely used in information retrieval.
- Direct Preference Optimization (DPO)
An alignment algorithm that optimizes policy networks directly using pairwise preference data without reward model training.
- Infrastructure as Code (IaC)
Defining cloud infrastructure in version-controlled files instead of console clicks — reviewable, repeatable, and recoverable.
- Inverted File Index (IVF)
A vector search optimization that clusters vector spaces to limit search scopes.
- Kahneman-Tversky Optimization (KTO)
An alignment algorithm that optimizes models using binary utility signals (good/bad) rather than pairwise preferences.
- NormalFloat 4 (NF4)
An information-theoretically optimal quantile quantization data type designed to compress parameters.
- Parameter-Efficient Fine-Tuning (PEFT)
A collection of fine-tuning techniques that adapt pre-trained models by modifying only a tiny fraction of parameters.
- Post-Training Quantization (PTQ)
An offline compression technique that converts weights to lower precision after model training finishes.
- Product Quantization (PQ)
A lossy vector compression technique that reduces database RAM footprints.
- Proximal Policy Optimization (PPO)
An on-policy reinforcement learning algorithm that restricts step updates to maintain training stability.
- QLoRA
A fine-tuning method that backpropagates gradients through frozen, 4-bit quantized base models into LoRA adapters.
- Quantization-Aware Training (QAT)
A training process that models quantization error during the forward pass to minimize precision loss.
- Reinforcement Learning from Human Feedback
The post-training technique that turns a raw next-token predictor into a helpful assistant by optimizing against human preference judgments.
- RLHF Reward Model
A scoring model trained on human feedback to evaluate language model generation quality.