Aliases: Model serving, LLM serving
Inference
Running a trained model to produce outputs in production — where nearly all AI compute cost and latency engineering lives.
Last reviewed: July 25, 2026
What is inference?
Inference is using a trained model to generate outputs. Training happens once; inference happens on every user request, forever — which is why inference, not training, dominates AI operating cost for almost every company that isn’t a frontier lab.
The two ways to run it
- API (managed) — call OpenAI/Anthropic/Google/AWS Bedrock. Zero ops, per-token pricing, frontier quality. The right default.
- Self-hosted — run open-weight models (Llama, Mistral, Qwen) on your own GPUs with a serving engine like vLLM or TensorRT-LLM. Justified by data residency, latency control, or sustained high volume on a narrow task — usually with a fine-tuned small model.
Key metrics
TTFT (time to first token — perceived responsiveness), tokens/sec per request (streaming speed), throughput per GPU (your unit economics), and cost per task. In 2026, cost-per-task is a board-level metric for AI-heavy SaaS: a single enterprise customer can drive five or six figures of monthly inference spend, making inference a first-class COGS line, not a feature flag.
Cost reality
GPU serving is a utilization game: an idle A100/H100 burns $1–4+/hour regardless. Batching (continuous batching in vLLM), quantization, prompt caching, and routing easy requests to small models are the standard levers — teams commonly cut inference bills 5–10× with a model-routing + caching pass before touching hardware.
What people get wrong
- Self-hosting to “save money” at low volume. Under sustained thousands of requests/day, APIs are almost always cheaper than idle GPUs plus an ML-ops engineer.
- Benchmarking one request at a time. Production throughput depends on batching behavior under concurrent load.
- Ignoring output length. Generation is sequential; capping and streaming output does more for perceived latency than faster hardware.
The Two Phases of Inference
LLM inference splits into two distinct computational phases with very different performance characteristics. The prefill phase processes the entire input prompt in parallel in a single forward pass, building the initial KV cache — this phase is compute-bound, meaning its speed scales with available raw GPU compute (FLOPs). The decoding phase then generates output tokens one at a time, with each new token requiring another forward pass that depends on the previous one — this phase is memory-bandwidth-bound rather than compute-bound, since generating each token mostly involves reading the model’s weights and the growing KV cache from memory rather than performing large amounts of new computation. Understanding this split explains several practical realities of LLM serving: why input tokens are typically cheaper to process than output tokens in most API pricing, and why serving optimizations tend to target the two phases with entirely different techniques.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.