Aliases: Sampling, Top-p
Temperature (Sampling)
The parameter controlling randomness in LLM output — low values pick likely tokens for consistency, high values spread probability for variety.
Last reviewed: July 25, 2026
What is temperature?
At each step an LLM produces a probability for every possible next token. Temperature reshapes that distribution before sampling: near 0, the model almost always picks the most likely token (deterministic-ish, focused); around 1, it samples proportionally (varied, creative); higher, it flattens toward randomness. Top-p (nucleus sampling) is the common companion — sample only from the smallest set of tokens whose probabilities sum to p.
Practical settings
- 0–0.3 — extraction, classification, structured output, code, anything you’ll parse programmatically.
- 0.5–0.8 — general chat and writing; most defaults live here.
- 0.9–1.2 — brainstorming, creative variety, generating diverse candidates.
- Adjust temperature or top-p, not both aggressively — they compound.
The determinism myth
Temperature 0 does not guarantee identical outputs. Floating-point non-associativity, batching effects, and infrastructure changes introduce variation even at zero, and providers update models behind stable names. If your system depends on exact reproducibility, you need output validation and evals — not a temperature setting.
A note on reasoning models
Many reasoning models ignore or restrict sampling parameters during their thinking phase — check provider docs before assuming your temperature knob does anything.
What people get wrong
- Blaming temperature for hallucination. Low temperature makes output consistent, not correct — a confidently wrong answer just gets repeated consistently (hallucination is a knowledge problem, not a sampling problem).
- One global setting. Temperature is per-request; route extraction calls at 0 and brainstorming calls at 1 from the same app.
- Cranking it for “better creativity.” Above ~1.3 output degrades into incoherence long before it becomes interestingly original.
Try it in code
See temperature reshape a distribution — same math the playground uses:
import math, random
logits = {"Paris": 8.0, "located": 5.2, "the": 4.9, "Lyon": 1.4}
def softmax(logits, temperature):
exps = {t: math.exp(l / temperature) for t, l in logits.items()}
total = sum(exps.values())
return {t: v / total for t, v in exps.items()}
for temp in (0.2, 1.0, 2.0):
probs = softmax(logits, temp)
sample = random.choices(list(probs), weights=probs.values(), k=10)
print(f"T={temp}: P(Paris)={probs['Paris']:.2f} samples={sample}")
# T=0.2: P(Paris)≈1.00 - deterministic-ish
# T=2.0: probability spread out - variety, then noise
Try it live in our temperature playground.
Practical Temperature Guidance
A temperature of 0 (or very close to it) makes generation effectively deterministic, always selecting the single most probable next token — appropriate for tasks with a clear correct answer, like classification, structured data extraction, or code generation, where consistency matters more than variety. Higher temperatures (commonly 0.7-1.0) are more appropriate for creative writing, brainstorming, or conversational applications where some variability in phrasing is desirable and there’s no single “correct” output to converge on. Temperatures much above 1.0 are rarely useful in production, since they push the model toward increasingly unlikely, often incoherent token choices — most applications settle on a value in the 0-1 range tuned to how much consistency versus variety the specific use case calls for.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.