Articles

Temperature, top-p and output limits

How one token is chosen out of a distribution, and what a token limit actually bounds.

Reading: 5 minAutomation & AI

Article cover: Temperature, top-p and output limits

A language model does not decide on a sentence and then type it out. At each step it produces a probability distribution over its vocabulary for the next token, and a separate step — decoding — picks one token from that distribution. Temperature, top-p (nucleus sampling) and top-k act on that step: they reshape and truncate the distribution before a token is taken. They add nothing to what the model knows and they check nothing. An output-token limit is a different control — a ceiling on how much may be generated, not a length the model aims for.

The mental model: a distribution, then a choice

The model’s job and the decoder’s job are two different jobs, and the order matters:

context tokens ──► model ──► probability for every token in the vocabulary
                                   │
                      decoding reshapes and truncates the distribution
                      temperature · top-p (nucleus) · top-k
                                   │
                     keep a candidate set, discard the tail
                                   │
                     greedy: take the maximum   |   sampling: draw
                                   │
                             one chosen token
                                   │
                     append it to the context, repeat (autoregressive)
                                   ▼
                     stop at end-of-sequence, or at the output limit

The model turns the context into a distribution; the parameters reshape and truncate it; one token is drawn from what is left and appended to the context, and the loop runs again. Everything to the right of the distribution is arithmetic over probabilities — no lookup, no check against the world.

Terminology the reader needs

  • Decoding — the algorithm that turns the model’s per-step distribution into tokens; the Hugging Face guide names three families: greedy search, beam search and sampling.
  • Greedy search — “the simplest decoding method”; it “selects the word with the highest probability as its next word” at each step. Beam search keeps several candidates at once but is “not guaranteed to find the most likely output”.
  • Sampling — “randomly picking the next word … according to its conditional probability distribution”; generation this way “is not deterministic anymore”.
  • Temperature — OpenAI: “sampling temperature … between 0 and 2”, default 1.0 in the examples; 0.8 makes output “more random”, 0.2 makes it “more focused and deterministic”.
  • Top-p (nucleus sampling) — OpenAI: “an alternative to sampling with temperature”; only “the tokens with top_p probability mass” are considered, so 0.1 means the top 10% probability mass.
  • Top-k — from the Hugging Face guide: “the K most likely next words are filtered and the probability mass is redistributed among only those K next words”; top_k=0 deactivates it.
  • Output limit — OpenAI’s max_completion_tokens: “an upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens”.

The mechanism, step by step

  1. The model reads the context and returns, for the next position, a probability over every token in its vocabulary; it is autoregressive, and what a token is belongs to a sibling sheet.
  2. The candidate set can be narrowed. Top-p keeps the smallest set of words whose cumulative probability exceeds p and redistributes the mass among them; top-k keeps the K most likely and does the same. Neither source read here states an order between this and temperature, and the two are separate operations on the same distribution.
  3. Temperature changes the shape of the distribution. Lowering it sharpens it — increasing high-probability likelihoods and decreasing low-probability ones; raising it flattens it.
  4. A token is chosen. Greedy takes the most probable; sampling draws one at random from the reshaped, truncated distribution.
  5. The token joins the context and the loop repeats, ending at an end-of-sequence token or at the output limit.
  6. On combining top-p with temperature, OpenAI’s note is: “We generally recommend altering this or temperature but not both.”

Greedy, temperature and top-p compared

Approach What it does Parameter (OpenAI spec unless noted) What the reader should expect
Greedy decoding takes the most probable next token at each step none — the guide describes it as the simplest method as the guide puts it, it can miss high-probability words hidden behind a low-probability one
Temperature sampling draws from the distribution after temperature has reshaped it temperature, between 0 and 2, default 1.0 in the examples lower values narrow the distribution, higher values spread it out
Top-p (nucleus) samples only from the smallest set of tokens whose cumulative probability exceeds p top_p, 0 to 1, default 1.0 in the examples the size of the candidate set adapts to how predictable the next token is
Top-k samples only from the K most likely tokens top_k in the Hugging Face guide; top_k=0 deactivates it a fixed-size candidate set, whatever the shape of the distribution

The output limit: a ceiling, not a target

max_completion_tokens is documented as “an upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens”. Two consequences follow. First, the bound counts everything the model generates, including reasoning tokens the caller may never see, so the visible answer can be shorter than the limit suggests. Second, an upper bound stops generation at the cap; it does not make the model conclude. Read that as my framing: the specification defines a bound, and the difference between “stopped because it finished” and “stopped because it hit the ceiling” is what a caller has to handle. The context window is a different limit, owned by another sheet.

Limits, and the common conceptual errors

  • Temperature is not a correctness dial. The specification’s wording is about random versus focused and deterministic — not about truth; neither source read here says a low temperature makes output more accurate, and this sheet claims no such thing. Why fluent output can still be false belongs to the hallucination and verification sheet.
  • Temperature 0 is not a promise of identical output. The guide says that in its limit, as temperature approaches 0, temperature-scaled sampling “becomes equal to greedy decoding”; the specification’s seed parameter says “our system will make a best effort to sample deterministically”, and “Determinism is not guaranteed”.
  • Sampling introduces variation on purpose. With sampling, the guide says, generation “is not deterministic anymore” — the mechanism behind two identical prompts producing two different answers.
  • A limit is not a length instruction. Asking for a long answer under a small cap yields text cut at the cap: it truncates, it does not summarise.

Level and prerequisites

L1 — the vocabulary of decoding and the difference between shaping a choice and bounding a length; no procedure and no tuning advice. Prerequisites: the sheets on tokens and autoregressive generation, and on prompts and messages; neither is published yet.

Where to go next

References

  • OpenAI, OpenAI API specification (openai-openapi.yaml) — the documented descriptions of temperature, top_p, max_completion_tokens and seed, and the note about altering temperature or top_p but not both.
  • Hugging Face, How to generate text: using different decoding methods for language generation with Transformers (hf-how-to-generate.html, edited July 2023) — the three decoding families, greedy search, beam search, sampling, temperature, top-k and top-p (nucleus) sampling.