Tokens, tokenizers and autoregressive generation
Why text becomes tokens, what a tokenizer and vocabulary are, and how generation works.

A language model never receives your words. It receives a sequence of numbers — tokens — produced by a tokenizer, which maps text to tokens and back against a fixed vocabulary. That one fact explains three surprises: why the same sentence can cost different amounts on different models, why “about 4 characters” per token is a rule of thumb and not a conversion factor, and why generation is a loop rather than one action. Autoregressive names that loop: the model predicts the next token conditioned on the tokens before it, then feeds its own output back in. Everything measured about a call — how much text fits, what it costs — is counted in tokens, not in characters or words.
The mental model: text in, one token out at a time
your text
│
▼ tokenizer (text → tokens, using the model's vocabulary)
token sequence [t1][t2][t3] … [tn]
│
▼ the model reads the token sequence, not the text
probability distribution over the whole vocabulary,
for the NEXT token, given the tokens so far
│
▼ one token is selected from it
the token is appended: [t1][t2][t3] … [tn][tn+1]
│
└────────────────► the longer sequence becomes the prefix; repeat
│
▼ detokenizer (tokens → text)
text you can read
The direction matters: text becomes tokens first, the model works only on tokens, and the prose you
read is decoded back at the end. Adjacent subjects are not developed here: how a token is selected
belongs to temperature-top-p-and-output-limits, the budget the loop consumes to
context-window-and-practical-limits.
Terminology the reader needs
Token — the unit the model operates on. tiktoken states it plainly: language models “don’t see
text like you and I, instead they see a sequence of numbers (known as tokens)”. A token is neither a
word nor a character: tokenization algorithms “split text into units between words and characters”, so
a common word stays whole while a rare one decomposes — the documentation offers two splits for
annoyingly, ["annoying", "ly"] and ["annoy", "ing", "ly"], depending on the vocabulary.
Tokenizer — the component that maps in both directions. tiktoken is “a fast BPE tokeniser for use with OpenAI’s models”; the Hugging Face documentation names three subword algorithms in this family — Byte pair encoding (BPE), Unigram and WordPiece.
Vocabulary — “the set of all tokens used by the model”, in the Gemini documentation’s words; splitting text into tokens is tokenization. It is finite and fixed for a given model: the same Hugging Face page records 50,257 tokens for GPT-2’s byte-level BPE vocabulary and 40,478 for an earlier GPT BPE vocabulary.
Byte-level BPE — including all Unicode characters would make the base vocabulary enormous, so byte-level BPE “uses 256 byte values as the base vocabulary instead, ensuring every word can be tokenized”.
Generation, step by step, in the correct direction
- Encode. The tokenizer turns input text into a token sequence. The mapping is reversible and lossless, so it also runs backwards.
- Read. The model turns tokens into vectors and processes the sequence; no word as such enters the computation.
- Predict the next token. This is what autoregressive names. NIST’s Generative AI Profile states the mechanism: “LLMs predict the next token or word in a sentence or phrase”. In a transformer this is the decoder’s output converted into “predicted next-token probabilities” — one probability per vocabulary entry.
- Select one token from that distribution. Which token, and with which parameters, is decoding — a separate subject, named above.
- Append and repeat. The chosen token joins the sequence, the longer sequence becomes the prefix, and the model predicts the next one. GPT-3’s paper calls a model built this way “an autoregressive language model”.
- Decode and stop. When the model emits a special end-of-text token, generation stops and the tokens are mapped back to text.
Why a token count is not a word count
The four units that get confused:
| Unit | What it is | Where the number comes from |
|---|---|---|
| Character | a letter, digit or symbol | the text editor’s string length |
| Word | a whitespace-delimited unit | a word count of the prose |
| Byte | an 8-bit value | the base layer of byte-level BPE |
| Token | the unit the model reads and the unit a call is billed on | the model’s own tokenizer |
None of these converts cleanly into another. The Gemini documentation states that “a token is equivalent to about 4 characters” and that “100 tokens is equal to about 60-80 English words”; tiktoken says that in practice “each token corresponds to about 4 bytes”. Neither is a factor to reuse across providers: the two tokenizers hold different vocabularies over different training text. Counts differ for the same text by language too — GPT-3 reuses GPT-2’s byte-level BPE tokenizer, which “was developed for an almost entirely English training dataset”.
The common conceptual error, and what it costs
The error is reading token as word. A token can be a whole word, a fragment, a punctuation mark or a single character, and the boundary moves with the model and the language.
Two consequences follow. First, an estimate is not a count: the “about 4 characters” figure is one vendor’s rule of thumb for its own models, so the reliable count comes from the tokenizer of the model you are calling. Second, tokens hide the characters inside them — GPT-3’s paper notes that tasks which “requires character-level manipulations” are awkward precisely because its BPE encoding “operates on significant fractions of a word” rather than on characters.
Because the model sees only tokens, the quantities that govern a call — how much text fits, what it
costs — are counted in tokens: the Gemini documentation notes the “cost of a call … is determined in
part by the number of input and output tokens”. The budget side is
context-window-and-practical-limits; the nesting of AI, machine learning, deep learning and LLM is
ai-machine-learning-deep-learning-and-llm.
Level and prerequisites
L1 — the vocabulary and the mechanism only: text → tokens → next-token prediction → tokens → text, with no tokenizer installed or trained, no decoding parameter and no benchmark. Prerequisites: none.
Where to go next
- Automation & AI — the area this sheet belongs to.
References
- OpenAI, tiktoken (README, local extract) — models see token numbers rather than text; BPE as the text-to-token conversion; losslessness; the “about 4 bytes” ratio; encodings named per model family.
- Hugging Face, Tokenization algorithms (Transformers documentation) — the three subword
algorithms; “units between words and characters”; the
annoyinglyexample; byte-level BPE and its 256-value base vocabulary; concrete vocabulary sizes. - Google, Gemini API — tokens and token counting — the definition of vocabulary and tokenization; the token character/word rule of thumb; cost determined by input and output tokens.
- NIST, AI 600-1, Generative AI Profile — next-token prediction as the mechanism generative models are designed around.
- Brown et al., Language Models are Few-Shot Learners (GPT-3, 2020) — “autoregressive language model”; the byte-level BPE tokenizer reused from GPT-2 and its English-skewed training data; the character-level-manipulation limitation.
- Vaswani et al., Attention Is All You Need (2017) — the transformer’s conversion of tokens to vectors and of the decoder output to predicted next-token probabilities.
- Hugging Face, How to generate text — the end-of-text (EOS) token that ends the loop.