Articles

Prompt, context and constraints: what the request envelope decides

The model sees what the envelope builds: template, budget and stopping conditions.

Reading: 6 minAutomation & AI

Article cover: Prompt, context and constraints: what the request envelope decides

The L1 sheet of this area explains what prompts, system instructions and messages mean to a model. This sheet is about the envelope an API puts around them, because the envelope — not the model — decides what the model actually sees: a chat template turns your turns into a token sequence, and a token budget decides where generation stops.

Three layers of envelope

1. Messages or a prompt. A chat API takes an array of turns with roles; a completion API takes a string. The engine converts the first into the second, and everything below happens to that conversion.

2. The chat template. Every instruction-tuned model was trained on one exact formatting of those turns — markers around the system message, role names, separators. The engine applies the template that the model’s own metadata declares: in a llama.cpp server, --jinja, --no-jinja selects “whether to use jinja template engine for chat (default: enabled)”, and --chat-template overrides the default, which is the “template taken from model’s metadata”. --chat-template-kwargs passes parameters into that template, and --skip-chat-parsing forces a pure content parser instead of splitting reasoning and tool calls out of the output.

Two things follow. First, the template is inspectable: POST /apply-template “Returns a JSON object with a field prompt containing a string of the input messages formatted according to the model’s chat template format” — the fastest way to see what the model is really going to read. Second, the template travels with the model — /props reports “chat_template - the model’s original Jinja2 prompt template” — which is one of the reasons a self-describing container matters (see the formats sheet in this batch). Using a different template from the one a model was trained with is a silent quality failure, not an error message.

3. The per-request constraints. These bound the output and stop the generation:

  • the output budget — -n, --predict, --n-predict N, documented as “number of tokens to predict (default: -1, -1 = infinity)”. Infinity is not a plan; server engines also accept the per-request equivalent, and Ollama exposes the same family through its options object alongside temperature and seed.
  • stopping conditions — the stop sequences an engine checks for, so that an answer ends where the caller wants rather than where the budget runs out.
  • sampling parameters — temperature and the rest, whose semantics belong to the L1 sheet; what matters here is that they travel in the request and can be set per call. Reproducibility is the same story: Ollama’s own guidance is “For reproducible outputs, set seed to a number”.
  • the system message, which a request may override: Ollama documents system as a system message that “overrides what is defined in the Modelfile”.

The budget is shared, and it is finite

The prompt and the output are paid from the same allowance, set by -c, --ctx-size — “size of the prompt context (default: 0, 0 = loaded from model)”. A request whose input plus requested output exceeds that allowance cannot be satisfied as written, and the engine’s behaviour in that case is defined by that engine’s documentation rather than by a general rule; the sheet that follows on streaming and timeouts assumes the request was accepted in the first place.

Concurrency makes the allowance smaller, not larger: server slots (-np, --parallel) share one key-value pool, and --kv-unified-per-slot N exists precisely to set a “context limit per parallel slot” when several requests are in flight. A client that sends a long document to a server sized for many short chats is choosing a different prompt length from the one it tested with.

The off switch, and why it is worth knowing

An engine can bypass formatting entirely: Ollama’s raw parameter lets the caller “bypass the templating system and provide a full prompt”, with the documented consequence that “raw mode will not return a context”. That option is the clearest demonstration that the envelope is a real, separable layer: template, roles and budget are decisions made around the model, and each of them can be replaced.

A checklist for the envelope

  1. Which model revision, and which chat template (its metadata’s, or an override you chose deliberately)?
  2. Do the roles and the message contents match what the template expects?
  3. Does prompt + output budget fit the configured context, at the concurrency you expect?
  4. Are stopping conditions set, or will generation end only when the budget is exhausted?
  5. Which sampling parameters are set per request, and are they pinned where reproducibility matters?
  6. What does the engine do when the allowance is exceeded — the answer is in its documentation.

Level and prerequisites. L2 — operational: the reader must be able to shape a request deliberately and to know which layer each parameter belongs to. The semantics of instructions and roles stay with the L1 sheet; the window’s meaning and practical limits stay with the L1 context sheet; structured output is the next sheet in this batch. Prerequisites: the L1 sheets on prompts and on the context window.

Where to go next

References

  • ggml-org — llama.cpp server — the Jinja template options and their defaults, POST /apply-template, the chat_template property, --ctx-size, --predict, --parallel and --kv-unified-per-slot.
  • Ollama — API reference — the options, system, raw and seed parameters and the streaming default.
  • The L1 sheets of this area — the meaning of prompts, system instructions and the context window.