Articles

Streaming, retry, timeout and idempotency

Streaming changes when you know; idempotency decides whether you may ask again.

Reading: 6 minAutomation & AI

Article cover: Streaming, retry, timeout and idempotency

Three mechanisms, one question. Streaming changes when a client learns what the model said; timeouts are the only control that turns silence into a decision; idempotency decides whether the client may ask again. They meet exactly where integrations break: in the moment a request has been sent and no answer has arrived.

Streaming

Streaming means the response arrives progressively instead of in one piece. Ollama documents it in one line — “Certain endpoints stream responses as JSON objects. Streaming can be disabled by providing {"stream": false} for these endpoints” — and llama.cpp’s server keeps the channel alive between events, --sse-ping-interval N, documented as a “server SSE ping interval in seconds (-1 = disabled, default: 30)”.

Four consequences follow, and the fourth is the one that causes bugs:

  1. The first token appears immediately. Time to first token stops being hidden behind total latency — which is why it is its own number in the measurement sheet.
  2. A stream is a sequence, not a document. It can end early: a dropped connection leaves a partial answer that looks like a shorter one. The client must know where the stream was meant to end — a completion marker, a closing event — and treat anything without it as incomplete.
  3. The last chunk is the only authority. Half-finished JSON at the moment of the cut is not JSON, which is where the schema check of the previous sheet earns its place.
  4. A total timeout is the wrong instrument. A long answer legitimately takes a long time, so a fixed deadline kills valid work; what streaming enables is an idle timeout, a limit on the gap since the last chunk.

Timeouts

The engineering rule is older than LLM APIs and applies to them unchanged: “A best practice in Amazon is to set a timeout on any remote call, and generally on any call across processes even on the same box. This includes both a connection timeout and a request timeout.” Two distinct timers, because two distinct things can hang: reaching the service, and waiting for it.

Choosing the value is the hard part: too high a timeout “reduces its usefulness, because resources are still consumed while the client waits”; too low it produces false failures, which are worse than they look because they cause retries. The documented starting point is the downstream service’s own latency measurements plus a chosen acceptable rate of false timeouts. On the server side the same idea exists from the other end: -to, --timeout N, the “server read/write timeout in seconds (default: 3600)”.

A timeout is not a failure. It is an unknown state. The request may have been received and completed, received and abandoned, or never received; the rest of this sheet is about acting on that uncertainty.

Idempotency

HTTP has a precise answer to “may I send this again”. A request method “is considered ‘idempotent’ if the intended effect on the server of multiple identical requests with that method is the same as the effect for a single such request”, and of the methods the specification defines, “PUT, DELETE, and safe request methods are idempotent”. It then states the rule a client must obey: “A client SHOULD NOT automatically retry a request with a non-idempotent method unless it has some means to know that the request semantics are actually idempotent, regardless of the method, or some means to detect that the original request was never applied.” Intermediaries are held to a stricter version still: “A proxy MUST NOT automatically retry non-idempotent requests.”

A generation request is a POST to a chat endpoint, so by that definition it is not idempotent. Two very different effects could be duplicated:

  • The computation and its cost. Sending the same prompt twice generates twice and bills twice. A duplicate here is a waste, not a corruption.
  • An action the request triggers. If the request causes something outside the model — a tool runs, a record is written, a message is sent — a duplicate is a real, visible error.

For the first, a retry with backoff is a cost decision. For the second, make the operation idempotent at the boundary you control, with an idempotency key — a client-generated identifier that lets the service recognise a repeat. The OpenAI API description defines exactly that for an endpoint that submits messages: an optional header, Idempotency-Key, “a client-generated key that makes retries of submitted messages idempotent”. Where no such key exists, the honest client does not retry the action automatically; it reports the unknown state to the caller.

Backoff, jitter and stopping

Retrying is not free and is not always safe: “It’s not always safe to retry. A retry can increase the load on the system being called, if the system is already failing because it’s approaching an overload. To avoid this problem, we implement our clients to use backoff.” The pattern the major platforms describe has four parts:

  • Retry only the transient class, properly classified per the previous sheet: the fault that is “typically self-correcting, and if the action that triggered a fault is repeated after a suitable delay it’s likely to be successful”.
  • Increase the delay between attempts, up to a maximum number of attempts rather than until it works.
  • Add jitter. Clients failing together retry in lockstep; randomised delays spread them out, which is why production clients rarely use a fixed interval.
  • Stop. “An aggressive retry policy with minimal delay between attempts, and a large number of retries, could further degrade a busy service”; after repeated failures it is better to stop sending and report, with a circuit breaker for a dependency that is down.

Two details are worth carrying into any client. Provider SDKs already implement the pattern — Anthropic’s documentation states that its SDK “automatically retries transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the retry-after header” — so a hand-rolled loop on top of such an SDK multiplies attempts rather than improving them. And a Retry-After value is the server telling you the earliest moment at which asking again is reasonable, not a hint to round down.

The order of decisions

  1. Classify the response: fix it, slow it, wait it, or unknown state.
  2. Ask whether the operation is idempotent, or can be made so with a key; if not, do not retry automatically.
  3. Set a connection timeout and a request timeout — idle, when streaming.
  4. Retry only transient classes, with increasing delays, jitter and capped attempts.
  5. Stop, and report the unknown state honestly, when the retries are exhausted.

Level and prerequisites. L2 — operational: the reader must be able to build a client that streams, times out and retries correctly. No client is implemented here and no library is prescribed; the retry policy of a specific pipeline belongs to Progetti. Prerequisites: the previous sheet on error classification, and the measurement sheet for the latency vocabulary.

Where to go next

References

  • IETF — RFC 9110, §9.2.2 — the definition of idempotency, the client rule for non-idempotent methods and the proxy rule.
  • Ollama — API reference — streaming as the default for certain endpoints and the stream: false opt-out.
  • ggml-org — llama.cpp server — the SSE ping interval and the server read/write timeout.
  • AWS — Timeouts, retries, and backoff with jitter — the timeout rule, both timers, the cost of choosing badly, and why retries need backoff.
  • Microsoft — Retry pattern — transient faults, increasing delays, the harmful aggressive policy, and the circuit breaker.
  • Anthropic — Errors — the SDK’s own retry behaviour and the retry-after header.
  • OpenAI — API description — the Idempotency-Key header and what it makes idempotent.