API authentication, rate limits and error handling
An endpoint refuses a request for three different reasons; the response says which.

An LLM endpoint can refuse a request for three very different reasons — who you are, how much you are asking for, and whether the service is working — and the response tells you which one it is. Treating them alike, by retrying everything, is how a client turns one throttled request into a self-inflicted outage. The classification is the skill; the retry mechanics belong to the next sheet.
Authentication: who is asking
A provider API authenticates with a bearer credential. The OpenAI API description declares its security
scheme as ApiKeyAuth of type: http and scheme: bearer, which is the standard way of saying that the
caller must send Authorization: Bearer <credential>; a second scheme covers administrative operations at
organisation level. Two failures are distinct and must be handled differently:
- 401 — not authenticated. The credential is missing, malformed, revoked or wrong; Anthropic names this
authentication_errorin its own error table. - 403 — authenticated but not permitted. The key is valid and lacks the right; Anthropic names this
permission_error.
A local engine inverts the default: authentication is something you turn on. In a llama.cpp server,
--api-key KEY is documented with “default: none” — and the health endpoint is explicitly “public (no API
key check)”. So an unauthenticated local server is not a misconfiguration that slipped through; it is what
happens when nothing is configured, and the same documentation offers --ssl-key-file and
--ssl-cert-file for TLS when the server leaves loopback. Where the credential lives is the subject of the
secrets sheet in this batch.
Rate limits and quotas: how much, and how fast
The status code for “too much, too quickly” is 429, defined in its own specification section: it
“indicates that the user has sent too many requests in a given amount of time (‘rate limiting’)”, and that
the response “representations SHOULD include details explaining the condition, and MAY include a
Retry-After header indicating how long to wait before making a new request”.
What a well-behaved client does with it is described in the provider’s own API description. Its rate-limited
response documents a Retry-After header defined as “The minimum number of seconds to wait before
retrying. This header is returned when the server has computed a retry delay and may be omitted” — with a
minimum value of one second — and explains the condition in the body of the error: a slow_down error
means traffic increased too quickly, so the client should reduce its request rate and then increase it
gradually. “Gradually” is the operative word: a client that waits the stated delay and then resumes at
exactly the rate that triggered the limit is asking for the same answer again.
Where the API documents them, per-endpoint rate-limit headers (X-RateLimit-Limit-*,
X-RateLimit-Remaining-*, X-RateLimit-Reset-*) let a client observe its budget instead of discovering it
by being refused.
429 is a message about you; a 5xx is a message about the service. Retrying a 429 immediately increases the load that caused it. A quota is not an outage.
A local engine has no quota in the provider’s sense, which does not mean it has no limits: it has a fixed
number of server slots (-np, --parallel), a read/write timeout (-to, --timeout, default 3600 seconds)
and a finite memory pool. Exceeding its capacity shows up as queueing or as a request that fails after
waiting — not as a 429 with a retry hint.
The service’s own failures
The 5xx class is defined by the HTTP specification as failures the client can usually do nothing about:
500 Internal Server Error, 502 Bad Gateway, 503 Service Unavailable and 504 Gateway Timeout. On top
of those, providers add codes of their own: Anthropic’s table lists 500 api_error and, distinctly,
529 overloaded_error — “the service is temporarily overloaded”, which in the core HTTP set has no exact
equivalent. When a provider invents a status code, the only reliable instruction is that provider’s own
documentation; a client that switches on the numeric range alone will misclassify it.
This is the class where retrying is reasonable — with backoff, a cap and jitter, which is the next sheet.
A classification you can implement
| Response | Meaning | What to do |
|---|---|---|
| 400, 422 | the request is malformed | fix the request; never retry unchanged |
| 401 | the credential is missing, invalid or revoked | refresh or replace it; stop |
| 403 | the credential lacks the right | escalate; retrying changes nothing |
| 404 | the route or the model identifier is wrong | fix the reference |
| 413 | the payload is too large | shorten the input |
| 429 | rate limit or quota | honour Retry-After, then slow the rate |
| 500, 502, 503, 504 | the service failed | retry with backoff, capped |
| 529 (provider-specific) | the service is overloaded | as 5xx, per that provider’s documentation |
| timeout, no response | unknown state | the next sheet: idempotency decides whether a retry is safe |
The last row is the one that matters most and the one most often ignored: a request that may have been executed is not the same as a request that failed.
The operational cost of getting this wrong
Retry logic is where a client’s behaviour becomes part of the service’s problem. OWASP names the class LLM10 Unbounded Consumption: excessive and uncontrolled inference, whether by design or by a client that multiplies its own traffic, producing denial of service, economic loss and service degradation. A retry policy is an amplifier; its controls — classification, backoff, cap, jitter, and a circuit breaker when a dependency is clearly down — are the subject of the next sheet.
Level and prerequisites. L2 — operational: the reader must be able to classify a refusal and choose a response policy for each class. Prerequisites: the serving sheet of this batch (what an endpoint is) and the secrets sheet for where the credential lives.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Applications — the node this sheet sits in.
References
- IETF — RFC 6585, §4 — the definition of 429 and the
Retry-Afterguidance. - IETF — RFC 9110 — the definitions of 500, 502, 503 and 504, and the retry rule for non-idempotent methods.
- OpenAI — API description — the bearer security scheme, the rate-limited response with its
Retry-Afterheader and itsslow_downcondition, and the per-endpoint rate-limit headers. - Anthropic — Errors — the error taxonomy including 401, 403, 413, 429, 500 and the provider-specific 529.
- OWASP — Top 10 for LLM Applications (2025), LLM10 — unbounded consumption as a risk class.