Articles

Model serving and compatible APIs: one endpoint, several contracts

A serving engine adds a queue, a lifecycle and an API surface to a weights file.

Reading: 6 minAutomation & AI

Article cover: Model serving and compatible APIs: one endpoint, several contracts

A weights file cannot answer a request. Something has to hold the model in memory, decide how many requests to run at once, queue the rest, and speak HTTP — that something is the serving engine, and it is where the phrase “OpenAI-compatible” does its work. The habit worth building is this: treat compatibility as a claim about shape, verify it against the engine’s own documentation, and remember that the same engine usually offers more than one contract at the same time.

What serving adds to a model

A model file plus a server is a different kind of object from a model file:

  • residency — the model stays loaded between requests; in Ollama the keep_alive parameter controls how long it stays in memory after the last request.
  • a queue and a batching strategy — server slots (-np, --parallel) and continuous batching (-cb, --cont-batching, enabled by default), which is what turns several requests into one batch.
  • a lifecycle — which model is loaded, when it is loaded, and what happens to the first request that arrives before it is ready.
  • an HTTP surface — endpoints, authentication, streaming, error shapes and, optionally, metrics.

The surface a local engine actually exposes

The llama.cpp server is a useful concrete case because its documentation lists every route. Its defaults are deliberately conservative: --host defaults to 127.0.0.1 and --port to 8080.

A health endpoint that needs no key. GET /health is documented as “public (no API key check)” with /v1/health as an alias, and it reports 503 while the model is not ready — which is precisely what a load balancer or a start-up script wants to poll.

OpenAI-compatible routes. GET /v1/models (“OpenAI-compatible Model Info API”), POST /v1/completions, POST /v1/chat/completions, POST /v1/embeddings and POST /v1/responses are each documented as OpenAI-compatible APIs.

A second contract. POST /v1/messages is documented as an “Anthropic-compatible Messages API” — the same server, the same loaded model, a different request and response shape.

Native routes. Alongside them the server exposes its own endpoints: /completion, /tokenize, /detokenize, /apply-template, /infill, /props, /slots, /lora-adapters. They are richer than the compatible ones — and, being engine-specific, they are not portable.

Observability and security. GET /metrics is a “Prometheus compatible metrics exporter”, disabled until --metrics is passed; authentication is off until you pass --api-key (a comma-separated list) or --api-key-file; TLS is optional (--ssl-key-file, --ssl-cert-file).

What “compatible” does and does not mean

It means the route and the JSON shape match closely enough that a client written for the OpenAI API works unchanged. The engine’s own documentation marks the seam explicitly: some endpoints are “not OAI-compatible” — /completion and /embeddings are described that way, next to their OpenAI-compatible counterparts /v1/completions and /v1/embeddings, which exist precisely because the response formats differ.

What the shape does not cover: default values for parameters the caller omits, fields the local engine adds to a response, the exact body and status of an error, rate-limit headers (a local engine has none in the provider’s sense — see the API sheet in this batch), model naming and aliases, and whether a parameter is implemented at all. Two engines can both be “compatible” and disagree about every one of those.

The consequence of running one

A local server is an HTTP service that, by default, listens only on loopback and accepts any caller. Making it reachable from the network is a second decision, and it is separate from the first: --host 0.0.0.0 publishes the port, while --api-key is what decides who may use it. Publishing without a key exposes the machine’s accelerator, its loaded model and its logs to anyone who can reach the port. The /health endpoint stays public by design, which is reasonable because it discloses only readiness — and it is worth knowing when reading an engine’s threat surface.

Level and prerequisites. L2 — operational: the reader must be able to bring a model up behind an HTTP interface, name what the serving layer adds, and check a compatibility claim instead of trusting it. No server is installed or configured in this sheet, and no deployment is described. Prerequisites: the memory and placement sheets of this batch.

Where to go next

References

  • ggml-org — llama.cpp server — the default host and port, the full endpoint list with each route’s own description, the API-key options, TLS options and the metrics endpoint.
  • Ollama — API reference — a second local engine’s surface: POST /api/generate and POST /api/chat on localhost:11434, with streaming by default.
  • vLLM — documentation — a serving engine described as exposing an “OpenAI-compatible API server, plus Anthropic Messages API and gRPC support”.