GGUF and its runtimes: one file, several engines
GGUF is a container; which engine reads it, and how, is a separate question.

GGUF is the file format that made local inference ordinary: one file, downloaded once, opened by an engine that needs nothing else from you. The specification explains why that works — it is “designed to be unambiguous by containing all the information needed to load a model” and “designed to be extensible, so that new information can be added to models without breaking compatibility”. It is also exactly why “it is a GGUF, so it will run” is not a safe assumption: GGUF is a container, and loading it is a negotiation between the container, the tensor types inside it and the engine doing the loading.
What is actually inside the file
The specification lays the file out sequentially: a header, the metadata key-value pairs it counts, one
description per tensor, and then the tensor data itself, padded to a multiple of the alignment declared in
general.alignment (which “must be a multiple of 8”). Files are little-endian by default. Two details
matter when an engine refuses a file:
- the quantization type is a property of the tensors, described per tensor in the tensor info entries;
general.file_typeis only “an enumerated value describing the type of the majority of the tensors in the file. Optional; can be inferred from the tensor types.” A file can therefore be heterogeneous. - the metadata section is what makes the file self-describing: it is where an engine finds the information it needs to interpret the tensors rather than guessing from the name.
What each runtime does with it
llama.cpp is the reference implementation of the format and the broadest demonstration of the point:
the project describes itself as providing “1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer
quantization for faster inference and reduced memory use”, across a table of backends from CUDA and Metal
to Vulkan, SYCL and WebGPU. Its quick start downloads a GGUF directly (llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF) and its server mode is described in the same breath (“Launch
OpenAI-compatible API server”).
Ollama is a local engine that manages models from its own library and serves them over HTTP; the same container is the unit it pulls and runs.
vLLM — an engine built for GPU serving — lists GGUF among the quantization formats it accepts, next to FP8, GPTQ/AWQ and compressed-tensors. The same file that a laptop engine opens can therefore be opened by a serving stack with a completely different execution strategy.
Hugging Face Transformers treats GGUF as what it is in the format’s definition — “a single-file format used to store models for inference with GGML, containing the model metadata and tensors” — and its own documentation states the trap precisely: installing the kernels matters, “otherwise the packed path falls back to full dequantization at load”.
“Compatible” has layers
- The container is read. Every engine above clears this one; GGUF is a public, versioned format.
- The tensor type is implemented, not merely recognised. The quoted fallback is the difference between the model runs quantized and the model is unpacked back to full precision at load, which silently gives back the memory the file was supposed to save.
- The architecture and tokenizer are implemented at the version you installed — the concern of the formats sheet in this batch.
- The engine exposes the parameters the workload needs — context size, sampling, chat templates, concurrency — which is the subject of the serving sheet.
What GGUF is not
It is not a training format: the specification is written for “storing models for inference with GGML”. Keeping a GGUF alongside the original checkpoint doubles the storage, so a batch of quantized variants is a decision about disk as well as about memory. And, as with any weights file, the container describes tensors and metadata — it is not a statement about who produced them.
Level and prerequisites. L2 — operational: enough to know what a GGUF file is, what can go wrong when an engine opens it, and where the layers of support sit. Prerequisites: the L1 sheets on the model/runtime boundary, plus the formats and quantization sheets of this batch. Converting a model to GGUF is a Progetti task and is not described here.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Model behaviour — the node this sheet sits in.
References
- ggml-org — GGUF specification — file structure, alignment, endianness, tensor info,
general.file_type. - ggml-org — llama.cpp README — quantization widths, backend table, direct download of a GGUF, server mode.
- Ollama — README and API reference — a local engine whose unit of work is a model it pulls and serves.
- vLLM — documentation — GGUF listed among the accepted quantization formats of a GPU serving engine.
- Hugging Face — GGUF documentation — the single-file definition and the dequantization fallback.