Inference on CPU and GPU: where the work actually goes
The same model runs on a CPU, a GPU or both; what changes is where the work is done.

The same GGUF file runs on a CPU with no accelerator at all, on a GPU, or on both at once. What changes is not the arithmetic the model performs — it is which device performs it and how quickly that device’s memory can feed it. Almost every misunderstanding about “CPU versus GPU” comes from skipping that second half, because the limiting factor during generation is usually memory bandwidth, not arithmetic.
Two phases with different costs
Every request has two phases, and they behave so differently that averaging them produces a meaningless number.
- Prefill processes the whole prompt at once, in parallel, and is where the prompt’s length shows up.
- Decode produces one token at a time, each conditioned on everything before it, and is where a conversation spends its wall-clock time.
Decode is not limited by how many arithmetic units a device has. “Computations in LLMs are mainly dominated by matrix-matrix multiplication operations; these operations with small dimensions are typically memory-bandwidth-bound on most hardware”, and at generation time the work is “bottlenecked on how quickly we can load model parameters from the device memory to the compute units”. The same source’s worked example makes the point arithmetic rather than intuitive: a 7-billion-parameter model at 16-bit precision with a time per output token of 14 ms is moving about 14 GB of parameters in 14 ms — a terabyte per second of bandwidth, whether or not the device can compute faster.
The practical consequence: tokens per second scale with memory bandwidth long before they scale with compute.
Where each device helps
CPU. llama.cpp describes itself as a “Plain C/C++ implementation without any dependencies” with
optimised kernels for x86 (AVX, AVX2, AVX512, AMX) and for ARM with “Apple silicon … optimized via ARM
NEON, Accelerate and Metal frameworks”. CPU inference is not a degraded mode; it is a first-class target.
Its knobs are thread counts — -t, --threads for generation and -tb, --threads-batch for prompt
processing — plus the batch sizes -b, --batch-size and -ub, --ubatch-size. Threads share one memory
bus, so adding them helps less than their number suggests, and the batch parameters shape how much working
memory the prompt phase uses.
GPU. Moving layers to the accelerator moves both the weights and the arithmetic: -ngl, --gpu-layers
(“max. number of layers to store in VRAM, either an exact number, ‘auto’, or ‘all’”). With more than one
device, -ts, --tensor-split sets the fraction offloaded to each and -sm, --split-mode {none,layer,row,tensor} chooses how the model is divided. Flash attention can be enabled explicitly
(-fa, --flash-attn, default auto), which is an attention implementation choice rather than a placement
one.
Both at once. The README states the feature plainly: “CPU+GPU hybrid inference to partially accelerate
models larger than the total VRAM capacity”. This is the important case for a reader with one modest card:
the layers that fit go to the GPU, the rest stay in system RAM, and the effect is a speed improvement on
a model that already fits in main memory — not permission to run something that does not fit anywhere. A
related case is a mixture-of-experts model, where -cmoe, --cpu-moe and -ncmoe, --n-cpu-moe N keep
expert weights on the CPU: a placement decision of a different kind, because those weights are not used by
every token.
What batching changes
A server does not serve one request at a time. llama.cpp enables continuous batching by default
(-cb, --cont-batching) and sizes the number of server slots with -np, --parallel; vLLM’s design notes
list “Continuous batching of incoming requests, chunked prefill, prefix caching” next to PagedAttention.
Continuous batching exists because decode is bandwidth-bound: several sequences can share one pass over
the weights, so the same memory traffic produces more output. That is why aggregate throughput and
per-user latency move in opposite directions — the subject of the next sheet.
What actually decides
Four things, in this order: whether the model fits where you are putting it, how long the prompt is, how long the output is, and how many requests share the device. Device class is not one of them as an independent property — “a GPU is faster” is a statement about a model, a context length and a batch size. A small model on a CPU can meet an interactive target; the same CPU with a large model will not, and the honest answer to “which is faster” is a measurement with its conditions attached.
Level and prerequisites. L2 — operational: the reader must be able to say which device holds which layer, why decode behaves differently from prefill, and what batching does to the numbers. No engine is installed, no benchmark is run and no hardware recommendation is made here. Prerequisites: the memory and quantization sheets of this batch.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- Model behaviour — the node this sheet sits in.
References
- Databricks — LLM Inference Performance Engineering — prefill versus decode, memory bandwidth as the decode bottleneck, and the 7B/14 ms worked example.
- ggml-org — llama.cpp README and llama.cpp server — the CPU kernels and backends, CPU+GPU hybrid inference, and the server parameters for threads, offloading, batching and MoE placement.
- vLLM — documentation — continuous batching, chunked prefill and PagedAttention in the engine’s own feature list.
- Hugging Face — GPU memory usage — activations as the phase-dependent part of the footprint.