Articles

HDD, SSD and NVMe: media, interfaces and the three numbers of storage performance

IOPS, throughput and latency: how the medium and the interface decide which one binds.

Reading: 6 minServer & Virtualization

Article cover: HDD, SSD and NVMe: media, interfaces and the three numbers of storage performance

“One terabyte” says how much data a device holds, not how fast it serves it. A 7.2k hard disk and an NVMe drive of the same capacity can move the same number of bytes per second in a large sequential copy and differ by two orders of magnitude on the same workload sliced into small random operations. Storage performance is described by three numbers — IOPS, throughput and latency — and both the medium and the interface in front of it decide which of the three you will actually hit.

The model: request, queue, medium

application read() / write()
        ↓
filesystem  →  block layer
        ↓
interface queue(s)        SATA: one command queue
        ↓                 NVMe: many submission/completion queue pairs
medium                    magnetic platters  |  NAND flash
        ↓
 IOPS  ·  throughput  ·  latency

Every storage figure is a property of this whole path. A number measured on the medium cannot be assumed to hold at the application, because the FS, the queue and the controller sit in between.

Terms used here

  • IOPS — I/O operations per second; the SNIA dictionary gives the term as shorthand for exactly that.
  • Throughput — bytes per second moved by those operations; the same IOPS at a larger block size is a larger throughput.
  • Latency — how long one operation takes to complete, including the time it waits in a queue.
  • Block size — how much data one operation carries; 4 KiB is the size most random-write workloads are described with.
  • Random vs sequential — whether successive operations address neighbouring data or scattered data.
  • Namespace (NVMe) — “a set of resources (e.g., formatted non-volatile storage) that may be accessed by a host”; the unit of storage NVMe exposes to the host.

The medium: mechanics first

A hard disk is mechanical. To read data the head must be positioned (a seek) and then the platter must rotate until the target data arrives under it; that second part is rotational latency, which the SNIA dictionary defines as “the time between the completion of a seek and the instant of arrival of the first” bit. Because the platters spin at a fixed rate, sequential access — where one seek buys many adjacent blocks — is what HDDs are good at, and random access pays the mechanical cost again and again.

NAND flash has no moving parts, so there is no seek and no rotational latency. Its cost model is different: data is written in pages that live inside erase blocks, and updating data in place requires erasing and rewriting more than the caller asked for. The ratio between what the device writes internally and what the host asked it to write is write amplification; Western Digital’s white paper on SSD endurance defines it exactly as “the ratio between the two”, notes that many factors influence it, and observes that large-block sequential workloads typically have low write amplification while “small-block random write workloads typically have” higher values. That is why an SSD’s endurance — how much writing it survives — depends on the workload, not only on the drive.

The interface: how many operations can be in flight

NVMe is designed around parallelism rather than around a single serialised command path. The NVMe Base Specification defines submission queues and completion queues, and its own text describes a queue pair per core “to avoid locking”, with one-to-one and n-to-one mappings between submission and completion queues. The point of the design is parallelism: the specification describes support for parallel operation “by supporting up to 65,535 I/O” submission queues, so the device can hold many commands in flight and hide the latency of each one. A path that can hold only a few cannot, whatever the media underneath it.

The three numbers, and how they interlock

Number What it answers Depends on
IOPS How many operations per second complete block size, randomness, queue depth, medium
Throughput How many bytes per second move IOPS × block size, interface width
Latency How long one operation takes medium mechanics, queueing, cache

The relationship is arithmetic and worth stating once: a device doing 500 random 4 KiB reads per second with a queue of one and a device doing 500,000 of them are both “doing IOPS” — the difference is where the wait happens. A workload with deep queues can reach high IOPS on a device with high per-operation latency; a workload with one outstanding request per thread cannot.

A common misconception: the spec sheet decides

Vendor numbers are measured under stated conditions (block size, queue depth, read/write mix, and usually sequential for throughput). A workload with a different mix, or a device placed behind a controller, a host bus or a hypervisor queue, will not reproduce them. The numbers that matter are the ones measured on the path the workload actually takes; the spec sheet tells you the ceiling of the medium, not the ceiling of your system.

What to remember

  • IOPS, throughput and latency are three views of the same path; a claim about one without its conditions says very little.
  • HDD performance is mechanical: seek plus rotational latency, favourable to sequential access.
  • SSD performance is electrical but writes cost more than they appear to, which is write amplification, and it is workload-dependent.
  • NVMe’s advantage is parallelism — many submission and completion queues, one per core is the design intent — not a different kind of flash.
  • Measure on the whole path: filesystem, queue, controller and hypervisor are part of the number.

Level and prerequisites

L1 — fundamentals: vocabulary and mechanism. Prerequisites: none. Measuring storage (fio, iostat), queue-depth tuning, filesystem choice and RAID are separate subjects, and RAID’s effect on these three numbers is treated in its own L1 sheet.

Where to go next

References

  • NVM Express — NVM Express Base Specification, Revision 2.3 (ratified 1 August 2025): submission and completion queues, queue-pair mappings, parallel operation with up to 65,535 I/O submission queues, and the namespace definition.
  • SNIA — SNIA Dictionary (storage networking industry terminology): IOPS, rotational latency, write amplification factor, RAID.
  • Western Digital — White Paper: SSD Endurance and HDD Workloads: write amplification definition, workload dependence, enterprise random-write mix.