Articles

CPU cores, hardware threads and virtualization support: what the words mean

Cores, hardware threads and vCPUs: three layers, three different numbers.

Reading: 6 minServer & Virtualization

Article cover: CPU cores, hardware threads and virtualization support: what the words mean

A server is usually described with four numbers that look interchangeable and are not: sockets, cores, hardware threads, and vCPUs. A machine quoted as “2 sockets, 32 cores, 64 threads” is not a machine with 64 cores, and the 48 vCPUs a hypervisor is happy to create do not all run at the same time. Reading those four numbers correctly is the difference between a consolidation plan and a latency problem.

The model: what executes, what is scheduled, what is presented

physical package (socket)
  └─ core  ── executes instructions, has its own L1/L2 caches
       ├─ hardware thread  (logical processor)   ┐ share the core's
       └─ hardware thread  (logical processor)   ┘ execution resources
                       ↓  presented to an operating system as CPUs
      hypervisor (VMM)  ──schedules──→  vCPU  ──→  guest OS sees a CPU

Three layers, three different things: execution happens in a core, scheduling happens on hardware threads, and what a guest operating system sees is a virtual CPU that the hypervisor decides when to run.

Terms used here

  • Socket / package — the physical processor package on the mainboard; a two-socket server has two.
  • Core — an independent execution unit inside the package, with its own first- and second-level caches.
  • Hardware thread (also logical processor) — one of the instruction streams a core can present.
  • SMT — simultaneous multithreading: the technique that produces those extra hardware threads; Intel’s implementation is Hyper-Threading Technology.
  • vCPU — a virtual CPU: the processor the hypervisor presents to a guest and schedules onto a hardware thread.
  • VMM / hypervisor — the software that runs guests and receives the privileged operations they attempt.
  • Virtualization extension — the CPU feature the hypervisor needs to do that efficiently, such as Intel Virtualization Technology (VT-x).

Cores and hardware threads: one executes, two can be interleaved

A core is a real execution unit. With SMT, the core presents two logical processors, and each of them has its own architectural state, so it can be halted, interrupted or directed to execute a thread independently of the other — but they share the core’s execution resources. Intel’s own technical guide puts the reason for SMT plainly: applications typically use only a fraction of a processor’s execution resources at any one time, and interleaving a second thread uses more of them.

Two consequences follow, and both matter on a host:

  1. The operating system schedules the two hardware threads as if they were two CPUs, because at that level they are two schedulable entities.
  2. The second thread does not add the core’s full throughput. Work that saturates the execution units gains little from it; work that stalls on memory or cache misses gains more.

So “64 threads” is a scheduling capacity of 64, not an execution capacity of 64 cores, and any estimation that multiplies threads by single-core performance overstates the machine.

What hardware virtualization support changes

Running several operating systems on one CPU requires that each guest believe it owns the machine — including the machine’s privileged instructions. The formal requirement for that was described by Popek and Goldberg in 1974: a virtual machine monitor must be able to intercept privileged operations and present a faithful illusion. Hardware support is what makes it practical on contemporary processors.

Intel’s software developer’s manual describes the model precisely: VMX operation has two kinds, VMX root operation and VMX non-root operation. The hypervisor runs in root operation; guest software runs in non-root operation. Transitions from root to non-root are VM entries, transitions from non-root back to root are VM exits, and the processor state controlling each guest is held in a virtual-machine control data structure (VMCS), with one VMCS per virtual machine and, for a guest with several virtual processors, one per virtual processor. The manual is equally clear about the consequence: in non-root operation some instructions and events do not behave normally, and they exit to the hypervisor instead.

AMD processors implement the same idea under the name AMD-V, documented in the AMD64 Architecture Programmer’s Manual as Secure Virtual Machine (SVM), “AMD’s virtualization architecture”: the manual distinguishes host mode from guest mode and keeps guest state in a VMCB (virtual machine control block), the counterpart of Intel’s VMCS.

From hardware threads to vCPUs

A vCPU is a software object until the hypervisor gives it time. In KVM, for example, a vCPU is created by an operation on the virtual machine — KVM_CREATE_VCPU — and the kernel then runs it on a hardware thread according to its own scheduling. The vCPU is therefore best understood as a request for processor time that the host may or may not be able to satisfy immediately.

That is the origin of the whole vocabulary of “CPU overcommitment”: a hypervisor can create more vCPUs than the host has hardware threads, because most vCPUs are idle most of the time — until they are not. How many vCPUs a host can carry, and how to read “CPU ready” or “steal time” when it cannot, is operational material (L4), not a definition.

A common misconception: “threads are cores”

Counting hardware threads as cores inflates the machine by the factor SMT adds, and then the same inflation is applied a second time when thread counts are compared with vCPU counts. The arithmetic that survives contact with a real workload uses cores for execution capacity, hardware threads for scheduling capacity, and treats a vCPU as work the host has promised to find time for.

What to remember

  • A core executes; a hardware thread is one of the streams a core can present; a vCPU is scheduled.
  • SMT adds scheduling slots, not a second core: the execution resources are shared.
  • Hardware virtualization adds a second privilege dimension (root/non-root on Intel processors) and a control structure per virtual processor, so privileged guest operations exit to the hypervisor.
  • More vCPUs than hardware threads is normal and is a scheduling decision, not free capacity.
  • The three numbers are read from different layers, which is why they cannot be compared one to one.

Level and prerequisites

L1 — fundamentals: the vocabulary and the mechanism. Prerequisites: none. Inspecting a host’s topology (lscpu, /proc/cpuinfo, Task Manager), sizing vCPUs, CPU pinning and overcommitment ratios belong to L2–L4, and CPU features used by virtual machines (nested virtualization, passthrough) to L4.

Where to go next

References

  • Intel, Intel Hyper-Threading Technology Technical User’s Guide (January 2003) — two logical processors per processor, each with its own architectural state, sharing the execution resources; independent halt/interrupt/execute; typical utilisation of execution resources.
  • Intel, Intel 64 and IA-32 Architectures Software Developer’s Manual, Volume 3C — VMX root and non-root operation, VM entries and VM exits, the VMCS and its one-per-virtual-processor rule, VMCS pointer management.
  • Linux kernel documentation, KVM API — KVM_CREATE_VCPU as the operation that creates a vCPU for a virtual machine (the fetchable reference for the software side of “vCPU”).
  • AMD, AMD64 Architecture Programmer’s Manual, Volume 2: System Programming — chapter 15, Secure Virtual Machine: SVM as AMD’s virtualization architecture, host and guest mode, and the VMCB.
  • Popek and Goldberg, Formal Requirements for Virtualizable Third Generation Architectures (1974) — the definition of a virtual machine monitor and the need to intercept privileged operations.