On a NUMA machine, reserving virtual address space does not necessarily choose the physical NUMA node that will supply every page. For anonymous memory, physical allocation commonly happens later, when a CPU first faults on a page. Under the default local allocation policy, that fault can make initialization order part of the memory-placement decision.

This behavior is often called first-touch placement. The thread that first writes or otherwise faults a page can cause Linux to allocate backing memory near the NUMA node on which that thread is running, subject to the active memory policy, cpuset constraints, available memory, and kernel fallback behavior.

The practical effect is subtle: two programs can reserve the same amount of memory and execute the same parallel computation yet begin with different locality if their pages were faulted by different CPUs.

Virtual reservation and physical placement are separate events

A process can obtain a large anonymous mapping without immediately receiving a distinct physical page for every virtual page in the range. Linux can defer physical allocation until access triggers a page fault.

At that fault, the kernel has concrete context: the task is running on a CPU, that CPU belongs to a NUMA node, and the process is subject to a memory policy and resource constraints. With the default policy, the kernel normally prefers memory local to the node associated with the faulting CPU.

This separates address-space construction from page placement. A call that creates a mapping establishes a virtual range, while later faults can determine where individual backing pages come from.

The distinction matters for large arrays. A single mapping can ultimately contain pages from several NUMA nodes because different worker threads faulted different portions while running on different CPUs.

Serial initialization can concentrate pages on one node

Consider a program that allocates a large array and initializes every element from one thread before starting worker threads. If that initialization faults the array’s pages while the thread remains on one NUMA node, many of those pages can be allocated from that node.

Later, workers scheduled on other nodes may repeatedly access those pages through remote memory paths. The parallel phase can therefore have balanced CPU work but asymmetric memory placement.

The issue is not that the worker threads are incapable of addressing remote memory. NUMA systems provide a coherent address space for ordinary CPU memory access. The cost difference comes from topology: an access to memory attached to another node can traverse additional interconnect resources and can have different latency and bandwidth characteristics from local access.

Exact costs depend on the machine, firmware topology, processor generation, workload, and traffic. First-touch placement does not imply a fixed remote-access penalty.

Parallel initialization can distribute placement

A common placement technique is to initialize each partition from the worker that will later process that partition. If workers are pinned or otherwise remain on intended NUMA nodes, their first faults can allocate backing pages near those CPUs under the default policy.

For example, a large array divided into contiguous regions can be initialized in parallel, with each worker touching its own region. The resulting page distribution can follow the workers’ NUMA placement rather than the location of a single initializer.

This is a placement consequence, not a universal optimization rule. A workload that later shares every page uniformly across nodes may need a different policy. Interleaving pages across nodes can be more suitable for some shared-access patterns, while explicit binding can be useful when ownership is stable.

The important point is that initialization is not merely a way to fill bytes. On demand-paged NUMA memory, it can also be the event that establishes physical locality.

CPU affinity can make placement more predictable

First-touch behavior depends on where the faulting thread actually executes. If the scheduler moves an initializing thread between NUMA nodes while it faults pages, the resulting placement can be spread across those nodes.

CPU affinity can reduce that variability by constraining workers to selected CPUs. NUMA-aware applications often combine CPU placement with a deliberate memory initialization strategy so that the CPU producing the first faults is also close to the pages it will use.

Affinity alone does not force a specific memory policy. It controls eligible CPUs, while memory policy controls allocation preferences or requirements. The two mechanisms interact because the default local policy uses the faulting CPU’s node as an allocation preference.

Container and cpuset configurations can further restrict which CPUs and memory nodes are available. An application’s apparent CPU placement therefore needs to be interpreted together with those constraints.

Memory policy can override the default locality preference

First-touch is most useful as a description of behavior under ordinary local allocation. Linux also exposes NUMA memory policies that can request other placement strategies.

A process or mapping can be bound to selected memory nodes, can prefer a node, or can use an interleaving policy. Tools and APIs built around Linux NUMA facilities can set such policies without relying solely on the node of the first fault.

These policies change the allocation decision made at fault time. A page first accessed on one CPU does not necessarily come from that CPU’s node if an explicit policy directs allocation elsewhere.

Resource pressure also matters. Even with a local preference, the kernel can face conditions in which the preferred node cannot satisfy an allocation under the applicable policy and constraints. Placement must therefore be observed rather than inferred from source code alone when locality is operationally important.

Page size changes the granularity of placement

Placement happens for physical pages, not individual application objects. With ordinary base pages, nearby bytes that share a page also share its NUMA location.

Larger pages increase that granularity. Transparent Huge Pages or explicit huge-page use can cause a much larger virtual region to be backed as one large page, subject to kernel behavior and configuration. The placement implications then apply at that larger unit.

This can matter when multiple workers operate on adjacent regions. An application-level partition boundary that falls inside one physical page cannot give each side an independent NUMA location for that page.

Page layout, alignment, huge-page policy, and access partitioning therefore affect how closely first-touch behavior can match logical ownership.

Placement can change after the initial fault

First touch describes initial allocation; it is not a permanent ownership label. Linux can migrate pages under explicit NUMA operations, and automatic NUMA balancing can sample access patterns and move tasks or pages in an effort to improve locality when that facility is active.

A snapshot taken long after startup may therefore differ from the initial distribution. Thread migration, page migration, memory pressure, policy changes, and balancing activity can all affect the observed state.

This also means startup placement and steady-state placement are separate questions. Deliberate parallel initialization can produce a useful starting layout, but long-running behavior still depends on where computation runs and which pages each CPU accesses.

Locality is a property to measure

NUMA placement is visible system state. Linux exposes node-level memory information, process NUMA maps, and NUMA-related counters that can help correlate page placement with CPU execution. Performance tooling can add evidence about remote accesses and memory behavior where the hardware and kernel expose suitable events.

The useful diagnostic question is not simply whether a program uses NUMA. On a multi-node system, the more precise questions are where its threads execute, where its physical pages reside, which threads access those pages, and whether those relationships remain stable over time.

First-touch placement connects the first two stages of that story. A page fault occurs on a particular CPU, and under the default local policy that location influences the preferred source of physical memory. Initialization order can therefore become part of a program’s NUMA topology before its main computation begins.