A pathname passed to openat() is resolved by the kernel, but the caller has limited control over traversal through intermediate components. Linux openat2() adds a resolve field that applies constraints to the complete lookup operation. The restriction is evaluated while components are traversed, rather than by validating a pathname in userspace and opening it later.
That distinction matters when path components can change concurrently. A userspace sequence that checks a path and then opens it creates separate observations of mutable filesystem state. openat2() places the selected lookup policy in the same system call that returns the file descriptor.
RESOLVE_BENEATH keeps traversal below dirfd
RESOLVE_BENEATH requires successful resolution to remain below the directory referenced by dirfd. An absolute input path is rejected, as is an absolute symbolic link encountered during traversal.
struct open_how how = {
.flags = O_RDONLY | O_CLOEXEC,
.resolve = RESOLVE_BENEATH | RESOLVE_NO_MAGICLINKS,
};
int fd = syscall(SYS_openat2, rootfd, path, &how, sizeof(how));A relative component such as .. is not intrinsically forbidden. It becomes a problem when resolution would escape the permitted subtree. The kernel tracks the lookup relative to the supplied directory and rejects an escape with EXDEV.
The interface can also return EAGAIN if the kernel cannot prove that a .. traversal remained safe because of a concurrent rename or related race. A caller may retry the operation. This is a runtime containment check, not lexical normalization of the input string.
RESOLVE_IN_ROOT changes the lookup root for one operation
RESOLVE_IN_ROOT gives dirfd root-like semantics for this lookup. Absolute paths are interpreted relative to that directory, absolute symbolic links are also resolved relative to it, and .. at the boundary remains at the boundary.
process root: /
dirfd: /srv/image
path: /etc/app.conf
RESOLVE_IN_ROOT lookup target:
/srv/image/etc/app.confThe effect resembles temporarily changing the root used for pathname resolution, but it is scoped to the individual openat2() call. It does not modify the process root directory or affect concurrent threads.
RESOLVE_BENEATH and RESOLVE_IN_ROOT therefore express different policies. The first rejects resolution that leaves the directory and rejects absolute paths. The second treats the directory as the root of the lookup and gives absolute paths a meaning inside that root.
Symlink policy can cover every component
O_NOFOLLOW applies to the final pathname component. It does not prevent an intermediate component from being a symbolic link. RESOLVE_NO_SYMLINKS is broader: it disallows symbolic-link resolution across all components and also implies RESOLVE_NO_MAGICLINKS.
Magic links are a distinct Linux mechanism associated notably with procfs entries such as /proc/pid/fd/*. RESOLVE_NO_MAGICLINKS blocks their resolution without banning ordinary symbolic links.
The current kernel behavior of RESOLVE_BENEATH and RESOLVE_IN_ROOT also disables magic-link resolution, but the documented interface does not make that an enduring property of those flags. A policy that specifically requires magic links to stay unresolved should set RESOLVE_NO_MAGICLINKS explicitly.
This separation lets callers state the actual boundary. A service may permit ordinary symlinks inside a controlled tree while excluding procfs-style magic links, or it may prohibit symlink traversal entirely.
Mount crossing is a separate constraint
RESOLVE_NO_XDEV rejects traversal across mount points, including bind mounts. A path can remain lexically and structurally beneath dirfd while crossing into another mounted filesystem, so subtree containment and mount containment are separate properties.
When RESOLVE_NO_XDEV detects a crossing, openat2() returns EXDEV. This flag can be restrictive on systems where bind mounts or ordinary mount points are intentional parts of a directory hierarchy. Its semantics are useful when the policy itself requires lookup to stay on one mount, rather than as a universal hardening switch.
RESOLVE_CACHED turns cache state into an explicit result
RESOLVE_CACHED requires the lookup to complete from the kernel’s path cache without revalidation or filesystem I/O. If that cannot be done, the call returns EAGAIN.
The flag does not mean that the target must permanently reside in cache, nor does it alter the target’s filesystem semantics. It makes cache-only completion a condition of this particular lookup. An application can use the result to keep a fast path nonblocking and route the fallback elsewhere.
This behavior differs from containment flags. RESOLVE_BENEATH, RESOLVE_IN_ROOT, and the NO_* flags constrain where or how traversal may proceed. RESOLVE_CACHED constrains the work the kernel may perform to complete it.
open_how is an extensible ABI structure
openat2() receives an open_how structure plus its size. The structure carries ordinary open flags, creation mode, and resolution flags:
struct open_how how = {
.flags = O_RDONLY | O_CLOEXEC,
.mode = 0,
.resolve = RESOLVE_IN_ROOT | RESOLVE_NO_MAGICLINKS,
};The structure must be zero-filled so fields added by later kernels have zero values unless the caller explicitly opts into them. The size argument acts as part of the versioning contract for the extensible structure.
Unlike openat(), openat2() rejects unknown or conflicting flag values rather than silently ignoring unknown bits. Invalid values in how.flags, how.mode, or how.resolve can therefore surface as EINVAL instead of being accepted with reduced semantics.
The security boundary is the kernel lookup, not a cleaned string
Path sanitization often operates on text: removing .., rejecting an absolute prefix, or inspecting symlink targets before an open. Those checks do not freeze the namespace between inspection and use. Renames, mount changes, and symlink replacement can make a later lookup observe different state.
openat2() does not make every filesystem operation safe by itself. The caller still needs a suitable directory file descriptor, correct access controls, appropriate open flags, and a policy matching the filesystem objects it permits. The relevant property is narrower: selected path-resolution constraints are enforced by the kernel during the lookup that produces the descriptor.
That moves containment from a pre-open pathname convention into the operation that resolves the path. For code that accepts paths relative to a trusted directory, the difference is an API-level boundary rather than a string-processing rule.
References
- Linux
openat2(2): https://man7.org/linux/man-pages/man2/openat2.2.html - Linux
path_resolution(7): https://man7.org/linux/man-pages/man7/path_resolution.7.html - Linux
open_how(2type): https://man7.org/linux/man-pages/man2/open_how.2type.html