The kernel's nf_tables netfilter subsystem has been a recurring source of local privilege escalations. It is reachable from a container that can create a user namespace and gain CAP_NET_ADMIN within it, because that unlocks the netfilter configuration API. A memory-corruption bug there yields kernel read and write, and so host code execution.
# Reachability check: can the container create a userns and get CAP_NET_ADMIN?
unshare -Urn sh -c 'nft list ruleset' 2>&1 | head -1
Exploitation notes#
- The key enabler is unprivileged user namespaces being allowed on the host (
kernel.unprivileged_userns_clone), which grantsCAP_NET_ADMINover the new netns. - Where user namespaces are restricted, this surface is largely closed from an unprivileged container.
- Exploits are version-specific kernel memory corruption; confirm the host kernel and the user-namespace policy first.