Erica Windisch warns that Linux user namespaces might not be secure enough. If a real root user has had SYS_CAP_ADMIN removed but then creates a user namespace, the capability is restored for the fake root user. Before creating the namespace, 'mount' would be denied, but following creation, the 'mount' syscall would work again in a limited fashion. This is significant enough that given a real root user and a kernel with user namespaces, Linux capabilities may be completely subverted.
The man page for user_namespaces confirms that the child process created by clone(2) with CLONE_NEWUSER starts with a complete set of capabilities. User namespaces allow for interesting intersections of security models, granting full root capabilities to the new namespace. This can allow CLONE_NEWUSER to effectively use CAP_NET_ADMIN over other network namespaces, especially if containers are not in use. Processes with CAP_NET_ADMIN have a large attack surface and have resulted in kernel vulnerabilities, potentially allowing an unprivileged user namespace to target the kernel networking subsystem.
Demonstrations show that on a host with user namespaces compiled in, a user can create a bridge using ioctl with SIOCBRADDBR, though setting file capabilities with cap_set_file fails. Ubuntu enables CONFIG_USER_NS but patches it so that unprivileged use can be disabled with the sysctl unprivileged_userns_clone. The kernel config defaults to 'n' for user namespaces, recommending MEMCG to limit memory for unprivileged users.
Source: Hacker News · Summarized by HeadlinesBriefing