Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,425 words · 1 segments analyzed
1 "Linux User Namespaces Might Not Be Secure Enough" by Erica Windisch: If a (real) root user has had the SYS_CAP_ADMIN capability removed, but then creates a user namespace, this capability is restored for the (fake) root user. That is, before creating the namespace, ‘mount’ would be denied, but following the creation of the user namespace, the ‘mount’ syscall would magically work again, albeit in a limited fashion. While limited in function, it’s significant enough that given a (real) root user and a kernel with user namespaces, Linux capabilities may be completely subverted. and man 7 user_namespaces says: The child process created by clone(2) with the CLONE_NEWUSER flag starts out with a complete set of capabilities in the new user namespace. and "Understanding and Hardening Linux Containers" again User namespaces also allows for ``interesting'' intersections of security models, whereas full root capabilities are granted to new namespace. This can allow CLONE_NEWUSER to effectively use CAP_NET_ADMIN over other network namespaces as they are exposed, and if containers are not in use. Additionally, as we have seen many times, processes with CAP_NET_ADMIN have a large attack surface and have resulted in a number of different kernel vulnerabilities. This may allow an unprivileged user namespace to target a large attack surface (the kernel networking subsystem) whereas a privileged container with reduced capabilities would not have such permissions. See Section 5.5 on page 39 for a more in-depth discussion on this topic. We can demonstrate this behavior (on a host with user namespaces compiled in) with Listing 1: subverting_networking.c/* Local Variables: */ /* compile-command: "gcc -Wall -Werror -static subverting_networking.c \*/ /* -o subverting_networking" */ /* End: */ #define _GNU_SOURCE #include <stdio.h> #include <unistd.h> #include <sched.h> #include <sys/ioctl.h> #include <sys/socket.h> #include <linux/sockios.h> int main (int argc, char **argv) { if (unshare(CLONE_NEWUSER | CLONE_NEWNET)) { fprintf(stderr, "++ unshare failed: %m\n"); return 1; } /* this is how you create a bridge... */ int sock = 0; if ((sock = socket(PF_LOCAL, SOCK_STREAM, 0)) == -1) { fprintf(stderr, "++ socket failed: %m\n"); return 1; } if (ioctl(sock, SIOCBRADDBR, "br0")) { fprintf(stderr, "++ ioctl failed: %m\n"); close(sock); return 1; } close(sock); fprintf(stderr, "++ success!\n"); return 0; } alpine-kernel-dev:~$ whoami lizzie alpine-kernel-dev:~$ ./subverting_networking ++ success! alpine-kernel-dev:~$ but we're not actually that powerful. Listing 2: subverting_setfcap.c/* Local Variables: */ /* compile-command: "gcc -Wall -Werror -lcap -static subverting_setfcap.c \*/ /* -o subverting_setfcap" */ /* End: */ #define _GNU_SOURCE #include <stdio.h> #include <sched.h> #include <linux/capability.h> #include <sys/capability.h> int main (int argc, char **argv) { if (unshare(CLONE_NEWUSER)) { fprintf(stderr, "++ unshare failed: %m\n"); return 1; } cap_t cap = cap_from_text("cap_net_admin+ep"); if (cap_set_file("example", cap)) { fprintf(stderr, "++ cap_set_file failed: %m\n"); cap_free(cap); return 1; } cap_free(cap); return 0; } alpine-kernel-dev:~$ whoami lizzie alpine-kernel-dev:~$ touch example alpine-kernel-dev:~$ ./subverting_setfcap ++ cap_set_file failed: Operation not permitted 2 init/Kconfig:1207@c8d2bc config USER_NS bool "User namespace" default n help This allows containers, i.e. vservers, to use user namespaces to provide different user info for different servers. When user namespaces are enabled in the kernel it is recommended that the MEMCG option also be enabled and that user-space use the memory control groups to limit the amount of memory a memory unprivileged users can use. If unsure, say N. 3 Ubuntu switches CONFIG_USER_NS on, but patches it so that it unprivileged use can be disabled with a sysctl, unpriviliged_userns_clone. Listing 3: 92e575e769cc50a9bfb50fb58fe94aab4f2a2bffcommit 92e575e769cc50a9bfb50fb58fe94aab4f2a2bff Author: Serge Hallyn <redacted> Date: Tue Jan 5 20:12:21 2016 +0000 UBUNTU: SAUCE: add a sysctl to disable unprivileged user namespace unsharing It is turned on by default, but can be turned off if admins prefer or, more importantly, if a security vulnerability is found. The intent is to use this as mitigation so long as Ubuntu is on the cutting edge of enablement for things like unprivileged filesystem mounting. (This patch is tweaked from the one currently still in Debian sid, which in turn came from the patch we had in saucy) Signed-off-by: Serge Hallyn <redacted> [bwh: Remove unneeded binary sysctl bits] Signed-off-by: Tim Gardner <redacted> Debian has the same behavior: Listing 4: debian/patches/debian/add-sysctl-to-allow-unprivileged-CLONE_NEWUSER-by-default.patchFrom: Serge Hallyn <redacted> Date: Fri, 31 May 2013 19:12:12 +0000 (+0100) Subject: add sysctl to disallow unprivileged CLONE_NEWUSER by default Origin: http://kernel.ubuntu.com/git?p=serge%2Fubuntu-saucy.git;a=commit;h=5c847404dcb2e3195ad0057877e1422ae90892b8 add sysctl to disallow unprivileged CLONE_NEWUSER by default This is a short-term patch. Unprivileged use of CLONE_NEWUSER is certainly an intended feature of user namespaces. However for at least saucy we want to make sure that, if any security issues are found, we have a fail-safe. Signed-off-by: Serge Hallyn <redacted> [bwh: Remove unneeded binary sysctl bits] --- Grsecurity disables it entirely for users without CAP_SYS_ADMIN, CAP_SETUID, and CAP_SETGID. Listing 5: https://grsecurity.net/test/grsecurity-3.1-4.7.9-201610200819.patch--- a/kernel/user_namespace.c +++ b/kernel/user_namespace.c @@ -84,6 +84,21 @@ int create_user_ns(struct cred *new) !kgid_has_mapping(parent_ns, group)) return -EPERM; +#ifdef CONFIG_GRKERNSEC + /* + * This doesn't really inspire confidence: + * http://marc.info/?l=linux-kernel&m=135543612731939&w=2 + * http://marc.info/?l=linux-kernel&m=135545831607095&w=2 + * Increases kernel attack surface in areas developers + * previously cared little about ("low importance due + * to requiring "root" capability") + * To be removed when this code receives *proper* review + */ + if (!capable(CAP_SYS_ADMIN) || !capable(CAP_SETUID) || + !capable(CAP_SETGID)) + return -EPERM; +#endif and Arch Linux has it off. Listing 6: {linux} 3.13 add CONFIG_USER_NSComment by William Kennington (Webhostbudd) - Sunday, 06 October 2013, 03:55 GMT I agree with Florian, allowing non-root users to take advantage of elevating themselves to a local root seems like a huge attack surface. Preferably this would be a sysctl with a huge warning attached to it when it is switched on. Comment by Daniel Micay (thestinger) - Monday, 24 November 2014, 03:55 GMT [...] Arch doesn't add new features via patches. If you want to see this feature enabled, then land something like this upstream. Note that CONFIG_USER_NS is already enabled in the linux-grsec package because it fully removes the ability to have unprivileged user namespaces. It would have been cool to include Red Hat's patches here, but I couldn't find them. 4 Most of this section is cribbed from the example at the bottom of man 2 clone. 5 Listing 11: clone_stack.c/* -*- compile-command: "gcc -Wall -Werror clone_stack.c -o clone_stack" -*- */ #define _GNU_SOURCE #include <sched.h> #include <sys/wait.h> #include <stdio.h> #include <stdlib.h> #include <unistd.h> #define STACK_SIZE (1024 * 1024) int child (void *_) { int stack_value = 0; fprintf(stderr, "pre-execve, stack is ~%p\n", &stack_value); execve("./show_stack", (char *[]) {",/show_stack", 0}, NULL); return 0; } int main (int argc, char **argv) { void *stack = malloc(STACK_SIZE); clone(child, stack + STACK_SIZE, SIGCHLD, NULL); wait(NULL); return 0; } Listing 12: show_stack.c/* -*- compile-command: "gcc -Wall -Werror -static show_stack.c -o show_stack" -*- */ #include <stdio.h> int main (int argc, char **argv) { int stack_value = 0; fprintf(stderr, "post-execve, stack is ~%p\n", &stack_value); return 0; } [lizzie@empress linux-containers-in-500-loc]$ ./clone_stack pre-execve, stack is ~0x7f3f98deefec post-execve, stack is ~0x7ffd14d2291c The stack grows down on x86, so the fact that the address is higher numerically post-execve means that a new stack has been allocated. 6 I thought this might be undefined behavior, since stack + STACK_SIZE does point past the last item of the array, but point 8 of 6.5.6 [Additive operators] in ISO-9899 has us covered: If both the pointer operand and the result point to elements of the same array object, or one past the last element of the array object, the evaluation shall not produce an overflow; otherwise, the behavior is undefined. If the result points one past the last element of the array object, it shall not be used as the operand of a unary * operator that is evaluated. i.e., the pointer addition is valid, but dereferencing it wouldn't be. 7 I wasn't confident that waitpid was enough to wait for the process and all of its children, but when the root of a pid namespace closes, all of its children get SIGKILL: man 7 pid_namespaces: If the "init" process of a PID namespace terminates, the kernel terminates all of the processes in the namespace via a SIGKILL signal. This behavior reflects the fact that the "init" process is essential for the correct operation of a PID namespace. Also verified this myself, before I found that: Listing 18: persistent_child.c/* -*- compile-command: "gcc -Wall -Werror -static persistent_child.c -o persistent_child" -*- */ #include <unistd.h> #include <stdio.h> #include <sys/types.h> #include <sys/stat.h> #include <fcntl.h> int main (int argc, char **argv) { switch (fork()) { case -1: fprintf(stderr, "++ fork failed: %m\n"); return 1; case 0:; int fd = 0; if ((fd = open("persistent_child.log", O_CREAT | O_APPEND | O_WRONLY, S_IRUSR | S_IWUSR)) == -1) { fprintf(stderr, "++ open failed: %m\n"); return 1; } size_t count = 0; while (count < 100) { if (dprintf(fd, "%lu\n", count++) < 0) { fprintf(stderr, "++