Mapping uid 0 of the parent into a user namespace requires CAP_SETFCAP, because a task in such a namespace can write file capabilities that are honoured in the parent. For a writer that already sits in the new namespace nothing can be read from its capability sets any more, so create_user_ns() records in ns->parent_could_setfcap whether the creator had the capability and verify_root_map() trusts that. The flag describes the creator, but it is applied to whoever opens /proc/self/uid_map from inside the namespace, and setns() lets other tasks in: task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP unshare(CLONE_NEWUSER) ns->parent_could_setfcap = 1 setns(A's user ns) allowed, euid == ns->owner write "0 0 1" to /proc/self/uid_map map_ns == file_ns and parent_could_setfcap -> allowed B could not have written that map from the parent namespace, nor into a namespace it unshared itself. The check for an opener in the parent namespace has the same weakness one level up. If B enters a namespace that already maps uid 0, it has CAP_SETFCAP there, may unshare again and map uid 0 once more, and the kuid behind that is still the root user of the initial namespace. Use cred->setfcap_level, which tells what the opener itself was allowed to do before it entered its namespace: CAP_SETFCAP has to reach up to the topmost namespace that has the same root user as the parent of the namespace being mapped. The old conditions stay, so nothing becomes allowed that was refused before. The change in behaviour is that a task which entered a namespace while lacking CAP_SETFCAP outside gets -EPERM when it tries to pass on uid 0 of the outside. The creator of a namespace, its children, tasks that join with the capability, helpers in the parent namespace and unprivileged users nesting namespaces below one of their own are not affected. Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0") Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/capability.h | 1 + kernel/user_namespace.c | 31 ++++++++++++++++++------------- security/commoncap.c | 23 +++++++++++++++++++++++ 3 files changed, 42 insertions(+), 13 deletions(-) diff --git a/include/linux/capability.h b/include/linux/capability.h index 026974a5111b..90f9f976d216 100644 --- a/include/linux/capability.h +++ b/include/linux/capability.h @@ -224,5 +224,6 @@ int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry, const void **ivalue, size_t size); int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns); +int cap_root_level(kuid_t kuid, struct user_namespace *ns); #endif /* !_LINUX_CAPABILITY_H */ diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c index 6b45df3a8d82..bd5f9cea7430 100644 --- a/kernel/user_namespace.c +++ b/kernel/user_namespace.c @@ -900,7 +900,7 @@ static bool verify_root_map(const struct file *file, struct user_namespace *map_ns, struct uid_gid_map *new_map) { - int idx; + int idx, level; const struct user_namespace *file_ns = file->f_cred->user_ns; struct uid_gid_extent *extent0 = NULL; @@ -918,24 +918,29 @@ static bool verify_root_map(const struct file *file, if (!extent0) return true; + /* The parent may in turn share its root user with its ancestors. */ + level = cap_root_level(make_kuid(map_ns->parent, 0), map_ns->parent); + if (map_ns == file_ns) { - /* The process unshared its ns and is writing to its own + /* The process is in the new ns and is writing to its own * /proc/self/uid_map. User already has full capabilites in - * the new namespace. Verify that the parent had CAP_SETFCAP - * when it unshared. - * */ + * the new namespace. Verify that the creator had CAP_SETFCAP + * when it unshared, and that the opener, which may have come + * in later with setns(), had it as well when it entered. + */ if (!file_ns->parent_could_setfcap) return false; - } else { - /* Process p1 is writing to uid_map of p2, who is in a child - * user namespace to p1's. Verify that the opener of the map - * file has CAP_SETFCAP against the parent of the new map - * namespace */ - if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP)) - return false; + return file->f_cred->setfcap_level <= level; } - return true; + /* Process p1 is writing to uid_map of p2, who is in a child + * user namespace to p1's. Verify that the opener of the map + * file has CAP_SETFCAP against the parent of the new map + * namespace, and not just because it entered that. + */ + if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP)) + return false; + return cap_setfcap_level(file->f_cred, map_ns->parent) <= level; } static ssize_t map_write(struct file *file, const char __user *buf, diff --git a/security/commoncap.c b/security/commoncap.c index 513473f859be..b406ede2fadc 100644 --- a/security/commoncap.c +++ b/security/commoncap.c @@ -157,6 +157,29 @@ int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns) return cred->user_ns->level + 1; } +/** + * cap_root_level - Find the topmost namespace in which a kuid is the root user + * @kuid: The kuid to look for + * @ns: The user namespace to start from + * + * Return: the lowest ->level among @ns and its ancestors in which @kuid is + * uid 0, which is how far up file capabilities with that root user are + * honoured; INT_MAX if there is no such namespace. + */ +int cap_root_level(kuid_t kuid, struct user_namespace *ns) +{ + int level = INT_MAX; + + for (;; ns = ns->parent) { + if (from_kuid(ns, kuid) == 0) + level = ns->level; + if (ns == &init_user_ns) + break; + } + + return level; +} + /** * cap_settime - Determine whether the current process may set the system clock * @ts: The time to set -- 2.55.0