File capabilities are tied to the kuid of the root user of a user namespace, and a namespace that maps uid 0 of its parent shares that kuid with the parent. So CAP_SETFCAP in such a namespace is worth as much as CAP_SETFCAP in the parent. But every task gets a full capability set when it enters a user namespace, and from then on its credentials say nothing about what it was allowed to do outside. The only trace that is kept is ns->parent_could_setfcap, which describes the task that created the namespace and not the ones that use it. Add cred->setfcap_level: CAP_SETFCAP of these credentials counts in cred->user_ns and in its ancestors down to that ->level. It is 0 in the initial namespace, where the capability sets speak for themselves. When credentials enter a namespace, set_cred_user_ns() keeps the value if they hold CAP_SETFCAP, and otherwise raises it to the first namespace on the way down over which they have it: the child they own, or else the namespace that is being entered. cap_setfcap_level() computes that. The value can only grow, and it is inherited over fork() and execve() like the rest of the cred. key_change_session_keyring() copies field by field and has to carry it over. Nothing looks at the new field yet. Assisted-by: LLM Signed-off-by: Josef Bacik --- include/linux/capability.h | 3 +++ include/linux/cred.h | 1 + kernel/user_namespace.c | 3 +++ security/commoncap.c | 26 ++++++++++++++++++++++++++ security/keys/process_keys.c | 1 + 5 files changed, 34 insertions(+) diff --git a/include/linux/capability.h b/include/linux/capability.h index 622137f66f09..026974a5111b 100644 --- a/include/linux/capability.h +++ b/include/linux/capability.h @@ -34,6 +34,7 @@ struct cpu_vfs_cap_data { #define _USER_CAP_HEADER_SIZE (sizeof(struct __user_cap_header_struct)) #define _KERNEL_CAP_T_SIZE (sizeof(kernel_cap_t)) +struct cred; struct file; struct inode; struct dentry; @@ -222,4 +223,6 @@ int get_vfs_caps_from_disk(const struct mnt_idmap *idmap, int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry, const void **ivalue, size_t size); +int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns); + #endif /* !_LINUX_CAPABILITY_H */ diff --git a/include/linux/cred.h b/include/linux/cred.h index 6ef1750c93e2..b95119099993 100644 --- a/include/linux/cred.h +++ b/include/linux/cred.h @@ -123,6 +123,7 @@ struct cred { kuid_t fsuid; /* UID for VFS ops */ kgid_t fsgid; /* GID for VFS ops */ unsigned securebits; /* SUID-less security management */ + int setfcap_level; /* how far up CAP_SETFCAP counts */ kernel_cap_t cap_inheritable; /* caps our children can inherit */ kernel_cap_t cap_permitted; /* caps we're permitted */ kernel_cap_t cap_effective; /* caps we can actually use */ diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c index 1b23d819d398..6b45df3a8d82 100644 --- a/kernel/user_namespace.c +++ b/kernel/user_namespace.c @@ -44,6 +44,9 @@ static void dec_user_namespaces(struct ucounts *ucounts) static void set_cred_user_ns(struct cred *cred, struct user_namespace *user_ns) { + /* The last chance to see what we can do outside of the new namespace. */ + cred->setfcap_level = cap_setfcap_level(cred, user_ns->parent); + /* Start with the same capabilities as init but useless for doing * anything as the capabilities are bound to the new user namespace. */ diff --git a/security/commoncap.c b/security/commoncap.c index d47ab3022343..513473f859be 100644 --- a/security/commoncap.c +++ b/security/commoncap.c @@ -131,6 +131,32 @@ int cap_capable(const struct cred *cred, struct user_namespace *target_ns, return ret; } +/** + * cap_setfcap_level - Determine how far up CAP_SETFCAP of a cred reaches + * @cred: The credentials to use + * @ns: The user namespace of @cred or one of its descendants + * + * File capabilities belong to the kuid of a namespace's root user, and the + * same kuid can be the root user of ancestors of that namespace. Every task + * gets CAP_SETFCAP when it enters a user namespace, so having it there says + * nothing about those ancestors. cred->setfcap_level does: it is handed down + * from namespace to namespace for as long as the capability is held. + * + * Return: the ->level of the topmost namespace, from @ns upwards, for which + * CAP_SETFCAP of @cred counts; @ns->level + 1 if it doesn't even over @ns. + */ +int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns) +{ + if (cap_capable(cred, ns, CAP_SETFCAP, CAP_OPT_NOAUDIT)) + return ns->level + 1; + + if (cap_raised(cred->cap_effective, CAP_SETFCAP)) + return cred->setfcap_level; + + /* All we have is that we own a child of our namespace. */ + return cred->user_ns->level + 1; +} + /** * cap_settime - Determine whether the current process may set the system clock * @ts: The time to set diff --git a/security/keys/process_keys.c b/security/keys/process_keys.c index a63c46bb2d14..f9cf3ce2426f 100644 --- a/security/keys/process_keys.c +++ b/security/keys/process_keys.c @@ -939,6 +939,7 @@ void key_change_session_keyring(struct callback_head *twork) new->group_info = get_group_info(old->group_info); new->securebits = old->securebits; + new->setfcap_level = old->setfcap_level; new->cap_inheritable = old->cap_inheritable; new->cap_permitted = old->cap_permitted; new->cap_effective = old->cap_effective; -- 2.55.0