rtnl_fill_vf() closes the IFLA_VFINFO_LIST nest with nla_nest_end(), which stores the accumulated length into nla_len. That field is a u16, so a nest larger than 65535 bytes is written truncated modulo 65536. The skb is sized by if_nlmsg_size(), which adds rtnl_vfinfo_size() for every VF, so the buffer really is large enough for the whole list and none of the nla_put() calls in rtnl_fill_vfinfo() fails. The overflow is therefore silent. Userspace then walks the message with RTA_NEXT(), which advances by the stored length, so parsing resumes inside VF payload and the top-level attributes that follow the nest are read out of VF data. Those are IFLA_VF_PORTS, IFLA_XDP, IFLA_LINKINFO, IFLA_PERM_ADDRESS and IFLA_PROP_LIST. iproute2 prints "!!!Deficit", and strictly validating parsers reject the message outright. rtnl_fill_vfinfo() emits 296 bytes per VF on a 64-bit kernel with efficient unaligned access, so the nest wraps at 222 VFs; without CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS it is 328 bytes and 200 VFs, and with ndo_get_vf_guid 336 bytes and 196 VFs. ice supports up to 256 VFs per PF (ICE_MAX_SRIOV_VFS), so this is reachable on shipping hardware. Requests that set RTEXT_FILTER_SKIP_STATS need 335 VFs, which the 256-VF limit puts out of reach, so plain "ip link show" is fine today while "ip -s link show" is not. On CONFIG_DEBUG_NET kernels nla_nest_end() now also splats via DEBUG_NET_WARN_ON_ONCE(), added along with nla_nest_end_safe() in commit 1346586a9ac9 ("netlink: add a nla_nest_end_safe() helper"). Using nla_nest_end_safe() here would not help: a nest that does not fit in a u16 will not fit in a retried skb either, so returning -EMSGSIZE would turn "ip link show" on such a device into a hard failure. Bound the nest in the writer instead, as suggested when this was last discussed: drop the VF that would push it past U16_MAX and end the list there. Userspace sees IFLA_NUM_VF unchanged and a shorter IFLA_VFINFO_LIST, and everything after the nest stays parsable. An empty nest is already emitted for a PF with no VFs, so a list shorter than IFLA_NUM_VF is not a new encoding. The other large nests in rtnl_fill_ifinfo() were audited and cannot overflow: IFLA_AF_SPEC is bounded by a handful of address families at about a kilobyte each, and IFLA_VF_PORTS would need more than 560 VFs. Fixes: c02db8c6290b ("rtnetlink: make SR-IOV VF interface symmetric") Reported-by: Jacob Keller Closes: https://lore.kernel.org/netdev/16b289f6-b025-5dd3-443d-92d4c167e79c@intel.com/ Assisted-by: Claude:claude-fable-5 Signed-off-by: Artem Lytkin --- net/core/rtnetlink.c | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/net/core/rtnetlink.c b/net/core/rtnetlink.c index 12aa3aa1688b1..bd1d65dcf42fd 100644 --- a/net/core/rtnetlink.c +++ b/net/core/rtnetlink.c @@ -1687,10 +1687,22 @@ static noinline_for_stack int rtnl_fill_vf(struct sk_buff *skb, return -EMSGSIZE; for (i = 0; i < num_vfs; i++) { + unsigned char *mark = skb_tail_pointer(skb); + if (rtnl_fill_vfinfo(skb, dev, i, ext_filter_mask)) { nla_nest_cancel(skb, vfinfo); return -EMSGSIZE; } + + /* An attribute length is a u16, so the nest cannot describe + * more than U16_MAX bytes. Drop the VF that would overflow it + * and stop: a truncated list keeps the rest of the message + * parsable, whereas a wrapped nest length does not. + */ + if (skb_tail_pointer(skb) - (unsigned char *)vfinfo > U16_MAX) { + nlmsg_trim(skb, mark); + break; + } } nla_nest_end(skb, vfinfo); base-commit: 248951ddc14de84de3910f9b13f51491a8cd91df -- 2.43.0