Commit 40b9d1ab63f5 ("ipvlan: hold lower dev to avoid possible
use-after-free") added a reference to the lower net_device owned by struct
ipvl_port. However, this reference is released when the ipvlan device is
unregistered, which can happen before the ipvlan device is actually freed.
If a stacked network device configuration is created (e.g., a lower
net_device, an ipvlan device on top, and another device on top of ipvlan),
the upper device holds a reference to the ipvlan device. When the lower
net_device is unregistered, the ipvlan device is unregistered as well. This
triggers unregistration of the upper device, which queues asynchronous work
to drop its reference to the ipvlan device.
During this process, netdev_run_todo() waits for the refcounts of both the
lower net_device and the ipvlan device to drop to 1. Since the ipvlan
device's refcount is elevated by the pending upper device unregistration,
it is kept alive. However, because the lower net_device's reference held by
the ipvlan port is released during unregistration, the lower net_device's
refcount drops to 1, and netdev_run_todo() proceeds to free it. This leaves
ipvlan->phy_dev as a dangling pointer.
If an operation (like querying port attributes via ethtool) is performed on
the ipvlan device while it is still alive, it can access the freed lower
net_device, triggering a use-after-free:
Call Trace:
netdev_need_ops_lock include/net/netdev_lock.h:30 [inline]
netdev_lock_ops include/net/netdev_lock.h:41 [inline]
__ethtool_get_link_ksettings+0x230/0x250 net/ethtool/ioctl.c:463
__ethtool_get_link_ksettings+0x11f/0x250 net/ethtool/ioctl.c:464
ib_get_eth_speed+0x180/0x7f0 drivers/infiniband/core/verbs.c:2052
rxe_query_port+0x93/0x3d0 drivers/infiniband/sw/rxe/rxe_verbs.c:56
__ib_query_port drivers/infiniband/core/device.c:2129 [inline]
ib_query_port+0x16e/0x830 drivers/infiniband/core/device.c:2161
smc_ib_remember_port_attr net/smc/smc_ib.c:364 [inline]
smc_ib_port_event_work+0x147/0x920 net/smc/smc_ib.c:388
Fix this by holding a reference to the lower net_device using a
netdevice_tracker in struct ipvl_dev. The reference is acquired in
ipvlan_init() and released in the priv_destructor callback
(ipvlan_dev_free). Releasing the reference in priv_destructor guarantees
that the lower net_device is held until the ipvlan device is actually
freed, after its refcount has dropped to 0. This mirrors the behavior of
other stacked devices like macvlan and vlan, and safely covers ipvtap
devices as well.
Fixes: 2ad7bf363841 ("ipvlan: Initial check-in of the IPVLAN driver.")
Assisted-by: Gemini:gemini-3.5-flash Gemini:gemini-3.1-pro-preview syzbot
Reported-by: syzbot+5fe14f2ff4ccbace9a26@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=5fe14f2ff4ccbace9a26
Link: https://syzkaller.appspot.com/ai_job?id=ecee4f7f-c35a-40b5-85ed-47bebd5cb0d1
To: "Andrew Lunn"
To: "David S. Miller"
To: "Eric Dumazet"
To: "Jakub Kicinski"
To:
To: "Paolo Abeni"
Cc: "Dmitry Skorodumov"
Cc: "Kees Cook"
Cc:
---
v2:
- Updated the patch subject to "ipvlan: keep lower device alive until private destruction"
- Rewrote the commit message to clarify that commit 40b9d1ab63f5 added a lower-device reference owned by struct ipvl_port
- Described the lower device generically as a lower net_device
- Replaced the full KASAN report with only the relevant call chain
- Corrected the description of priv_destructor
v1:
https://lore.kernel.org/all/e66f374b-905f-471b-989e-60a3e0505873@mail.kernel.org/T/
---
diff --git a/drivers/net/ipvlan/ipvlan.h b/drivers/net/ipvlan/ipvlan.h
index 80f84fc87..13cdad002 100644
--- a/drivers/net/ipvlan/ipvlan.h
+++ b/drivers/net/ipvlan/ipvlan.h
@@ -64,6 +64,7 @@ struct ipvl_dev {
struct list_head pnode;
struct ipvl_port *port;
struct net_device *phy_dev;
+ netdevice_tracker dev_tracker;
struct list_head addrs;
struct ipvl_pcpu_stats __percpu *pcpu_stats;
DECLARE_BITMAP(mac_filters, IPVLAN_MAC_FILTER_SIZE);
diff --git a/drivers/net/ipvlan/ipvlan_main.c b/drivers/net/ipvlan/ipvlan_main.c
index ed46439a9..b1435296a 100644
--- a/drivers/net/ipvlan/ipvlan_main.c
+++ b/drivers/net/ipvlan/ipvlan_main.c
@@ -162,6 +162,9 @@ static int ipvlan_init(struct net_device *dev)
}
port = ipvlan_port_get_rtnl(phy_dev);
port->count += 1;
+
+ netdev_hold(phy_dev, &ipvlan->dev_tracker, GFP_KERNEL);
+
return 0;
}
@@ -673,6 +676,13 @@ void ipvlan_link_delete(struct net_device *dev, struct list_head *head)
}
EXPORT_SYMBOL_GPL(ipvlan_link_delete);
+static void ipvlan_dev_free(struct net_device *dev)
+{
+ struct ipvl_dev *ipvlan = netdev_priv(dev);
+
+ netdev_put(ipvlan->phy_dev, &ipvlan->dev_tracker);
+}
+
void ipvlan_link_setup(struct net_device *dev)
{
ether_setup(dev);
@@ -682,6 +692,7 @@ void ipvlan_link_setup(struct net_device *dev)
dev->priv_flags |= IFF_UNICAST_FLT | IFF_NO_QUEUE;
dev->netdev_ops = &ipvlan_netdev_ops;
dev->needs_free_netdev = true;
+ dev->priv_destructor = ipvlan_dev_free;
dev->header_ops = &ipvlan_header_ops;
dev->ethtool_ops = &ipvlan_ethtool_ops;
}
base-commit: 8cdeaa50eae8dad34885515f62559ee83e7e8dda
--
This is an AI-generated patch subject to moderation.
Reply with '#syz upstream' to Sign-off the patch as a human author
and send it to the upstream kernel mailing lists.
Reply with '#syz reject' to reject it ('#syz unreject' to undo).
See https://goo.gle/syzbot-ai-patches for information about AI-generated patches.
You can comment on the patch as usual, syzbot will try to address
the comments and send a new version of the patch if necessary.
syzbot engineers can be reached at syzkaller@googlegroups.com.