ibmveth_change_mtu(), veth_pool_store(), ibmveth_set_csum_offload() and ibmveth_set_tso() call close() and open() directly. If open() fails, NAPI is left disabled while IFF_UP stays set, so the next close() (ifdown, unregister or another reconfiguration) calls napi_disable() again and waits forever with RTNL held. Networking and shutdown hang; only a reboot recovers. ethtool -L in that state also wakes queues that have no TX buffer and dereferences NULL in ibmveth_start_xmit(). Any open() failure on those paths triggers it, for example an allocation failure on an MTU change to jumbo frames. Track a successful open in adapter->opened. close() returns early when it is clear, and set_channels() checks it instead of IFF_UP. The open() error-path leaks were fixed separately in net by commit af0524bf4ce1 ("ibmveth: h_free logical LAN on open-fail after register") and commit 84bec0bf0352 ("ibmveth: fix TX LTB and filter unwind on open-fail"), which are now in net-next too, so a failed open() no longer leaves TX buffers or the logical LAN registration behind for the early return to skip. Found by AI-assisted review of the ibmveth multi-queue RX series and confirmed by code inspection. Tested on a POWER10 LPAR with ibmveth_open() forced to fail by a test-only module parameter (not part of this patch): 'ip link set dev eth1 mtu 9000' fails, then 'ip link set dev eth1 down' returns at once and 'ip link set dev eth1 up' recovers the interface. No kernel selftests cover ibmveth. Fixes: 860f242eb534 ("[PATCH] ibmveth change buffer pools dynamically") Signed-off-by: Mingming Cao --- Changes in v3: - commit message: the two open() fixes are now in net-next Changes in v2: - commit message: say what a failed open() leaves behind without the two net fixes drivers/net/ethernet/ibm/ibmveth.c | 24 +++++++++++++++++------- drivers/net/ethernet/ibm/ibmveth.h | 2 ++ 2 files changed, 19 insertions(+), 7 deletions(-) diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c index 4a5869183be6..4665447997ea 100644 --- a/drivers/net/ethernet/ibm/ibmveth.c +++ b/drivers/net/ethernet/ibm/ibmveth.c @@ -739,6 +739,7 @@ static int ibmveth_open(struct net_device *netdev) netif_tx_start_all_queues(netdev); + adapter->opened = true; netdev_dbg(netdev, "open complete\n"); return 0; @@ -781,6 +782,14 @@ static int ibmveth_close(struct net_device *netdev) long lpar_rc; int i; + /* change_mtu, pool sysfs, set_csum and set_tso call close() and + * open() directly. If that open() fails, IFF_UP stays set and + * NAPI is disabled; a second close() would hang in napi_disable(). + */ + if (!adapter->opened) + return 0; + adapter->opened = false; + netdev_dbg(netdev, "close starting\n"); napi_disable(&adapter->napi); @@ -832,10 +841,10 @@ static int ibmveth_close(struct net_device *netdev) * * @w: pointer to work_struct embedded in adapter structure * - * Context: This routine acquires rtnl_mutex and disables its NAPI through - * ibmveth_close. It can't be called directly in a context that has - * already acquired rtnl_mutex or disabled its NAPI, or directly from - * a poll routine. + * Context: This routine acquires rtnl_mutex and, if the device is open, + * disables its NAPI through ibmveth_close. It can't be called + * directly in a context that has already acquired rtnl_mutex or + * disabled its NAPI, or directly from a poll routine. * * Return: void */ @@ -1129,10 +1138,11 @@ static int ibmveth_set_channels(struct net_device *netdev, goal = channels->tx_count; int rc, i; - /* If ndo_open has not been called yet then don't allocate, just set - * desired netdev_queue's and return + /* If the device is not open (including a failed close/open with + * IFF_UP still set) then don't allocate, just set desired + * netdev_queue's and return */ - if (!(netdev->flags & IFF_UP)) + if (!adapter->opened) return netif_set_real_num_tx_queues(netdev, goal); /* We have IBMVETH_MAX_QUEUES netdev_queue's allocated diff --git a/drivers/net/ethernet/ibm/ibmveth.h b/drivers/net/ethernet/ibm/ibmveth.h index d87713668ed3..3f2240823f6a 100644 --- a/drivers/net/ethernet/ibm/ibmveth.h +++ b/drivers/net/ethernet/ibm/ibmveth.h @@ -172,6 +172,8 @@ struct ibmveth_adapter { int rx_csum; int large_send; bool is_active_trunk; + /* Set by a successful ibmveth_open(), cleared by ibmveth_close(). */ + bool opened; unsigned int rx_buffers_per_hcall; u64 fw_ipv6_csum_support; -- 2.50.1 (Apple Git-155)