napi_disable() returns once ibmveth_poll() has called napi_complete_done(), but the poll is not finished: it then re-enables the interrupt and calls ibmveth_rxq_pending_buffer(), which reads the RX queue. ibmveth_close() can free that queue in the meantime, and the late enable can leave the interrupt unmasked after close() masked it. The request_irq() failure path in ibmveth_open() frees the same memory after napi_disable() too. Call synchronize_net() after napi_disable() on both paths. Every caller of ibmveth_poll() runs it with bottom halves or interrupts disabled, so this waits for the poll to return. Found by AI-assisted review of the ibmveth multi-queue RX series and confirmed by code inspection; also raised by the Sashiko AI review of the first version of this series. The race was not reproduced. Tested on a POWER10 LPAR under an incoming ping flood with 30 rapid link down/up cycles and 30 rapid MTU cycles (1500 <-> 9000); ran cleanly with no warnings or faults. No kernel selftests cover ibmveth. Fixes: bea3348eef27 ("[NET]: Make NAPI polling independent of struct net_device objects.") Signed-off-by: Mingming Cao --- Changes in v2: - new patch; raised by the Sashiko review of v1 patch 1 as a pre-existing bug drivers/net/ethernet/ibm/ibmveth.c | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/drivers/net/ethernet/ibm/ibmveth.c b/drivers/net/ethernet/ibm/ibmveth.c index 7dec7ca9753b..fc90ee6663c3 100644 --- a/drivers/net/ethernet/ibm/ibmveth.c +++ b/drivers/net/ethernet/ibm/ibmveth.c @@ -769,6 +769,7 @@ static int ibmveth_open(struct net_device *netdev) netdev); if (rc != 0) { napi_disable(&adapter->napi); + synchronize_net(); netdev_err(netdev, "unable to request irq 0x%x, rc %d\n", netdev->irq, rc); goto out_free_buffer_pools; @@ -835,6 +836,11 @@ static int ibmveth_close(struct net_device *netdev) netdev_dbg(netdev, "close starting\n"); napi_disable(&adapter->napi); + /* napi_disable() returns once ibmveth_poll() has called + * napi_complete_done(), but the poll still re-enables the + * interrupt and reads the RX queue after that. + */ + synchronize_net(); netif_tx_disable(netdev); -- 2.50.1 (Apple Git-155)