| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/09/25 06:22 | flow | patch-triage |
0mArgs:null Results: {
"EnableConfigs": null,
"FocusSymbols": [],
"KMSANReasoning": "",
"NeedsKMSAN": false,
"Reasoning": "The patch introduces a new network driver (MPNIC) for Meta Platforms Network Interface Controllers. This is a vendor-specific PCIe NIC requiring dedicated physical ASIC hardware (PCI Vendor ID 0x1d9b, Device ID 0x0014) not emulated in standard virtualized environments (QEMU/GCE). As such, the code is structurally unreachable during fuzzing.",
"WorthFuzzing": false
} |
| 1/1 | 2026/09/25 06:22 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 3b134b7aeaaa9d7b2da52bb108b65ee735fd8e6c\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Fri Sep 25 06:22:52 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/drivers/net/ethernet/meta/Kconfig b/drivers/net/ethernet/meta/Kconfig\nindex ca5c7ac2a5bc2..f8b5fe7e31b73 100644\n--- a/drivers/net/ethernet/meta/Kconfig\n+++ b/drivers/net/ethernet/meta/Kconfig\n@@ -35,4 +35,17 @@ config FBNIC\n \t To compile this driver as a module, choose M here. The module\n \t will be called fbnic. MSI-X interrupt support is required.\n \n+config MPNIC\n+\ttristate \"Meta Platforms Network Interface Controller\"\n+\tdepends on 64BIT || COMPILE_TEST\n+\tdepends on !S390\n+\tdepends on PCI_MSI\n+\tselect PAGE_POOL\n+\thelp\n+\t This driver supports the Meta Platforms Network Interface\n+\t Controller.\n+\n+\t To compile this driver as a module, choose M here. The module\n+\t will be called mpnic.\n+\n endif # NET_VENDOR_META\ndiff --git a/drivers/net/ethernet/meta/Makefile b/drivers/net/ethernet/meta/Makefile\nindex 88804f3de9633..633973419c21b 100644\n--- a/drivers/net/ethernet/meta/Makefile\n+++ b/drivers/net/ethernet/meta/Makefile\n@@ -4,3 +4,4 @@\n #\n \n obj-$(CONFIG_FBNIC) += fbnic/\n+obj-$(CONFIG_MPNIC) += mpnic/\ndiff --git a/drivers/net/ethernet/meta/mpnic/Makefile b/drivers/net/ethernet/meta/mpnic/Makefile\nnew file mode 100644\nindex 0000000000000..d5bdbd5bd1c7b\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/Makefile\n@@ -0,0 +1,16 @@\n+# SPDX-License-Identifier: GPL-2.0\n+# Copyright (c) Meta Platforms, Inc. and affiliates.\n+\n+#\n+# Makefile for the Meta(R) Platforms Network Interface Controller\n+#\n+\n+obj-$(CONFIG_MPNIC) += mpnic.o\n+\n+mpnic-y := \\\n+\tmpnic_init.o \\\n+\tmpnic_irq.o \\\n+\tmpnic_netdev.o \\\n+\tmpnic_pci.o \\\n+\tmpnic_txrx.o \\\n+# End of mpnic-y\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic.h b/drivers/net/ethernet/meta/mpnic/mpnic.h\nnew file mode 100644\nindex 0000000000000..78359ab6abb12\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic.h\n@@ -0,0 +1,66 @@\n+/* SPDX-License-Identifier: GPL-2.0 */\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#ifndef _MPNIC_H_\n+#define _MPNIC_H_\n+\n+#include \u003clinux/interrupt.h\u003e\n+#include \u003clinux/io-64-nonatomic-lo-hi.h\u003e\n+#include \u003clinux/types.h\u003e\n+\n+#include \"mpnic_csr.h\"\n+\n+#define MPNIC_DRV_NAME\t\t\"mpnic\"\n+\n+#define MPNIC_MAX_TXQS\t\t1024u\n+#define MPNIC_MAX_RXQS\t\t1024u\n+\n+/* misc IRQ entries are allocated before the completion queue IRQs */\n+enum {\n+\tMPNIC_FW_MSIX_ENTRY,\n+\tMPNIC_NON_NAPI_VECTORS\n+};\n+\n+struct mpnic_dev {\n+\tstruct device *dev;\n+\tstruct net_device *netdev;\n+\n+\tu32 __iomem *uc_addr0;\n+\n+\tu16 num_irqs;\n+\n+\tu64 dsn;\n+\tu32 mps;\n+\tu32 readrq;\n+\tu8 relaxed_ord;\n+};\n+\n+u64 mpnic_rd64(struct mpnic_dev *mpd, u32 reg);\n+\n+int mpnic_dev_init(struct mpnic_dev *mpd);\n+\n+int mpnic_request_irq(struct mpnic_dev *mpd, int nr, irq_handler_t handler,\n+\t\t unsigned long flags, const char *name, void *data);\n+void mpnic_free_irq(struct mpnic_dev *mpd, int nr, void *data);\n+void mpnic_free_irqs(struct mpnic_dev *mpd);\n+int mpnic_alloc_irqs(struct mpnic_dev *mpd);\n+\n+static inline void mpnic_wr64(struct mpnic_dev *mpd, u32 reg, u64 val)\n+{\n+\tu32 __iomem *csr = READ_ONCE(mpd-\u003euc_addr0);\n+\n+\tif (csr)\n+\t\twriteq(val, csr + reg);\n+}\n+\n+static inline void mpnic_wrfl(struct mpnic_dev *mpd)\n+{\n+\tmpnic_rd64(mpd, MPNIC_BDQ_SPARE);\n+}\n+\n+static inline bool mpnic_present(struct mpnic_dev *mpd)\n+{\n+\treturn !!READ_ONCE(mpd-\u003euc_addr0);\n+}\n+\n+#endif /* _MPNIC_H_ */\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_csr.h b/drivers/net/ethernet/meta/mpnic/mpnic_csr.h\nnew file mode 100644\nindex 0000000000000..423378ca802c5\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_csr.h\n@@ -0,0 +1,321 @@\n+/* SPDX-License-Identifier: GPL-2.0 */\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#ifndef _MPNIC_CSR_H_\n+#define _MPNIC_CSR_H_\n+\n+#include \u003clinux/bits.h\u003e\n+\n+#define CSR_BIT(nr)\t\tBIT_ULL(nr)\n+#define CSR_GENMASK(h, l)\tGENMASK_ULL(h, l)\n+\n+#define DESC_BIT(nr)\t\tBIT_ULL(nr)\n+#define DESC_GENMASK(h, l)\tGENMASK_ULL(h, l)\n+\n+/* Transmit Work Descriptor Format */\n+#define MPNIC_TWD_L2_HLEN\t\tDESC_GENMASK(5, 0)\n+#define MPNIC_TWD_FLAG_REQ_COMPLETION\tDESC_BIT(37)\n+#define MPNIC_TWD_FLAG_DEST_MAC\t\tDESC_BIT(43)\n+#define MPNIC_TWD_TYPE\t\t\tDESC_GENMASK(47, 46)\n+enum {\n+\tMPNIC_TWD_TYPE_META\t= 0,\n+\tMPNIC_TWD_TYPE_AL\t= 2,\n+\tMPNIC_TWD_TYPE_LAST_AL\t= 3,\n+};\n+\n+#define MPNIC_TWD_ADDR\t\t\tDESC_GENMASK(45, 0)\n+#define MPNIC_TWD_LEN\t\t\tDESC_GENMASK(63, 48)\n+\n+/* Tx Completion Descriptor Format */\n+#define MPNIC_TCD_TYPE0_HEAD0\t\tDESC_GENMASK(15, 0)\n+#define MPNIC_TCD_DONE\t\t\tDESC_BIT(63)\n+\n+/* Rx Buffer Descriptor Format */\n+#define MPNIC_BD_DESC_ADDR\t\tDESC_GENMASK(39, 2)\n+#define MPNIC_BD_DESC_ID\t\tDESC_GENMASK(57, 40)\n+#define MPNIC_BD_DESC_BUF_SZ_LOG2\tDESC_GENMASK(62, 58)\n+\n+/* Rx Completion Queue Descriptors */\n+#define MPNIC_RCD_TYPE\t\t\tDESC_GENMASK(62, 61)\n+enum {\n+\tMPNIC_RCD_TYPE_HDR_AL\t= 0,\n+\tMPNIC_RCD_TYPE_PAY_AL\t= 1,\n+\tMPNIC_RCD_TYPE_META\t= 3,\n+};\n+\n+#define MPNIC_RCD_DONE\t\t\tDESC_BIT(63)\n+\n+#define MPNIC_RCD_HDR_SUBTYPE\t\tDESC_GENMASK(60, 59)\n+enum {\n+\tMPNIC_RCD_HDR_SUBTYPE_HDR\t= 2,\n+};\n+\n+/* Address/Length Completion Descriptors */\n+#define MPNIC_RCD_AL_BUFF_OFF\t\tDESC_GENMASK(15, 0)\n+#define MPNIC_RCD_AL_BUFF_ID\t\tDESC_GENMASK(33, 16)\n+#define MPNIC_RCD_AL_BUFF_LEN\t\tDESC_GENMASK(47, 34)\n+#define MPNIC_RCD_AL_PAGE_FIN\t\tDESC_BIT(53)\n+\n+/* Metadata Completion Descriptors */\n+#define MPNIC_RCD_META_ERR_MAC_EOP\t\tDESC_BIT(53)\n+#define MPNIC_RCD_META_ERR_TRUNCATED_FRAME\tDESC_BIT(54)\n+#define MPNIC_RCD_META_UNCORRECTABLE_ERR_MASK\t\\\n+\t(MPNIC_RCD_META_ERR_MAC_EOP | MPNIC_RCD_META_ERR_TRUNCATED_FRAME)\n+\n+/* Common fields for all DESC_CFG CSRs */\n+#define MPNIC_DESC_CFG_NUM_DESCS\tCSR_GENMASK(2, 0)\n+#define MPNIC_DESC_CFG_START_ADDR\tCSR_GENMASK(19, 8)\n+\n+/* Register Definitions\n+ *\n+ * The register file is addressed as an array of le32, so the byte address of\n+ * a register is 4 times the index below. Each register is listed with its\n+ * name, index and byte address.\n+ *\n+ *\tName\t\t\t\tIndex\t\t\tAddress\n+ *****************************************************************************/\n+\n+/* NIC_CORE_TDF */\n+#define MPNIC_TWQ_CTL(i, j)\t\t(0x0 + 1024 * (i) + 2 * (j))\n+\t\t\t\t\t\t\t\t/* 0x0 */\n+#define MPNIC_TWQ_CTL_RESET\t\t\tCSR_BIT(0)\n+#define MPNIC_TWQ_CTL_ENABLE\t\t\tCSR_BIT(1)\n+#define MPNIC_TWQ_TAIL(i, j)\t\t(0x4 + 1024 * (i) + 2 * (j))\n+\t\t\t\t\t\t\t\t/* 0x10 */\n+#define MPNIC_TWQ_SIZE(i, j)\t\t(0x10 + 1024 * (i) + 2 * (j))\n+\t\t\t\t\t\t\t\t/* 0x40 */\n+#define MPNIC_TWQ_SIZE_SIZE\t\t\tCSR_GENMASK(3, 0)\n+#define MPNIC_TWQ_BASE_ADDR(i, j)\t(0x1c + 1024 * (i) + 2 * (j))\n+\t\t\t\t\t\t\t\t/* 0x70 */\n+\n+/* NIC_CORE_TCM */\n+#define MPNIC_TCQ_CTL(i)\t\t(0x80 + 1024 * (i))\t/* 0x200 */\n+#define MPNIC_TCQ_CTL_RESET\t\t\tCSR_BIT(0)\n+#define MPNIC_TCQ_CTL_ENABLE\t\t\tCSR_BIT(1)\n+#define MPNIC_TCQ_BASE_ADDR(i)\t\t(0x86 + 1024 * (i))\t/* 0x218 */\n+#define MPNIC_TCQ_HEAD(i)\t\t(0x8e + 1024 * (i))\t/* 0x238 */\n+#define MPNIC_TCQ_SIZE(i)\t\t(0x94 + 1024 * (i))\t/* 0x250 */\n+#define MPNIC_TCQ_SIZE_SIZE\t\t\tCSR_GENMASK(4, 0)\n+\n+/* NIC_CORE_TIM */\n+#define MPNIC_TIM_CTL1(i)\t\t(0xc0 + 1024 * (i))\t/* 0x300 */\n+#define MPNIC_TIM_CTL1_UPD_IGN_LONG_EVENT_CNT\tCSR_BIT(48)\n+#define MPNIC_TIM_CTL1_UPD_IGN_LONG_TIME_CNT\tCSR_BIT(49)\n+#define MPNIC_TIM_CTL1_UPD_IGN_SHORT_TIME_CNT\tCSR_BIT(50)\n+#define MPNIC_TIM_CTL1_MASK\t\t\tCSR_BIT(51)\n+#define MPNIC_TIM_CTL1_MASK_EN\t\t\tCSR_BIT(52)\n+#define MPNIC_TIM_CTL1_TRIGGER\t\t\tCSR_BIT(53)\n+#define MPNIC_TIM_INTR_MASK(i)\t\t(0xc8 + 1024 * (i))\t/* 0x320 */\n+#define MPNIC_TIM_INTR_MASK_MASK\t\tCSR_BIT(0)\n+\n+/* NIC_CORE_RBP */\n+#define MPNIC_BDQ_CTL(i)\t\t(0x200 + 1024 * (i))\t/* 0x800 */\n+#define MPNIC_BDQ_CTL_RESET\t\t\tCSR_BIT(0)\n+#define MPNIC_BDQ_CTL_ENABLE\t\t\tCSR_BIT(1)\n+#define MPNIC_BDQ_CTL_ENABLE_PPQ\t\tCSR_BIT(3)\n+#define MPNIC_HPQ_TAIL(i)\t\t(0x202 + 1024 * (i))\t/* 0x808 */\n+#define MPNIC_PPQ_TAIL(i)\t\t(0x204 + 1024 * (i))\t/* 0x810 */\n+#define MPNIC_HPQ_SIZE(i)\t\t(0x20a + 1024 * (i))\t/* 0x828 */\n+#define MPNIC_HPQ_SIZE_SIZE\t\t\tCSR_GENMASK(4, 0)\n+#define MPNIC_PPQ_SIZE(i)\t\t(0x20c + 1024 * (i))\t/* 0x830 */\n+#define MPNIC_PPQ_SIZE_SIZE\t\t\tCSR_GENMASK(4, 0)\n+#define MPNIC_HPQ_BASE_ADDR(i)\t\t(0x216 + 1024 * (i))\t/* 0x858 */\n+#define MPNIC_PPQ_BASE_ADDR(i)\t\t(0x218 + 1024 * (i))\t/* 0x860 */\n+\n+/* NIC_CORE_RCM */\n+#define MPNIC_RCQ_CTL(i)\t\t(0x280 + 1024 * (i))\t/* 0xa00 */\n+#define MPNIC_RCQ_CTL_RESET\t\t\tCSR_BIT(0)\n+#define MPNIC_RCQ_CTL_ENABLE\t\t\tCSR_BIT(1)\n+#define MPNIC_RCQ_BASE_ADDR(i)\t\t(0x286 + 1024 * (i))\t/* 0xa18 */\n+#define MPNIC_RCQ_HEAD(i)\t\t(0x28e + 1024 * (i))\t/* 0xa38 */\n+#define MPNIC_RCQ_SIZE(i)\t\t(0x294 + 1024 * (i))\t/* 0xa50 */\n+#define MPNIC_RCQ_SIZE_SIZE\t\t\tCSR_GENMASK(4, 0)\n+\n+/* NIC_CORE_RIM */\n+#define MPNIC_RIM_INTR_MASK(i)\t\t(0x2c8 + 1024 * (i))\t/* 0xb20 */\n+#define MPNIC_RIM_INTR_MASK_MASK\t\tCSR_BIT(0)\n+\n+/* NIC_CORE_TIM_PRV */\n+#define MPNIC_TIM_CTL(i)\t\t(0x100100 + 1024 * (i))\t/* 0x400400 */\n+\n+/* NIC_CORE_RDE */\n+#define MPNIC_RDE_CFG(i)\t\t(0x10021c + 1024 * (i))\t/* 0x400870 */\n+#define MPNIC_RDE_CFG_MIN_TAIL_ROOM\t\tCSR_GENMASK(9, 0)\n+#define MPNIC_RDE_CFG_MIN_HEAD_ROOM\t\tCSR_GENMASK(18, 10)\n+#define MPNIC_RDE_CFG_MAX_HEADER_BYTES\t\tCSR_GENMASK(45, 32)\n+\n+/* NIC_CORE_RIM_PRV */\n+#define MPNIC_RIM_CTL(i)\t\t(0x100280 + 1024 * (i))\t/* 0x400a00 */\n+\n+/* NIC_CORE_RBP_HP_GLBL */\n+#define MPNIC_HPQ_IDLE(i)\t\t(0x420000 + 2 * (i))\t/* 0x1080000 */\n+#define MPNIC_HPQ_IDLE_CNT\t\t16\n+#define MPNIC_PPQ_IDLE(i)\t\t(0x420060 + 2 * (i))\t/* 0x1080180 */\n+#define MPNIC_PPQ_IDLE_CNT\t\t16\n+#define MPNIC_BDQ_GLBL_CTL0\t\t0x420080\t\t/* 0x1080200 */\n+#define MPNIC_BDQ_GLBL_CTL0_MAX_REQ_SIZE\tCSR_GENMASK(26, 18)\n+#define MPNIC_BDQ_GLBL_CTL0_PREFETCH_SPACE_THRESH \\\n+\t\t\t\t\t\tCSR_GENMASK(42, 32)\n+#define MPNIC_RDE_CTL\t\t\t0x420082\t\t/* 0x1080208 */\n+#define MPNIC_RDE_CTL_HPQ_DROP_THRESHOLD\tCSR_GENMASK(10, 0)\n+#define MPNIC_RDE_CTL_PPQ_DROP_THRESHOLD\tCSR_GENMASK(21, 11)\n+#define MPNIC_RDE_CTL_HPQ_LOCAL_DROP_THRESHOLD\tCSR_GENMASK(42, 32)\n+#define MPNIC_RDE_CTL_PPQ_LOCAL_DROP_THRESHOLD\tCSR_GENMASK(53, 43)\n+#define MPNIC_BDQ_MEM_INIT_REQ\t\t0x42013a\t\t/* 0x10804e8 */\n+#define MPNIC_BDQ_MEM_INIT_DONE\t\t0x42013c\t\t/* 0x10804f0 */\n+#define MPNIC_BDQ_SPARE\t\t\t0x42013e\t\t/* 0x10804f8 */\n+#define MPNIC_HPQ_DESC_CFG(i)\t\t(0x420140 + 2 * (i))\t/* 0x1080500 */\n+#define MPNIC_PPQ_DESC_CFG(i)\t\t(0x420940 + 2 * (i))\t/* 0x1082500 */\n+\n+/* NIC_CORE_RDE_GLBL */\n+#define MPNIC_RDE_MEM_INIT_REQ\t\t0x4240e6\t\t/* 0x1090398 */\n+#define MPNIC_RDE_MEM_INIT_DONE\t\t0x4240e8\t\t/* 0x10903a0 */\n+\n+/* NIC_CORE_RCM_GLBL */\n+#define MPNIC_RCQ_IDLE(i)\t\t(0x42505e + 2 * (i))\t/* 0x1094178 */\n+#define MPNIC_RCQ_IDLE_CNT\t\t16\n+#define MPNIC_RCM_MEM_INIT_REQ\t\t0x42507e\t\t/* 0x10941f8 */\n+#define MPNIC_RCM_MEM_INIT_DONE\t\t0x425080\t\t/* 0x1094200 */\n+\n+/* NIC_CORE_RNI_GLBL */\n+#define MPNIC_RNI_RBP_CTL\t\t0x427000\t\t/* 0x109c000 */\n+#define MPNIC_RNI_RDE_CTL\t\t0x427002\t\t/* 0x109c008 */\n+#define MPNIC_RNI_RDE_CTL_MPS\t\t\tCSR_GENMASK(1, 0)\n+#define MPNIC_RNI_RDE_CTL_CLS\t\t\tCSR_GENMASK(3, 2)\n+#define MPNIC_RNI_RCM_CTL\t\t0x427004\t\t/* 0x109c010 */\n+\n+/* NIC_CORE_TDF_GLBL */\n+#define MPNIC_TWQ_IDLE(i)\t\t(0x428042 + 2 * (i))\t/* 0x10a0108 */\n+#define MPNIC_TWQ_IDLE_CNT\t\t32\n+#define MPNIC_TWQ_DEF_PRI_TWD\t\t0x428082\t\t/* 0x10a0208 */\n+#define MPNIC_TDF_MEM_INIT_REQ\t\t0x42813a\t\t/* 0x10a04e8 */\n+#define MPNIC_TDF_MEM_INIT_DONE\t\t0x42813c\t\t/* 0x10a04f0 */\n+#define MPNIC_TDF_DESC_CFG(i)\t\t(0x428140 + 2 * (i))\t/* 0x10a0500 */\n+\n+/* NIC_CORE_TQS_GLBL */\n+#define MPNIC_TQS_GLBL_CTL0\t\t0x42a000\t\t/* 0x10a8000 */\n+#define MPNIC_TQS_GLBL_CTL0_TWD_ERROR_CHECK_EN\tCSR_BIT(2)\n+#define MPNIC_TQS_GLBL_P0_0\t\t0x42a002\t\t/* 0x10a8008 */\n+#define MPNIC_TQS_GLBL_P0_0_TXB_MAX_CRDTS_0\tCSR_GENMASK(63, 48)\n+#define MPNIC_TQS_GLBL_P0_1\t\t0x42a004\t\t/* 0x10a8010 */\n+#define MPNIC_TQS_GLBL_BMC\t\t0x42a012\t\t/* 0x10a8048 */\n+#define MPNIC_TQS_GLBL_BMC_TXB_MAX_CRDTS\tCSR_GENMASK(15, 0)\n+#define MPNIC_TQS_SLOWDOWN_CTL\t\t0x42a026\t\t/* 0x10a8098 */\n+#define MPNIC_TQS_SLOWDOWN_CTL_ENABLE\t\tCSR_BIT(6)\n+#define MPNIC_TQS_MTU_CTL0\t\t0x42a030\t\t/* 0x10a80c0 */\n+#define MPNIC_TQS_MTU_CTL1\t\t0x42a032\t\t/* 0x10a80c8 */\n+#define MPNIC_TQS_IDLE(i)\t\t(0x42a040 + 2 * (i))\t/* 0x10a8100 */\n+#define MPNIC_TQS_IDLE_CNT\t\t32\n+#define MPNIC_TQS_SET_P0_MAP0(i)\t(0x42a082 + 2 * (i))\t/* 0x10a8208 */\n+#define MPNIC_TQS_SET_P0_MAP1(i)\t(0x42a092 + 2 * (i))\t/* 0x10a8248 */\n+#define MPNIC_TQS_GLBL_SHAPING\t\t0x42a108\t\t/* 0x10a8420 */\n+#define MPNIC_TQS_GLBL_SHAPING_DISABLE\t\tCSR_BIT(0)\n+#define MPNIC_TQS_ARB_CTL\t\t0x42a122\t\t/* 0x10a8488 */\n+#define MPNIC_TQS_ARB_CTL_SET_CRDT_BUCKET_EN\tCSR_BIT(9)\n+#define MPNIC_TQS_ARB_CTL_SET_IMM_DECR_EN\tCSR_BIT(8)\n+#define MPNIC_TQS_ARB_CTL_GROUP_CRDT_BUCKET_EN\tCSR_BIT(5)\n+#define MPNIC_TQS_ARB_CTL_GROUP_IMM_DECR_EN\tCSR_BIT(4)\n+#define MPNIC_TQS_ARB_CTL_QUEUE_CRDT_BUCKET_EN\tCSR_BIT(1)\n+#define MPNIC_TQS_ARB_CTL_QUEUE_IMM_DECR_EN\tCSR_BIT(0)\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0 \\\n+\t\t\t\t\t0x42a124\t\t/* 0x10a8490 */\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_QUEUE\tCSR_GENMASK(19, 0)\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_GROUP\tCSR_GENMASK(59, 32)\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1 \\\n+\t\t\t\t\t0x42a126\t\t/* 0x10a8498 */\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_SET\tCSR_GENMASK(27, 0)\n+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_PORT\tCSR_GENMASK(63, 32)\n+#define MPNIC_TQS_SRAM_INIT_CTL\t\t0x42a128\t\t/* 0x10a84a0 */\n+#define MPNIC_TQS_SRAM_INIT_CTL_QUANTUM\t\tCSR_GENMASK(31, 20)\n+#define MPNIC_TQS_SRAM_INIT_CTL_INIT\t\tCSR_BIT(32)\n+#define MPNIC_TQS_GROUP_INIT_CTL\t0x42a12a\t\t/* 0x10a84a8 */\n+#define MPNIC_TQS_GROUP_INIT_CTL_QUANTUM\tCSR_GENMASK(51, 32)\n+#define MPNIC_TQS_GROUP_INIT_CTL_INIT\t\tCSR_BIT(52)\n+#define MPNIC_TQS_SET_INIT_CTL\t\t0x42a12c\t\t/* 0x10a84b0 */\n+#define MPNIC_TQS_SET_INIT_CTL_QUANTUM\t\tCSR_GENMASK(51, 32)\n+#define MPNIC_TQS_SET_INIT_CTL_INIT\t\tCSR_BIT(52)\n+#define MPNIC_TQS_PORT_INIT_CTL\t\t0x42a12e\t\t/* 0x10a84b8 */\n+#define MPNIC_TQS_PORT_INIT_CTL_QUANTUM\t\tCSR_GENMASK(55, 32)\n+#define MPNIC_TQS_PORT_INIT_CTL_INIT\t\tCSR_BIT(56)\n+#define MPNIC_TQS_SRAM_STS\t\t0x42a130\t\t/* 0x10a84c0 */\n+#define MPNIC_TQS_PORT_CTL(i)\t\t(0x42a1e4 + 2 * (i))\t/* 0x10a8790 */\n+\n+/* NIC_CORE_TDE_GLBL */\n+#define MPNIC_TDE_IDLE(i)\t\t(0x42b000 + 2 * (i))\t/* 0x10ac000 */\n+#define MPNIC_TDE_IDLE_CNT\t\t32\n+#define MPNIC_TDE_MEM_INIT_REQ\t\t0x42b1ee\t\t/* 0x10ac7b8 */\n+#define MPNIC_TDE_MEM_INIT_DONE\t\t0x42b1f0\t\t/* 0x10ac7c0 */\n+\n+/* NIC_CORE_TCM_GLBL */\n+#define MPNIC_TCQ_IDLE(i)\t\t(0x42c09e + 2 * (i))\t/* 0x10b0278 */\n+#define MPNIC_TCQ_IDLE_CNT\t\t16\n+#define MPNIC_TCM_MEM_INIT_REQ\t\t0x42c0be\t\t/* 0x10b02f8 */\n+#define MPNIC_TCM_MEM_INIT_DONE\t\t0x42c0c0\t\t/* 0x10b0300 */\n+\n+/* NIC_CORE_TNI_GLBL */\n+#define MPNIC_TNI_GLBL_TDF_CTL\t\t0x42e000\t\t/* 0x10b8000 */\n+#define MPNIC_TNI_GLBL_TDF_CTL_MRRS\t\tCSR_GENMASK(2, 0)\n+#define MPNIC_TNI_GLBL_TDF_CTL_CLS\t\tCSR_GENMASK(5, 3)\n+#define MPNIC_TNI_GLBL_TDE_CTL\t\t0x42e002\t\t/* 0x10b8008 */\n+#define MPNIC_TNI_GLBL_TCM_CTL\t\t0x42e004\t\t/* 0x10b8010 */\n+\n+/* NIC_CORE_TXB */\n+#define MPNIC_TXB_PORT_CONFIG\t\t0x600000\t\t/* 0x1800000 */\n+#define MPNIC_TXB_PORT_CONFIG_PORT_MODE\t\tCSR_GENMASK(15, 13)\n+#define MPNIC_TXB_BMC\t\t\t0x600122\t\t/* 0x1800488 */\n+#define MPNIC_TXB_P0(i)\t\t\t(0x600124 + 2 * (i))\t/* 0x1800490 */\n+#define MPNIC_TXB_P0_CNT\t\t17\n+#define MPNIC_TXB_BMC_THRESH\t\t0x60025c\t\t/* 0x1800970 */\n+#define MPNIC_TXB_P0_THRESH(i)\t\t(0x60025e + 2 * (i))\t/* 0x1800978 */\n+#define MPNIC_TXB_P0_ARB_WEIGHTS(i)\t(0x600378 + 2 * (i))\t/* 0x1800de0 */\n+\n+/* NIC_CORE_RXB */\n+#define MPNIC_RXB_MEM_INIT_REQ\t\t0x620002\t\t/* 0x1880008 */\n+#define MPNIC_RXB_MEM_INIT_DONE\t\t0x620004\t\t/* 0x1880010 */\n+#define MPNIC_RXB_PORT_CFG(i)\t\t(0x620006 + 2 * (i))\t/* 0x1880018 */\n+#define MPNIC_RXB_PORT_CFG_FCS_STRIP_MODE\tCSR_GENMASK(22, 22)\n+enum {\n+\tMPNIC_FCS_MODE_KEEP\t= 0,\n+\tMPNIC_FCS_MODE_STRIP\t= 1,\n+};\n+\n+#define MPNIC_RXB_PORT_CLASS_CFG(i)\t(0x620020 + 2 * (i))\t/* 0x1880080 */\n+#define MPNIC_RXB_PORT_CLASS_CFG_DEFAULT_L2_ACTION \\\n+\t\t\t\t\t\tCSR_GENMASK(0, 0)\n+enum {\n+\tMPNIC_L2_ACTION_DROP\t= 0,\n+\tMPNIC_L2_ACTION_PASS\t= 1,\n+};\n+\n+#define MPNIC_RXB_TC_CRDTS(i)\t\t(0x620418 + 2 * (i))\t/* 0x1881060 */\n+#define MPNIC_RXB_POOL_COMMON_CRDTS(i)\t(0x620458 + 2 * (i))\t/* 0x1881160 */\n+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC(i) \\\n+\t\t\t\t\t(0x620472 + 2 * (i))\t/* 0x18811c8 */\n+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_THRESH\tCSR_GENMASK(15, 0)\n+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_TC_EN\tCSR_BIT(29)\n+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_MAX_CRDTS\tCSR_GENMASK(47, 32)\n+#define MPNIC_RXB_HOST_DROP_THRESH(i)\t(0x6205b0 + 2 * (i))\t/* 0x18816c0 */\n+#define MPNIC_RXB_BMC_CRDTS(i)\t\t(0x6205d0 + 2 * (i))\t/* 0x1881740 */\n+\n+/* NIC_CORE_RPC */\n+#define MPNIC_RPC_MEM_INIT_REQ\t\t0x780440\t\t/* 0x1e01100 */\n+#define MPNIC_RPC_MEM_INIT_DONE\t\t0x780442\t\t/* 0x1e01108 */\n+\n+/* NIC_CORE_ROF */\n+#define MPNIC_RSC_GLOBAL_CONF\t\t0x7e2002\t\t/* 0x1f88008 */\n+#define MPNIC_RSC_GLOBAL_CONF_RSC_DISABLE\tCSR_BIT(0)\n+\n+/* NIC_CORE_TOF */\n+#define MPNIC_TOF_TCAM_DEST_REMAP\t0x7e3022\t\t/* 0x1f8c088 */\n+\n+/* PEMO_WRAPPER */\n+#define MPNIC_OB_ATTR_RO\t\t\tCSR_BIT(1)\n+#define MPNIC_OB_ATTR_TDE_H\t\t0x9a000e\t\t/* 0x2680038 */\n+#define MPNIC_OB_ATTR_TDE_P\t\t0x9a0010\t\t/* 0x2680040 */\n+#define MPNIC_OB_ATTR_TDF\t\t0x9a0012\t\t/* 0x2680048 */\n+#define MPNIC_OB_ATTR_RBP_HPQ\t\t0x9a0014\t\t/* 0x2680050 */\n+#define MPNIC_OB_ATTR_RBP_PPQ\t\t0x9a0016\t\t/* 0x2680058 */\n+#define MPNIC_OB_ATTR_RDE_H\t\t0x9a0018\t\t/* 0x2680060 */\n+#define MPNIC_OB_ATTR_RDE_P\t\t0x9a001a\t\t/* 0x2680068 */\n+\n+#endif /* _MPNIC_CSR_H_ */\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_init.c b/drivers/net/ethernet/meta/mpnic/mpnic_init.c\nnew file mode 100644\nindex 0000000000000..f8ebb19766731\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_init.c\n@@ -0,0 +1,553 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#include \u003clinux/bitfield.h\u003e\n+#include \u003clinux/bits.h\u003e\n+#include \u003clinux/cache.h\u003e\n+#include \u003clinux/if_ether.h\u003e\n+#include \u003clinux/iopoll.h\u003e\n+#include \u003clinux/log2.h\u003e\n+#include \u003clinux/sizes.h\u003e\n+\n+#include \"mpnic.h\"\n+\n+#define MPNIC_MEM_INIT_POLL_US\t\t500\n+#define MPNIC_MEM_INIT_TO_US\t\t5000\n+\n+/* BDQ mem init:\n+ * bit 1: fifo_wrptr_mem\n+ * bit 0: fifo_rdptr_mem\n+ */\n+#define MPNIC_MEM_INIT_BDQ_VAL\t\t0x3\n+\n+/* RCM mem init:\n+ * bit 3: cq_base_addr\n+ * bit 2: cd_fifo_rptr_stats\n+ * bit 1: cd_fifo_wptr_stats\n+ * bit 0: cq_head_ptr_stats\n+ */\n+#define MPNIC_MEM_INIT_RCM_VAL\t\t0xf\n+\n+/* RDE mem init, bits 0-18. Bits 0-12 are the per-queue packet, error and\n+ * drop counters, bits 13-18 the two context memories of each of the HPQ,\n+ * PPQ and SPQ descriptor prefetchers.\n+ */\n+#define MPNIC_MEM_INIT_RDE_VAL\t\t0x7ffff\n+\n+/* RPC mem init, bit 0 covers the whole classifier. */\n+#define MPNIC_MEM_INIT_RPC_VAL\t\t0x1\n+\n+/* TCM mem init:\n+ * bit 3: cq_head_ptr_stats\n+ * bit 2: cd_fifo_wptr_stats\n+ * bit 1: cd_fifo_rptr_stats\n+ * bit 0: cq_base_addr\n+ */\n+#define MPNIC_MEM_INIT_TCM_VAL\t\t0xf\n+\n+/* TDE mem init:\n+ * bit 1: stats mem\n+ * bit 0: dma_head_ptr SRAM\n+ */\n+#define MPNIC_MEM_INIT_TDE_VAL\t\t0x3\n+\n+/* TDF mem init, bit 0 covers the descriptor fetch SRAM. */\n+#define MPNIC_MEM_INIT_TDF_VAL\t\t0x1\n+\n+/* RXB mem init, bit 0 covers the DMAC TCAM statistics RAM. */\n+#define MPNIC_MEM_INIT_RXB_VAL\t\t0x1\n+\n+/* TQS arbiter SRAM init done, one bit per DWRR level. */\n+#define MPNIC_TQS_ARB_INIT_VAL\t\t0xf\n+\n+/* On-chip SRAM allocated to each queue for descriptor fetch, in units of\n+ * descriptors. Valid values are 64, 128, 256, 512 and 1024.\n+ *\n+ * The partition sizes are chosen for the maximum number of queues the\n+ * device supports, so they do not have to be adjusted when the active\n+ * queue count changes:\n+ *\n+ * BDQ: 1 MiB / (8 B/DESC) / (1024 HPQ + 1024 PPQ) = 64 DESC/QUEUE\n+ * TWQ: 2 MiB / (8 B/DESC) / (1024 TXQ * 2 TWQ) = 128 DESC/QUEUE\n+ */\n+#define MPNIC_BDQ_SRAM_DESCS\t\t64u\n+#define MPNIC_TDF_SRAM_DESCS\t\t128u\n+\n+/* A total of 1 MiB worth of Tx credits is available, in units of 128 B.\n+ * The BMC gets a guaranteed share of them whether or not the host is\n+ * routing anything its way, everything else goes to MAC TC0.\n+ */\n+#define MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL\t\t800\n+#define MPNIC_TXB_P0_MAC_PVT_CRDT_INIT_VAL\t\\\n+\t(SZ_1M / 128 - 2 * MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL)\n+\n+/* The recommended lower bound for the TXB threshold is 80, based on a 10K\n+ * MTU. Round up by 20% to stay on the defensive side. The same reasoning\n+ * applies to the arbitration weight, which has to exceed the full packet\n+ * size.\n+ */\n+#define MPNIC_TXB_INIT_BMC_THRESH\t\t100\n+#define MPNIC_TXB_INIT_P0_THRESH\t\t100\n+#define MPNIC_TXB_INIT_P0_ARB_WEIGHTS\t\t0x64\n+\n+/* RXB host drop threshold in units of 128 B beats. Packets targeting a\n+ * queue are dropped when the available credits fall below it. 80 beats is\n+ * 10 KB, which is also the largest frame the device is configured for.\n+ */\n+#define MPNIC_RXB_INIT_HOST_DROP_THRESH\t\t0x50\n+\n+/* A total of 8 MiB of Rx buffer is available. The recommended per-TC pool\n+ * for a 800G configuration is 420 KB with all 8 TCs enabled. Only one TC\n+ * is in use, so give it 8 * 420 KB (in units of 128 B) and push the rest\n+ * to the common pool. Start drawing from the common pool as soon as the\n+ * TC0 credits fall below one max sized frame. The BMC keeps a small\n+ * reserve of its own whether or not the host talks to it.\n+ */\n+#define MPNIC_RXB_INIT_POOL_TC_CRDTS_P0\t\t(0xd20 * 8)\n+#define MPNIC_RXB_INIT_BMC_CRDTS\t\t0x20\n+#define MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH\t\\\n+\t(SZ_8M / 128 - MPNIC_RXB_INIT_POOL_TC_CRDTS_P0 - \\\n+\t MPNIC_RXB_INIT_BMC_CRDTS)\n+#define MPNIC_RXB_INIT_COMMON_CRDT_THRSH\t0x50\n+\n+/* TXB port mode selects the number of active MAC ports for Tx buffer\n+ * credit distribution and arbitration. The hardware only supports single,\n+ * dual and quad port, encoded as b'001, b'010 and b'100.\n+ */\n+#define MPNIC_TXB_PORT_MODE_SINGLE\t1\n+\n+/* The MAC traffic classes start at index 8 in the TXB arrays, the BMC\n+ * sits above them.\n+ */\n+#define MPNIC_TXB_TC_IDX_MAC_0\t\t8\n+#define MPNIC_TXB_TC_IDX_BMC\t\t16\n+\n+/* Largest frame the Tx queue scheduler will pass through. Anything above\n+ * it gets truncated.\n+ */\n+#define MPNIC_TQS_MTU_CTL0_MAX\t\t0x2800\n+\n+/* The unit of the DWRR quantum is 256 B. It has to be large enough for at\n+ * least 11 MTUs to be transmitted in one quantum; use 15 for headroom.\n+ */\n+#define MPNIC_TQS_DWRR_INIT_QUANTUM\t(15 * MPNIC_TQS_MTU_CTL0_MAX / 256)\n+\n+/* Lower bound on the credit available to a queue, group, set or port\n+ * before the scheduler stops issuing requests for it. 15 MTUs, matching\n+ * the DWRR quantum above.\n+ */\n+#define MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH\t(15 * MPNIC_TQS_MTU_CTL0_MAX)\n+\n+/* The 1 MiB Tx buffer is partitioned between the MAC and the BMC in units\n+ * of 1 KiB. The BMC portion is fixed at 100 KB.\n+ */\n+#define MPNIC_TQS_GLBL_TXB_CRDT_BMC\t100\n+#define MPNIC_TQS_GLBL_TXB_CRDT_MAC\t(SZ_1M / SZ_1K - \\\n+\t\t\t\t\t MPNIC_TQS_GLBL_TXB_CRDT_BMC)\n+\n+struct mpnic_init_poll {\n+\tu64 exp_val;\n+\tu32 addr;\n+};\n+\n+struct mpnic_poll_state {\n+\tint poll_idx;\n+\tu64 val;\n+};\n+\n+static void mpnic_tdf_glbl_init(struct mpnic_dev *mpd)\n+{\n+\t/* Default metadata descriptor, used for frames the driver did not\n+\t * prepend one to.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_TWQ_DEF_PRI_TWD,\n+\t\t FIELD_PREP(MPNIC_TWD_L2_HLEN, ETH_HLEN) |\n+\t\t MPNIC_TWD_FLAG_REQ_COMPLETION);\n+}\n+\n+static void mpnic_txb_init(struct mpnic_dev *mpd)\n+{\n+\tint i;\n+\n+\tmpnic_wr64(mpd, MPNIC_TXB_BMC, MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL);\n+\n+\t/* Zero the private credits of every traffic class, then hand the\n+\t * unreserved ones to MAC TC0.\n+\t */\n+\tfor (i = 0; i \u003c MPNIC_TXB_P0_CNT; i++)\n+\t\tmpnic_wr64(mpd, MPNIC_TXB_P0(i), 0);\n+\tmpnic_wr64(mpd, MPNIC_TXB_P0(MPNIC_TXB_TC_IDX_MAC_0),\n+\t\t MPNIC_TXB_P0_MAC_PVT_CRDT_INIT_VAL);\n+\tmpnic_wr64(mpd, MPNIC_TXB_P0(MPNIC_TXB_TC_IDX_BMC),\n+\t\t MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL);\n+\n+\tmpnic_wr64(mpd, MPNIC_TXB_BMC_THRESH, MPNIC_TXB_INIT_BMC_THRESH);\n+\n+\tmpnic_wr64(mpd, MPNIC_TXB_P0_THRESH(MPNIC_TXB_TC_IDX_MAC_0),\n+\t\t MPNIC_TXB_INIT_P0_THRESH);\n+\tmpnic_wr64(mpd, MPNIC_TXB_P0_ARB_WEIGHTS(MPNIC_TXB_TC_IDX_MAC_0),\n+\t\t MPNIC_TXB_INIT_P0_ARB_WEIGHTS);\n+\tmpnic_wr64(mpd, MPNIC_TXB_PORT_CONFIG,\n+\t\t FIELD_PREP(MPNIC_TXB_PORT_CONFIG_PORT_MODE,\n+\t\t\t MPNIC_TXB_PORT_MODE_SINGLE));\n+}\n+\n+static void mpnic_rxb_init(struct mpnic_dev *mpd)\n+{\n+\t/* Accept all packets that miss dmac tcam, until l2 filtering is\n+\t * implemented.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_RXB_PORT_CLASS_CFG(0),\n+\t\t FIELD_PREP(MPNIC_RXB_PORT_CLASS_CFG_DEFAULT_L2_ACTION,\n+\t\t\t MPNIC_L2_ACTION_PASS));\n+\tmpnic_wr64(mpd, MPNIC_RXB_PORT_CFG(0),\n+\t\t FIELD_PREP(MPNIC_RXB_PORT_CFG_FCS_STRIP_MODE,\n+\t\t\t MPNIC_FCS_MODE_STRIP));\n+\n+\tmpnic_wr64(mpd, MPNIC_RXB_HOST_DROP_THRESH(0),\n+\t\t MPNIC_RXB_INIT_HOST_DROP_THRESH);\n+\n+\tmpnic_wr64(mpd, MPNIC_RXB_TC_CRDTS(0),\n+\t\t MPNIC_RXB_INIT_POOL_TC_CRDTS_P0);\n+\tmpnic_wr64(mpd, MPNIC_RXB_COMMON_CRDT_CTRL_TC(0),\n+\t\t FIELD_PREP(MPNIC_RXB_COMMON_CRDT_CTRL_TC_THRESH,\n+\t\t\t MPNIC_RXB_INIT_COMMON_CRDT_THRSH) |\n+\t\t FIELD_PREP(MPNIC_RXB_COMMON_CRDT_CTRL_TC_MAX_CRDTS,\n+\t\t\t MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH) |\n+\t\t MPNIC_RXB_COMMON_CRDT_CTRL_TC_TC_EN);\n+\n+\t/* Only pool 0 is used, it gets all of the common credits */\n+\tmpnic_wr64(mpd, MPNIC_RXB_POOL_COMMON_CRDTS(0),\n+\t\t MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH);\n+\n+\tmpnic_wr64(mpd, MPNIC_RXB_BMC_CRDTS(0), MPNIC_RXB_INIT_BMC_CRDTS);\n+\n+\tmpnic_wr64(mpd, MPNIC_RXB_MEM_INIT_REQ, MPNIC_MEM_INIT_RXB_VAL);\n+}\n+\n+static u64 mpnic_desc_cfg(unsigned int sram_descs, unsigned int q_idx)\n+{\n+\t/* The hardware encodes the partition size as 64 * 2^n descriptors,\n+\t * and its start address in units of 64 descriptors.\n+\t */\n+\treturn FIELD_PREP(MPNIC_DESC_CFG_NUM_DESCS, __ffs(sram_descs) - 6) |\n+\t FIELD_PREP(MPNIC_DESC_CFG_START_ADDR,\n+\t\t\t q_idx * (sram_descs \u003e\u003e 6));\n+}\n+\n+static void mpnic_desc_sram_init(struct mpnic_dev *mpd)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c MPNIC_MAX_TXQS * 2; i++)\n+\t\tmpnic_wr64(mpd, MPNIC_TDF_DESC_CFG(i),\n+\t\t\t mpnic_desc_cfg(MPNIC_TDF_SRAM_DESCS, i));\n+\n+\tfor (i = 0; i \u003c MPNIC_MAX_RXQS; i++) {\n+\t\tmpnic_wr64(mpd, MPNIC_HPQ_DESC_CFG(i),\n+\t\t\t mpnic_desc_cfg(MPNIC_BDQ_SRAM_DESCS, i));\n+\t\tmpnic_wr64(mpd, MPNIC_PPQ_DESC_CFG(i),\n+\t\t\t mpnic_desc_cfg(MPNIC_BDQ_SRAM_DESCS,\n+\t\t\t\t\t MPNIC_MAX_RXQS + i));\n+\t}\n+}\n+\n+static void mpnic_rxglb_init(struct mpnic_dev *mpd)\n+{\n+\t/* Descriptor prefetch reads are only issued once 32 descriptors\n+\t * worth of FIFO space is available, and no more than 64 descriptors\n+\t * are fetched for one queue at a time so that a single queue cannot\n+\t * monopolize the bus. Both have to be multiples of 16 to keep the\n+\t * reads 128 B aligned.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_BDQ_GLBL_CTL0,\n+\t\t FIELD_PREP(MPNIC_BDQ_GLBL_CTL0_PREFETCH_SPACE_THRESH, 32) |\n+\t\t FIELD_PREP(MPNIC_BDQ_GLBL_CTL0_MAX_REQ_SIZE, 64));\n+\n+\t/* Minimum number of descriptors which has to be available before\n+\t * the descriptor engine considers a queue usable, globally and in\n+\t * the per-queue prefetch FIFO.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_RDE_CTL,\n+\t\t FIELD_PREP(MPNIC_RDE_CTL_HPQ_DROP_THRESHOLD, 16) |\n+\t\t FIELD_PREP(MPNIC_RDE_CTL_PPQ_DROP_THRESHOLD, 16) |\n+\t\t FIELD_PREP(MPNIC_RDE_CTL_HPQ_LOCAL_DROP_THRESHOLD, 16) |\n+\t\t FIELD_PREP(MPNIC_RDE_CTL_PPQ_LOCAL_DROP_THRESHOLD, 16));\n+\n+\t/* Receive side coalescing is not supported yet */\n+\tmpnic_wr64(mpd, MPNIC_RSC_GLOBAL_CONF,\n+\t\t MPNIC_RSC_GLOBAL_CONF_RSC_DISABLE);\n+\n+\tmpnic_wr64(mpd, MPNIC_BDQ_MEM_INIT_REQ, MPNIC_MEM_INIT_BDQ_VAL);\n+\tmpnic_wr64(mpd, MPNIC_RCM_MEM_INIT_REQ, MPNIC_MEM_INIT_RCM_VAL);\n+\tmpnic_wr64(mpd, MPNIC_RDE_MEM_INIT_REQ, MPNIC_MEM_INIT_RDE_VAL);\n+\tmpnic_wr64(mpd, MPNIC_RPC_MEM_INIT_REQ, MPNIC_MEM_INIT_RPC_VAL);\n+}\n+\n+static void mpnic_txglb_init(struct mpnic_dev *mpd)\n+{\n+\t/* Nothing is redirected to the BMC until the Tx offload TCAM gets\n+\t * programmed with its addresses.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_TOF_TCAM_DEST_REMAP, 0);\n+\n+\tmpnic_wr64(mpd, MPNIC_TCM_MEM_INIT_REQ, MPNIC_MEM_INIT_TCM_VAL);\n+\tmpnic_wr64(mpd, MPNIC_TDE_MEM_INIT_REQ, MPNIC_MEM_INIT_TDE_VAL);\n+\tmpnic_wr64(mpd, MPNIC_TDF_MEM_INIT_REQ, MPNIC_MEM_INIT_TDF_VAL);\n+}\n+\n+/* Fill the DWRR arbiter memories. Setting the INIT bit makes the hardware\n+ * write the given credit and quantum into every entry at the queue, group,\n+ * set and port level, so nothing has to be programmed per queue.\n+ */\n+static void mpnic_tqs_sram_init(struct mpnic_dev *mpd)\n+{\n+\tmpnic_wr64(mpd, MPNIC_TQS_GROUP_INIT_CTL,\n+\t\t FIELD_PREP(MPNIC_TQS_GROUP_INIT_CTL_QUANTUM,\n+\t\t\t MPNIC_TQS_DWRR_INIT_QUANTUM) |\n+\t\t MPNIC_TQS_GROUP_INIT_CTL_INIT);\n+\tmpnic_wr64(mpd, MPNIC_TQS_SET_INIT_CTL,\n+\t\t FIELD_PREP(MPNIC_TQS_SET_INIT_CTL_QUANTUM,\n+\t\t\t MPNIC_TQS_DWRR_INIT_QUANTUM) |\n+\t\t MPNIC_TQS_SET_INIT_CTL_INIT);\n+\tmpnic_wr64(mpd, MPNIC_TQS_PORT_INIT_CTL,\n+\t\t FIELD_PREP(MPNIC_TQS_PORT_INIT_CTL_QUANTUM,\n+\t\t\t MPNIC_TQS_DWRR_INIT_QUANTUM) |\n+\t\t MPNIC_TQS_PORT_INIT_CTL_INIT);\n+\tmpnic_wr64(mpd, MPNIC_TQS_SRAM_INIT_CTL,\n+\t\t FIELD_PREP(MPNIC_TQS_SRAM_INIT_CTL_QUANTUM,\n+\t\t\t MPNIC_TQS_DWRR_INIT_QUANTUM) |\n+\t\t MPNIC_TQS_SRAM_INIT_CTL_INIT);\n+}\n+\n+static void mpnic_tqs_init(struct mpnic_dev *mpd)\n+{\n+\tu64 val;\n+\n+\t/* Initialize to the largest frame we support, the scheduler\n+\t * truncates anything above it. The BMC gets the same limit.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_TQS_MTU_CTL0, MPNIC_TQS_MTU_CTL0_MAX);\n+\tmpnic_wr64(mpd, MPNIC_TQS_MTU_CTL1, MPNIC_TQS_MTU_CTL0_MAX);\n+\n+\tmpnic_wr64(mpd, MPNIC_TQS_GLBL_CTL0,\n+\t\t MPNIC_TQS_GLBL_CTL0_TWD_ERROR_CHECK_EN);\n+\n+\tmpnic_tqs_sram_init(mpd);\n+\n+\t/* Only port 0 is used. A single traffic class is in use as well, so\n+\t * give all of the Tx buffer credits to TC0.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_TQS_GLBL_P0_0,\n+\t\t FIELD_PREP(MPNIC_TQS_GLBL_P0_0_TXB_MAX_CRDTS_0,\n+\t\t\t MPNIC_TQS_GLBL_TXB_CRDT_MAC));\n+\tmpnic_wr64(mpd, MPNIC_TQS_GLBL_P0_1, 0);\n+\tmpnic_wr64(mpd, MPNIC_TQS_GLBL_BMC,\n+\t\t FIELD_PREP(MPNIC_TQS_GLBL_BMC_TXB_MAX_CRDTS,\n+\t\t\t MPNIC_TQS_GLBL_TXB_CRDT_BMC));\n+\n+\tmpnic_wr64(mpd, MPNIC_TQS_PORT_CTL(0), 0);\n+\n+\t/* Map all sets to port 0 */\n+\tmpnic_wr64(mpd, MPNIC_TQS_SET_P0_MAP0(0), ~0ULL);\n+\tmpnic_wr64(mpd, MPNIC_TQS_SET_P0_MAP1(0), ~0ULL);\n+\n+\tmpnic_wr64(mpd, MPNIC_TQS_CEV_MIN_SCHED_THRESH_0,\n+\t\t FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_QUEUE,\n+\t\t\t MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH) |\n+\t\t FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_GROUP,\n+\t\t\t MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH));\n+\tmpnic_wr64(mpd, MPNIC_TQS_CEV_MIN_SCHED_THRESH_1,\n+\t\t FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_SET,\n+\t\t\t MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH) |\n+\t\t FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_PORT,\n+\t\t\t MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH));\n+\n+\t/* The rate limiters are left uninitialized, so shaping has to stay\n+\t * off or nothing would ever get scheduled.\n+\t */\n+\tmpnic_wr64(mpd, MPNIC_TQS_GLBL_SHAPING, MPNIC_TQS_GLBL_SHAPING_DISABLE);\n+\n+\t/* Enable fairness protection (phantom eligibility). Read modify\n+\t * write so that the reset default slowdown cycle is preserved.\n+\t */\n+\tval = mpnic_rd64(mpd, MPNIC_TQS_SLOWDOWN_CTL);\n+\tval |= MPNIC_TQS_SLOWDOWN_CTL_ENABLE;\n+\tmpnic_wr64(mpd, MPNIC_TQS_SLOWDOWN_CTL, val);\n+\n+\t/* Use immediate credit decrement at every DWRR level so that the\n+\t * credit reflects a grant in the same cycle.\n+\t */\n+\tval = mpnic_rd64(mpd, MPNIC_TQS_ARB_CTL);\n+\tval |= MPNIC_TQS_ARB_CTL_SET_CRDT_BUCKET_EN |\n+\t MPNIC_TQS_ARB_CTL_SET_IMM_DECR_EN |\n+\t MPNIC_TQS_ARB_CTL_GROUP_CRDT_BUCKET_EN |\n+\t MPNIC_TQS_ARB_CTL_GROUP_IMM_DECR_EN |\n+\t MPNIC_TQS_ARB_CTL_QUEUE_CRDT_BUCKET_EN |\n+\t MPNIC_TQS_ARB_CTL_QUEUE_IMM_DECR_EN;\n+\tmpnic_wr64(mpd, MPNIC_TQS_ARB_CTL, val);\n+}\n+\n+/* The MPS and CLS fields sit at the same bit positions in every block, so\n+ * one set of masks covers both the RNI and the TNI registers.\n+ */\n+static void mpnic_mps_init(struct mpnic_dev *mpd, u32 reg, unsigned int mps,\n+\t\t\t unsigned int cls)\n+{\n+\tu64 val = mpnic_rd64(mpd, reg);\n+\n+\tval \u0026= ~(MPNIC_RNI_RDE_CTL_MPS | MPNIC_RNI_RDE_CTL_CLS);\n+\tval |= FIELD_PREP(MPNIC_RNI_RDE_CTL_MPS, mps) |\n+\t FIELD_PREP(MPNIC_RNI_RDE_CTL_CLS, cls);\n+\n+\tmpnic_wr64(mpd, reg, val);\n+}\n+\n+/* Likewise for the MRRS and CLS fields, which have their own common\n+ * layout.\n+ */\n+static void mpnic_mrrs_init(struct mpnic_dev *mpd, u32 reg, unsigned int mrrs,\n+\t\t\t unsigned int cls)\n+{\n+\tu64 val = mpnic_rd64(mpd, reg);\n+\n+\tval \u0026= ~(MPNIC_TNI_GLBL_TDF_CTL_MRRS | MPNIC_TNI_GLBL_TDF_CTL_CLS);\n+\tval |= FIELD_PREP(MPNIC_TNI_GLBL_TDF_CTL_MRRS, mrrs) |\n+\t FIELD_PREP(MPNIC_TNI_GLBL_TDF_CTL_CLS, cls);\n+\n+\tmpnic_wr64(mpd, reg, val);\n+}\n+\n+/**\n+ * mpnic_axi_init - Configure AXI bus parameters from host PCIe capabilities\n+ * @mpd: Device to configure\n+ *\n+ * Programs the max read request size, max payload size and cache line size\n+ * of the DMA engines. The hardware encodes all three as a power of 2 index,\n+ * MRRS and MPS relative to 128 B and CLS relative to 64 B.\n+ *\n+ * MAX_OT and MAX_OB are left at their hardware defaults.\n+ */\n+static void mpnic_axi_init(struct mpnic_dev *mpd)\n+{\n+\tint mps, cls, mrrs;\n+\n+\tmps = clamp(ilog2(mpd-\u003emps) - 7, 0, 3);\n+\tcls = clamp(ilog2(L1_CACHE_BYTES) - 6, 0, 3);\n+\n+\tmpnic_mps_init(mpd, MPNIC_RNI_RDE_CTL, mps, cls);\n+\tmpnic_mps_init(mpd, MPNIC_RNI_RCM_CTL, mps, cls);\n+\tmpnic_mps_init(mpd, MPNIC_TNI_GLBL_TCM_CTL, mps, cls);\n+\n+\tmrrs = clamp(ilog2(mpd-\u003ereadrq) - 7, 0, 3);\n+\tmpnic_mrrs_init(mpd, MPNIC_RNI_RBP_CTL, mrrs, cls);\n+\tmpnic_mrrs_init(mpd, MPNIC_TNI_GLBL_TDF_CTL, mrrs, cls);\n+\n+\t/* TDE supports a wider range of MRRS encodings. */\n+\tmrrs = clamp(ilog2(mpd-\u003ereadrq) - 7, 0, 5);\n+\tmpnic_mrrs_init(mpd, MPNIC_TNI_GLBL_TDE_CTL, mrrs, cls);\n+}\n+\n+/**\n+ * mpnic_ro_init - Set relaxed ordering on the outbound TLP attributes\n+ * @mpd: Device to configure\n+ *\n+ * Completions must stay ordered so that they are not observed before the\n+ * payload DMA they describe has landed, so RCM and TCM are left alone.\n+ */\n+static void mpnic_ro_init(struct mpnic_dev *mpd)\n+{\n+\tu64 attr = mpd-\u003erelaxed_ord ? MPNIC_OB_ATTR_RO : 0;\n+\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_TDE_H, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_TDE_P, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_TDF, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_RBP_HPQ, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_RBP_PPQ, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_RDE_H, attr);\n+\tmpnic_wr64(mpd, MPNIC_OB_ATTR_RDE_P, attr);\n+}\n+\n+static bool mpnic_init_status_ready(struct mpnic_dev *mpd,\n+\t\t\t\t const struct mpnic_init_poll *polls,\n+\t\t\t\t struct mpnic_poll_state *state)\n+{\n+\tu64 val;\n+\tint i;\n+\n+\tfor (i = state-\u003epoll_idx; polls[i].addr; i++) {\n+\t\tval = mpnic_rd64(mpd, polls[i].addr);\n+\n+\t\tif ((val \u0026 polls[i].exp_val) != polls[i].exp_val) {\n+\t\t\tstate-\u003epoll_idx = i;\n+\t\t\tstate-\u003eval = val;\n+\t\t\treturn false;\n+\t\t}\n+\t}\n+\n+\treturn true;\n+}\n+\n+/**\n+ * mpnic_mem_init_poll - Wait for the memory initializations to complete\n+ * @mpd: Device to poll\n+ *\n+ * The blocks initialize their memories in parallel, so walk the status\n+ * registers in order and only go back to sleep on the first one which is\n+ * not done yet.\n+ *\n+ * Return: 0 on success, -ETIMEDOUT if not everything completed in time\n+ */\n+static int mpnic_mem_init_poll(struct mpnic_dev *mpd)\n+{\n+\tstatic const struct mpnic_init_poll polls[] = {\n+\t\t{ MPNIC_TQS_ARB_INIT_VAL, MPNIC_TQS_SRAM_STS },\n+\t\t{ MPNIC_MEM_INIT_BDQ_VAL, MPNIC_BDQ_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_RCM_VAL, MPNIC_RCM_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_RDE_VAL, MPNIC_RDE_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_RPC_VAL, MPNIC_RPC_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_TCM_VAL, MPNIC_TCM_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_TDE_VAL, MPNIC_TDE_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_TDF_VAL, MPNIC_TDF_MEM_INIT_DONE },\n+\t\t{ MPNIC_MEM_INIT_RXB_VAL, MPNIC_RXB_MEM_INIT_DONE },\n+\t\t{ 0 },\n+\t};\n+\tstruct mpnic_poll_state state = {};\n+\tbool done;\n+\tint err;\n+\n+\terr = read_poll_timeout(mpnic_init_status_ready, done, done,\n+\t\t\t\tMPNIC_MEM_INIT_POLL_US, MPNIC_MEM_INIT_TO_US,\n+\t\t\t\tfalse, mpd, polls, \u0026state);\n+\tif (err)\n+\t\tdev_err(mpd-\u003edev, \"Poll timeout for reg 0x%x: 0x%llx\\n\",\n+\t\t\tpolls[state.poll_idx].addr, state.val);\n+\n+\treturn err;\n+}\n+\n+int mpnic_dev_init(struct mpnic_dev *mpd)\n+{\n+\tint err;\n+\n+\tmpnic_tdf_glbl_init(mpd);\n+\tmpnic_txb_init(mpd);\n+\tmpnic_rxb_init(mpd);\n+\tmpnic_desc_sram_init(mpd);\n+\tmpnic_axi_init(mpd);\n+\tmpnic_ro_init(mpd);\n+\tmpnic_rxglb_init(mpd);\n+\tmpnic_txglb_init(mpd);\n+\tmpnic_tqs_init(mpd);\n+\n+\terr = mpnic_mem_init_poll(mpd);\n+\tif (err) {\n+\t\tdev_err(mpd-\u003edev, \"Device initialization failed: %d\\n\", err);\n+\t\treturn err;\n+\t}\n+\n+\tif (!mpnic_present(mpd))\n+\t\treturn -EIO;\n+\n+\treturn 0;\n+}\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_irq.c b/drivers/net/ethernet/meta/mpnic/mpnic_irq.c\nnew file mode 100644\nindex 0000000000000..bcc33655cbbea\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_irq.c\n@@ -0,0 +1,64 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#include \u003clinux/cpumask.h\u003e\n+#include \u003clinux/interrupt.h\u003e\n+#include \u003clinux/minmax.h\u003e\n+#include \u003clinux/pci.h\u003e\n+\n+#include \"mpnic.h\"\n+\n+int mpnic_request_irq(struct mpnic_dev *mpd, int nr, irq_handler_t handler,\n+\t\t unsigned long flags, const char *name, void *data)\n+{\n+\tstruct pci_dev *pdev = to_pci_dev(mpd-\u003edev);\n+\tint irq = pci_irq_vector(pdev, nr);\n+\n+\tif (irq \u003c 0)\n+\t\treturn irq;\n+\n+\treturn request_irq(irq, handler, flags, name, data);\n+}\n+\n+void mpnic_free_irq(struct mpnic_dev *mpd, int nr, void *data)\n+{\n+\tstruct pci_dev *pdev = to_pci_dev(mpd-\u003edev);\n+\tint irq = pci_irq_vector(pdev, nr);\n+\n+\tif (irq \u003c 0)\n+\t\treturn;\n+\n+\tfree_irq(irq, data);\n+}\n+\n+void mpnic_free_irqs(struct mpnic_dev *mpd)\n+{\n+\tstruct pci_dev *pdev = to_pci_dev(mpd-\u003edev);\n+\n+\tmpd-\u003enum_irqs = 0;\n+\tpci_free_irq_vectors(pdev);\n+}\n+\n+int mpnic_alloc_irqs(struct mpnic_dev *mpd)\n+{\n+\tunsigned int wanted_irqs = MPNIC_NON_NAPI_VECTORS;\n+\tstruct pci_dev *pdev = to_pci_dev(mpd-\u003edev);\n+\tint num_irqs;\n+\n+\twanted_irqs += min_t(unsigned int, num_online_cpus(), MPNIC_MAX_RXQS);\n+\tnum_irqs = pci_alloc_irq_vectors(pdev, MPNIC_NON_NAPI_VECTORS + 1,\n+\t\t\t\t\t wanted_irqs, PCI_IRQ_MSIX);\n+\tif (num_irqs \u003c 0) {\n+\t\tdev_err(mpd-\u003edev, \"Failed to allocate MSI-X entries: %d\\n\",\n+\t\t\tnum_irqs);\n+\t\treturn num_irqs;\n+\t}\n+\n+\tif (num_irqs \u003c wanted_irqs)\n+\t\tdev_warn(mpd-\u003edev, \"Allocated %d IRQs, expected %u\\n\",\n+\t\t\t num_irqs, wanted_irqs);\n+\n+\tmpd-\u003enum_irqs = num_irqs;\n+\n+\treturn 0;\n+}\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c\nnew file mode 100644\nindex 0000000000000..fd34f48a654ca\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c\n@@ -0,0 +1,181 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#include \u003clinux/etherdevice.h\u003e\n+#include \u003clinux/ipv6.h\u003e\n+#include \u003clinux/netdevice.h\u003e\n+#include \u003clinux/pci.h\u003e\n+#include \u003clinux/types.h\u003e\n+\n+#include \"mpnic.h\"\n+#include \"mpnic_netdev.h\"\n+#include \"mpnic_txrx.h\"\n+\n+static int mpnic_open(struct net_device *netdev)\n+{\n+\tstruct mpnic_net *mpn = netdev_priv(netdev);\n+\tint err;\n+\n+\terr = mpnic_alloc_napi_vectors(mpn);\n+\tif (err)\n+\t\treturn err;\n+\n+\terr = mpnic_alloc_resources(mpn);\n+\tif (err)\n+\t\tgoto err_free_napi_vectors;\n+\n+\terr = mpnic_set_netif_queues(mpn);\n+\tif (err)\n+\t\tgoto err_free_resources;\n+\n+\tmpnic_enable(mpn);\n+\tmpnic_fill(mpn);\n+\tmpnic_napi_enable(mpn);\n+\n+\tnetif_tx_wake_all_queues(netdev);\n+\tnetif_carrier_on(netdev);\n+\n+\treturn 0;\n+\n+err_free_resources:\n+\tmpnic_free_resources(mpn);\n+err_free_napi_vectors:\n+\tmpnic_free_napi_vectors(mpn);\n+\treturn err;\n+}\n+\n+static int mpnic_stop(struct net_device *netdev)\n+{\n+\tstruct mpnic_net *mpn = netdev_priv(netdev);\n+\n+\tnetif_carrier_off(netdev);\n+\n+\tmpnic_napi_disable(mpn);\n+\tnetif_tx_disable(netdev);\n+\n+\tmpnic_disable(mpn);\n+\tmpnic_wait_all_queues_idle(mpn-\u003empd);\n+\tmpnic_flush(mpn);\n+\n+\tmpnic_reset_netif_queues(mpn);\n+\tmpnic_free_resources(mpn);\n+\tmpnic_free_napi_vectors(mpn);\n+\n+\treturn 0;\n+}\n+\n+static const struct net_device_ops mpnic_netdev_ops = {\n+\t.ndo_open\t\t= mpnic_open,\n+\t.ndo_stop\t\t= mpnic_stop,\n+\t.ndo_validate_addr\t= eth_validate_addr,\n+\t.ndo_start_xmit\t\t= mpnic_xmit_frame,\n+};\n+\n+/**\n+ * mpnic_netdev_free - Free the netdev associated with mpnic\n+ * @mpd: Driver specific structure to free netdev from\n+ **/\n+void mpnic_netdev_free(struct mpnic_dev *mpd)\n+{\n+\tfree_netdev(mpd-\u003enetdev);\n+\tmpd-\u003enetdev = NULL;\n+}\n+\n+/**\n+ * mpnic_netdev_alloc - Allocate a netdev and associate it with mpnic\n+ * @mpd: Driver specific structure to associate the netdev with\n+ *\n+ * Return: NULL on failure.\n+ **/\n+struct net_device *mpnic_netdev_alloc(struct mpnic_dev *mpd)\n+{\n+\tstruct net_device *netdev;\n+\tstruct mpnic_net *mpn;\n+\tunsigned int queues;\n+\n+\tnetdev = alloc_etherdev_mq(sizeof(*mpn), MPNIC_MAX_RXQS);\n+\tif (!netdev)\n+\t\treturn NULL;\n+\n+\tSET_NETDEV_DEV(netdev, mpd-\u003edev);\n+\tmpd-\u003enetdev = netdev;\n+\n+\tnetdev-\u003enetdev_ops = \u0026mpnic_netdev_ops;\n+\tnetdev-\u003erequest_ops_lock = true;\n+\n+\tmpn = netdev_priv(netdev);\n+\tmpn-\u003enetdev = netdev;\n+\tmpn-\u003empd = mpd;\n+\n+\tmpn-\u003etxq_size = MPNIC_TXQ_SIZE_DEFAULT;\n+\tmpn-\u003ehpq_size = MPNIC_HPQ_SIZE_DEFAULT;\n+\tmpn-\u003eppq_size = MPNIC_PPQ_SIZE_DEFAULT;\n+\tmpn-\u003ercq_size = MPNIC_RCQ_SIZE_DEFAULT;\n+\n+\tqueues = min(netif_get_num_default_rss_queues(),\n+\t\t mpd-\u003enum_irqs - MPNIC_NON_NAPI_VECTORS);\n+\tmpn-\u003enum_tx_queues = queues;\n+\tmpn-\u003enum_rx_queues = queues;\n+\tmpn-\u003enum_napi = queues;\n+\n+\tnetdev-\u003efeatures |= NETIF_F_SG;\n+\tnetdev-\u003ehw_features |= netdev-\u003efeatures;\n+\tnetdev-\u003evlan_features |= netdev-\u003efeatures;\n+\n+\tnetdev-\u003emin_mtu = IPV6_MIN_MTU;\n+\tnetdev-\u003emax_mtu = MPNIC_MAX_JUMBO_FRAME_SIZE - ETH_HLEN;\n+\n+\tnetif_carrier_off(netdev);\n+\tnetif_tx_stop_all_queues(netdev);\n+\n+\treturn netdev;\n+}\n+\n+static int mpnic_dsn_to_mac_addr(u64 dsn, char *addr)\n+{\n+\taddr[0] = (dsn \u003e\u003e 56) \u0026 0xFF;\n+\taddr[1] = (dsn \u003e\u003e 48) \u0026 0xFF;\n+\taddr[2] = (dsn \u003e\u003e 40) \u0026 0xFF;\n+\taddr[3] = (dsn \u003e\u003e 16) \u0026 0xFF;\n+\taddr[4] = (dsn \u003e\u003e 8) \u0026 0xFF;\n+\taddr[5] = dsn \u0026 0xFF;\n+\n+\treturn is_valid_ether_addr(addr) ? 0 : -EINVAL;\n+}\n+\n+/**\n+ * mpnic_netdev_register - Assign the MAC address and register the netdev\n+ * @netdev: Netdev to register\n+ *\n+ * The permanent address is derived from the PCIe device serial number, the\n+ * same way the firmware and the BMC derive it. A random address would break\n+ * provisioning, so refuse to spawn the interface if the serial number does\n+ * not yield a valid one.\n+ *\n+ * Return: non-zero on failure.\n+ **/\n+int mpnic_netdev_register(struct net_device *netdev)\n+{\n+\tstruct mpnic_net *mpn = netdev_priv(netdev);\n+\tstruct mpnic_dev *mpd = mpn-\u003empd;\n+\tu8 addr[ETH_ALEN];\n+\tint err;\n+\n+\terr = mpnic_dsn_to_mac_addr(mpd-\u003edsn, addr);\n+\tif (err) {\n+\t\tdev_err(mpd-\u003edev, \"MAC addr %pM invalid\\n\", addr);\n+\t\treturn err;\n+\t}\n+\n+\tether_addr_copy(netdev-\u003eperm_addr, addr);\n+\teth_hw_addr_set(netdev, addr);\n+\n+\t/* Abort if MMIO has failed. This has to be the last check before\n+\t * registration, the register accessors can only detach the device\n+\t * once it has been registered.\n+\t */\n+\tif (!mpnic_present(mpd))\n+\t\treturn -EIO;\n+\n+\treturn register_netdev(netdev);\n+}\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h\nnew file mode 100644\nindex 0000000000000..ccb0929f9180a\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h\n@@ -0,0 +1,35 @@\n+/* SPDX-License-Identifier: GPL-2.0 */\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#ifndef _MPNIC_NETDEV_H_\n+#define _MPNIC_NETDEV_H_\n+\n+#include \u003clinux/types.h\u003e\n+\n+#include \"mpnic.h\"\n+#include \"mpnic_txrx.h\"\n+\n+struct mpnic_net {\n+\tstruct mpnic_ring *tx[MPNIC_MAX_TXQS];\n+\tstruct mpnic_ring *rx[MPNIC_MAX_RXQS];\n+\n+\tstruct mpnic_napi_vector *napi[MPNIC_MAX_NAPI_VECTORS];\n+\n+\tstruct net_device *netdev;\n+\tstruct mpnic_dev *mpd;\n+\n+\tu32 txq_size;\n+\tu32 hpq_size;\n+\tu32 ppq_size;\n+\tu32 rcq_size;\n+\n+\tu16 num_napi;\n+\tu16 num_tx_queues;\n+\tu16 num_rx_queues;\n+};\n+\n+struct net_device *mpnic_netdev_alloc(struct mpnic_dev *mpd);\n+void mpnic_netdev_free(struct mpnic_dev *mpd);\n+int mpnic_netdev_register(struct net_device *netdev);\n+\n+#endif /* _MPNIC_NETDEV_H_ */\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_pci.c b/drivers/net/ethernet/meta/mpnic/mpnic_pci.c\nnew file mode 100644\nindex 0000000000000..968cd611b8eab\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_pci.c\n@@ -0,0 +1,187 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#include \u003clinux/dma-mapping.h\u003e\n+#include \u003clinux/err.h\u003e\n+#include \u003clinux/module.h\u003e\n+#include \u003clinux/netdevice.h\u003e\n+#include \u003clinux/pci.h\u003e\n+#include \u003clinux/slab.h\u003e\n+#include \u003clinux/types.h\u003e\n+\n+#include \"mpnic.h\"\n+#include \"mpnic_netdev.h\"\n+\n+#define PCI_DEVICE_ID_META_MPNIC\t0x0014\n+\n+static void mpnic_mmio_err(struct mpnic_dev *mpd, u32 reg)\n+{\n+\t/* Hardware is giving us all 1's reads, assume it is gone */\n+\tWRITE_ONCE(mpd-\u003euc_addr0, NULL);\n+\n+\tdev_err(mpd-\u003edev,\n+\t\t\"Failed read (idx 0x%x AKA addr 0x%x), disabled CSR access, awaiting reset\\n\",\n+\t\treg, reg \u003c\u003c 2);\n+\n+\t/* Tell the stack the device has lost its PCIe link */\n+\tif (mpd-\u003enetdev)\n+\t\tnetif_device_detach(mpd-\u003enetdev);\n+}\n+\n+u64 mpnic_rd64(struct mpnic_dev *mpd, u32 reg)\n+{\n+\tu32 __iomem *csr = READ_ONCE(mpd-\u003euc_addr0);\n+\tu64 value;\n+\n+\tif (!csr)\n+\t\treturn ~0ULL;\n+\n+\tvalue = readq(csr + reg);\n+\n+\t/* If any bits are 0 value should be valid */\n+\tif (~value)\n+\t\treturn value;\n+\n+\t/* All ones can be a valid value, so confirm against a register\n+\t * which never reads that way on a live device.\n+\t */\n+\tif (reg != MPNIC_BDQ_SPARE \u0026\u0026 ~readq(csr + MPNIC_BDQ_SPARE))\n+\t\treturn value;\n+\n+\tmpnic_mmio_err(mpd, reg);\n+\n+\treturn ~0ULL;\n+}\n+\n+static struct mpnic_dev *mpnic_alloc(struct pci_dev *pdev)\n+{\n+\tstruct mpnic_dev *mpd;\n+\n+\tmpd = kzalloc_obj(*mpd);\n+\tif (!mpd)\n+\t\treturn NULL;\n+\n+\tpci_set_drvdata(pdev, mpd);\n+\tmpd-\u003edev = \u0026pdev-\u003edev;\n+\n+\tmpd-\u003edsn = pci_get_dsn(pdev);\n+\tmpd-\u003emps = pcie_get_mps(pdev);\n+\tmpd-\u003ereadrq = pcie_get_readrq(pdev);\n+\tmpd-\u003erelaxed_ord = pcie_relaxed_ordering_enabled(pdev);\n+\n+\treturn mpd;\n+}\n+\n+/**\n+ * mpnic_probe - Device initialization routine\n+ * @pdev: PCI device information struct\n+ * @ent: entry in mpnic_pci_tbl\n+ *\n+ * Return: 0 on success, negative on failure\n+ **/\n+static int mpnic_probe(struct pci_dev *pdev, const struct pci_device_id *ent)\n+{\n+\tstruct net_device *netdev;\n+\tvoid __iomem *uc_addr0;\n+\tstruct mpnic_dev *mpd;\n+\tint err;\n+\n+\tif (pdev-\u003eerror_state != pci_channel_io_normal) {\n+\t\tdev_err(\u0026pdev-\u003edev,\n+\t\t\t\"PCI device still in an error state. Unable to load...\\n\");\n+\t\treturn -EIO;\n+\t}\n+\n+\terr = pcim_enable_device(pdev);\n+\tif (err) {\n+\t\tdev_err(\u0026pdev-\u003edev, \"PCI enable device failed: %d\\n\", err);\n+\t\treturn err;\n+\t}\n+\n+\terr = dma_set_mask_and_coherent(\u0026pdev-\u003edev, DMA_BIT_MASK(46));\n+\tif (err) {\n+\t\tdev_err(\u0026pdev-\u003edev, \"DMA configuration failed: %d\\n\", err);\n+\t\treturn err;\n+\t}\n+\n+\tmpd = mpnic_alloc(pdev);\n+\tif (!mpd)\n+\t\treturn -ENOMEM;\n+\n+\tuc_addr0 = pcim_iomap_region(pdev, 0, MPNIC_DRV_NAME);\n+\tif (IS_ERR(uc_addr0)) {\n+\t\terr = PTR_ERR(uc_addr0);\n+\t\tdev_err(\u0026pdev-\u003edev, \"Mapping the register file failed: %d\\n\",\n+\t\t\terr);\n+\t\tgoto err_free_mpd;\n+\t}\n+\tmpd-\u003euc_addr0 = uc_addr0;\n+\n+\tpci_set_master(pdev);\n+\tpci_save_state(pdev);\n+\n+\terr = mpnic_alloc_irqs(mpd);\n+\tif (err)\n+\t\tgoto err_free_mpd;\n+\n+\terr = mpnic_dev_init(mpd);\n+\tif (err)\n+\t\tgoto err_free_irqs;\n+\n+\tnetdev = mpnic_netdev_alloc(mpd);\n+\tif (!netdev) {\n+\t\tdev_err(\u0026pdev-\u003edev, \"Netdev allocation failed\\n\");\n+\t\terr = -ENOMEM;\n+\t\tgoto err_free_irqs;\n+\t}\n+\n+\terr = mpnic_netdev_register(netdev);\n+\tif (err) {\n+\t\tdev_err(\u0026pdev-\u003edev, \"Netdev registration failed: %d\\n\", err);\n+\t\tgoto err_free_netdev;\n+\t}\n+\n+\treturn 0;\n+\n+err_free_netdev:\n+\tmpnic_netdev_free(mpd);\n+err_free_irqs:\n+\tmpnic_free_irqs(mpd);\n+err_free_mpd:\n+\tkfree(mpd);\n+\n+\treturn err;\n+}\n+\n+/**\n+ * mpnic_remove - Device removal routine\n+ * @pdev: PCI device information struct\n+ **/\n+static void mpnic_remove(struct pci_dev *pdev)\n+{\n+\tstruct mpnic_dev *mpd = pci_get_drvdata(pdev);\n+\n+\tunregister_netdev(mpd-\u003enetdev);\n+\tmpnic_netdev_free(mpd);\n+\tmpnic_free_irqs(mpd);\n+\tkfree(mpd);\n+}\n+\n+static const struct pci_device_id mpnic_pci_tbl[] = {\n+\t{ PCI_VDEVICE(META, PCI_DEVICE_ID_META_MPNIC) },\n+\t/* required last entry */\n+\t{}\n+};\n+MODULE_DEVICE_TABLE(pci, mpnic_pci_tbl);\n+\n+static struct pci_driver mpnic_driver = {\n+\t.name\t\t= MPNIC_DRV_NAME,\n+\t.id_table\t= mpnic_pci_tbl,\n+\t.probe\t\t= mpnic_probe,\n+\t.remove\t\t= mpnic_remove,\n+};\n+\n+module_pci_driver(mpnic_driver);\n+\n+MODULE_DESCRIPTION(\"Meta Platforms Network Interface Controller\");\n+MODULE_LICENSE(\"GPL\");\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c\nnew file mode 100644\nindex 0000000000000..9878ea5a2f8e1\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c\n@@ -0,0 +1,1398 @@\n+// SPDX-License-Identifier: GPL-2.0\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#include \u003clinux/bitfield.h\u003e\n+#include \u003clinux/dma-mapping.h\u003e\n+#include \u003clinux/iopoll.h\u003e\n+#include \u003clinux/pci.h\u003e\n+#include \u003clinux/slab.h\u003e\n+#include \u003cnet/page_pool/helpers.h\u003e\n+\n+#include \"mpnic.h\"\n+#include \"mpnic_netdev.h\"\n+#include \"mpnic_txrx.h\"\n+\n+struct mpnic_xmit_cb {\n+\tu32 bytecount;\n+\tu8 desc_count;\n+};\n+\n+#define MPNIC_XMIT_CB(__skb) ((struct mpnic_xmit_cb *)((__skb)-\u003ecb))\n+#define MPNIC_TWD_TYPE_MASK(_type) \\\n+\tcpu_to_le64(FIELD_PREP(MPNIC_TWD_TYPE, MPNIC_TWD_TYPE_##_type))\n+\n+/* Leave the interrupt moderation counters alone when arming or masking */\n+#define MPNIC_TIM_PARAM_CFG_PRESERVE_MASK \\\n+\t(MPNIC_TIM_CTL1_UPD_IGN_LONG_EVENT_CNT | \\\n+\t MPNIC_TIM_CTL1_UPD_IGN_LONG_TIME_CNT | \\\n+\t MPNIC_TIM_CTL1_UPD_IGN_SHORT_TIME_CNT)\n+\n+static void mpnic_nv_irq_disable(struct mpnic_napi_vector *nv)\n+{\n+\tmpnic_wr64(nv-\u003empd, MPNIC_TIM_CTL1(nv-\u003eqt[0].cmpl.q_idx),\n+\t\t MPNIC_TIM_PARAM_CFG_PRESERVE_MASK |\n+\t\t MPNIC_TIM_CTL1_MASK_EN | MPNIC_TIM_CTL1_MASK);\n+}\n+\n+static void mpnic_nv_irq_rearm(struct mpnic_napi_vector *nv)\n+{\n+\t/* Rearming a single queue on a given IRQ rearms all the other\n+\t * queues mapped to the same IRQ.\n+\t */\n+\tmpnic_wr64(nv-\u003empd, MPNIC_TIM_CTL1(nv-\u003eqt[0].cmpl.q_idx),\n+\t\t MPNIC_TIM_PARAM_CFG_PRESERVE_MASK | MPNIC_TIM_CTL1_MASK_EN);\n+}\n+\n+static unsigned int mpnic_desc_unused(struct mpnic_ring *ring)\n+{\n+\treturn (ring-\u003ehead - ring-\u003etail - 1) \u0026 ring-\u003esize_mask;\n+}\n+\n+static struct netdev_queue *mpnic_txring_txq(const struct net_device *dev,\n+\t\t\t\t\t const struct mpnic_ring *ring)\n+{\n+\treturn netdev_get_tx_queue(dev, ring-\u003eq_idx);\n+}\n+\n+static void mpnic_tx_doorbell(struct mpnic_ring *ring, __le64 *meta)\n+{\n+\t*meta |= cpu_to_le64(MPNIC_TWD_FLAG_REQ_COMPLETION);\n+\tring-\u003edeferred_meta = -1;\n+\n+\t/* Force DMA writes to flush before writing to tail */\n+\tdma_wmb();\n+\n+\twriteq(ring-\u003etail, ring-\u003edoorbell);\n+}\n+\n+/* Packets handed to us with xmit_more set are left in the ring without a\n+ * doorbell, and without a completion request, in the expectation that the\n+ * packet ending the burst will ring for all of them. If that packet gets\n+ * dropped instead we have to ring here, otherwise the descriptors sit in\n+ * the ring until the next transmit, which may never come.\n+ */\n+static void mpnic_tx_flush_doorbell(struct mpnic_ring *ring)\n+{\n+\tif (ring-\u003edeferred_meta \u003e= 0)\n+\t\tmpnic_tx_doorbell(ring, \u0026ring-\u003edesc[ring-\u003edeferred_meta]);\n+}\n+\n+static void mpnic_unmap_single_twd(struct device *dev, __le64 *twd)\n+{\n+\tu64 raw_twd = le64_to_cpu(*twd);\n+\n+\tdma_unmap_single(dev, FIELD_GET(MPNIC_TWD_ADDR, raw_twd),\n+\t\t\t FIELD_GET(MPNIC_TWD_LEN, raw_twd), DMA_TO_DEVICE);\n+}\n+\n+static void mpnic_unmap_page_twd(struct device *dev, __le64 *twd)\n+{\n+\tu64 raw_twd = le64_to_cpu(*twd);\n+\n+\tdma_unmap_page(dev, FIELD_GET(MPNIC_TWD_ADDR, raw_twd),\n+\t\t FIELD_GET(MPNIC_TWD_LEN, raw_twd), DMA_TO_DEVICE);\n+}\n+\n+static bool\n+mpnic_tx_map(struct mpnic_ring *ring, struct sk_buff *skb, __le64 *meta)\n+{\n+\tstruct device *dev = skb-\u003edev-\u003edev.parent;\n+\tunsigned int tail = ring-\u003etail, first;\n+\tunsigned int size, data_len;\n+\tskb_frag_t *frag;\n+\tdma_addr_t dma;\n+\t__le64 *twd;\n+\n+\ttail++;\n+\ttail \u0026= ring-\u003esize_mask;\n+\tfirst = tail;\n+\n+\tsize = skb_headlen(skb);\n+\tdata_len = skb-\u003edata_len;\n+\n+\tif (size \u003e FIELD_MAX(MPNIC_TWD_LEN))\n+\t\tgoto err_dma;\n+\n+\tdma = dma_map_single(dev, skb-\u003edata, size, DMA_TO_DEVICE);\n+\n+\tfor (frag = \u0026skb_shinfo(skb)-\u003efrags[0];; frag++) {\n+\t\ttwd = \u0026ring-\u003edesc[tail];\n+\n+\t\tif (dma_mapping_error(dev, dma))\n+\t\t\tgoto err_dma;\n+\n+\t\t*twd = cpu_to_le64(FIELD_PREP(MPNIC_TWD_ADDR, dma) |\n+\t\t\t\t FIELD_PREP(MPNIC_TWD_LEN, size) |\n+\t\t\t\t FIELD_PREP(MPNIC_TWD_TYPE,\n+\t\t\t\t\t MPNIC_TWD_TYPE_AL));\n+\n+\t\ttail++;\n+\t\ttail \u0026= ring-\u003esize_mask;\n+\n+\t\tif (!data_len)\n+\t\t\tbreak;\n+\n+\t\tsize = skb_frag_size(frag);\n+\t\tdata_len -= size;\n+\n+\t\tif (size \u003e FIELD_MAX(MPNIC_TWD_LEN))\n+\t\t\tgoto err_dma;\n+\n+\t\tdma = skb_frag_dma_map(dev, frag, 0, size, DMA_TO_DEVICE);\n+\t}\n+\n+\t*twd |= MPNIC_TWD_TYPE_MASK(LAST_AL);\n+\n+\tMPNIC_XMIT_CB(skb)-\u003edesc_count = ((twd - meta) + 1) \u0026 ring-\u003esize_mask;\n+\n+\tskb_tx_timestamp(skb);\n+\n+\tring-\u003etail = tail;\n+\n+\t/* Verify there is room for another packet */\n+\tnetif_txq_maybe_stop(mpnic_txring_txq(skb-\u003edev, ring),\n+\t\t\t mpnic_desc_unused(ring), MPNIC_MAX_SKB_DESC,\n+\t\t\t MPNIC_TX_DESC_WAKEUP);\n+\n+\tif (__netdev_tx_sent_queue(mpnic_txring_txq(skb-\u003edev, ring),\n+\t\t\t\t MPNIC_XMIT_CB(skb)-\u003ebytecount,\n+\t\t\t\t netdev_xmit_more()))\n+\t\tmpnic_tx_doorbell(ring, meta);\n+\telse\n+\t\tring-\u003edeferred_meta = meta - ring-\u003edesc;\n+\n+\treturn false;\n+err_dma:\n+\tif (net_ratelimit())\n+\t\tnetdev_err(skb-\u003edev, \"TX DMA map failed\\n\");\n+\n+\twhile (tail != first) {\n+\t\ttail--;\n+\t\ttail \u0026= ring-\u003esize_mask;\n+\t\ttwd = \u0026ring-\u003edesc[tail];\n+\t\tif (tail == first)\n+\t\t\tmpnic_unmap_single_twd(dev, twd);\n+\t\telse\n+\t\t\tmpnic_unmap_page_twd(dev, twd);\n+\t}\n+\n+\treturn true;\n+}\n+\n+#define MPNIC_MIN_FRAME_LEN\t60\n+\n+static netdev_tx_t mpnic_xmit_frame_ring(struct sk_buff *skb,\n+\t\t\t\t\t struct mpnic_ring *ring)\n+{\n+\t__le64 *meta = \u0026ring-\u003edesc[ring-\u003etail];\n+\tu32 tail = ring-\u003etail;\n+\n+\tif (skb_put_padto(skb, MPNIC_MIN_FRAME_LEN))\n+\t\tgoto err_drop;\n+\n+\tif (!netif_txq_maybe_stop(mpnic_txring_txq(skb-\u003edev, ring),\n+\t\t\t\t mpnic_desc_unused(ring), MPNIC_MAX_SKB_DESC,\n+\t\t\t\t MPNIC_TX_DESC_WAKEUP)) {\n+\t\tmpnic_tx_flush_doorbell(ring);\n+\t\treturn NETDEV_TX_BUSY;\n+\t}\n+\n+\tring-\u003etx_buf[tail] = skb;\n+\t*meta = cpu_to_le64(MPNIC_TWD_FLAG_DEST_MAC);\n+\n+\tMPNIC_XMIT_CB(skb)-\u003ebytecount = skb-\u003elen;\n+\tMPNIC_XMIT_CB(skb)-\u003edesc_count = 0;\n+\n+\tif (mpnic_tx_map(ring, skb, meta))\n+\t\tgoto err_free;\n+\n+\treturn NETDEV_TX_OK;\n+\n+err_free:\n+\tdev_kfree_skb_any(skb);\n+\tring-\u003etx_buf[tail] = NULL;\n+\tring-\u003etail = tail;\n+err_drop:\n+\tmpnic_tx_flush_doorbell(ring);\n+\n+\treturn NETDEV_TX_OK;\n+}\n+\n+netdev_tx_t mpnic_xmit_frame(struct sk_buff *skb, struct net_device *dev)\n+{\n+\tstruct mpnic_net *mpn = netdev_priv(dev);\n+\n+\treturn mpnic_xmit_frame_ring(skb, mpn-\u003etx[skb_get_queue_mapping(skb)]);\n+}\n+\n+static void mpnic_clean_twq0(struct mpnic_napi_vector *nv, int napi_budget,\n+\t\t\t struct mpnic_ring *ring, bool discard,\n+\t\t\t unsigned int hw_head)\n+{\n+\tu64 total_bytes = 0, total_packets = 0;\n+\tunsigned int head = ring-\u003ehead;\n+\tstruct netdev_queue *txq;\n+\tunsigned int clean_desc;\n+\n+\tclean_desc = (hw_head - head) \u0026 ring-\u003esize_mask;\n+\n+\twhile (clean_desc) {\n+\t\tstruct sk_buff *skb = ring-\u003etx_buf[head];\n+\t\tunsigned int desc_cnt;\n+\n+\t\tdesc_cnt = MPNIC_XMIT_CB(skb)-\u003edesc_count;\n+\t\tif (desc_cnt \u003e clean_desc)\n+\t\t\tbreak;\n+\n+\t\tring-\u003etx_buf[head] = NULL;\n+\n+\t\tclean_desc -= desc_cnt;\n+\n+\t\t/* Step over the metadata descriptor */\n+\t\thead++;\n+\t\thead \u0026= ring-\u003esize_mask;\n+\t\tdesc_cnt--;\n+\n+\t\tmpnic_unmap_single_twd(nv-\u003edev, \u0026ring-\u003edesc[head]);\n+\t\thead++;\n+\t\thead \u0026= ring-\u003esize_mask;\n+\t\tdesc_cnt--;\n+\n+\t\twhile (desc_cnt--) {\n+\t\t\tmpnic_unmap_page_twd(nv-\u003edev, \u0026ring-\u003edesc[head]);\n+\t\t\thead++;\n+\t\t\thead \u0026= ring-\u003esize_mask;\n+\t\t}\n+\n+\t\ttotal_bytes += MPNIC_XMIT_CB(skb)-\u003ebytecount;\n+\t\ttotal_packets++;\n+\n+\t\tnapi_consume_skb(skb, napi_budget);\n+\t}\n+\n+\tif (!total_bytes)\n+\t\treturn;\n+\n+\tring-\u003ehead = head;\n+\n+\tif (discard)\n+\t\treturn;\n+\n+\ttxq = mpnic_txring_txq(nv-\u003enapi.dev, ring);\n+\tnetif_txq_completed_wake(txq, total_packets, total_bytes,\n+\t\t\t\t mpnic_desc_unused(ring),\n+\t\t\t\t MPNIC_TX_DESC_WAKEUP);\n+}\n+\n+static void mpnic_commit_cq_head(struct mpnic_ring *cmpl)\n+{\n+\tu32 head = cmpl-\u003ehead;\n+\n+\t/* The tail shadows the last value written to the doorbell, so a\n+\t * completion queue which has not moved costs no MMIO write.\n+\t */\n+\tif (cmpl-\u003etail != head) {\n+\t\tcmpl-\u003etail = head;\n+\t\twriteq(head \u0026 cmpl-\u003esize_mask, cmpl-\u003edoorbell);\n+\t}\n+}\n+\n+static void mpnic_clean_tcq(struct mpnic_napi_vector *nv,\n+\t\t\t struct mpnic_q_triad *qt, int napi_budget)\n+{\n+\tstruct mpnic_ring *cmpl = \u0026qt-\u003ecmpl;\n+\t__le64 *raw_tcd, done;\n+\tu32 head = cmpl-\u003ehead;\n+\ts32 head0 = -1;\n+\n+\tdone = (head \u0026 (cmpl-\u003esize_mask + 1)) ? 0 : cpu_to_le64(MPNIC_TCD_DONE);\n+\traw_tcd = \u0026cmpl-\u003edesc[head \u0026 cmpl-\u003esize_mask];\n+\n+\t/* Walk the completion queue collecting the heads reported by NIC.\n+\t * Only the first work queue is enabled and no packet asks for a\n+\t * timestamp, so every completion is a plain head update and the\n+\t * descriptor type does not have to be decoded.\n+\t */\n+\twhile ((*raw_tcd \u0026 cpu_to_le64(MPNIC_TCD_DONE)) == done) {\n+\t\tu64 tcd;\n+\n+\t\tdma_rmb();\n+\n+\t\ttcd = le64_to_cpu(*raw_tcd);\n+\t\thead0 = FIELD_GET(MPNIC_TCD_TYPE0_HEAD0, tcd);\n+\n+\t\traw_tcd++;\n+\t\thead++;\n+\n+\t\tif (unlikely(!(head \u0026 cmpl-\u003esize_mask))) {\n+\t\t\tdone ^= cpu_to_le64(MPNIC_TCD_DONE);\n+\t\t\traw_tcd = \u0026cmpl-\u003edesc[0];\n+\t\t}\n+\t}\n+\n+\tcmpl-\u003ehead = head;\n+\n+\tif (head0 \u003e= 0)\n+\t\tmpnic_clean_twq0(nv, napi_budget, \u0026qt-\u003esub0, false, head0);\n+}\n+\n+static void mpnic_bd_prep(struct mpnic_ring *bdq, u32 idx, struct page *page)\n+{\n+\tdma_addr_t dma = page_pool_get_dma_addr(page);\n+\n+\tbdq-\u003edesc[idx] = cpu_to_le64(FIELD_PREP(MPNIC_BD_DESC_ADDR, dma \u003e\u003e 10) |\n+\t\t\t\t FIELD_PREP(MPNIC_BD_DESC_ID, idx) |\n+\t\t\t\t FIELD_PREP(MPNIC_BD_DESC_BUF_SZ_LOG2,\n+\t\t\t\t\t\tpage_shift(page) - 10));\n+}\n+\n+/* Descriptors are only handed to the device in whole batches, so the slot\n+ * the device is working on and everything up to the next batch boundary\n+ * stay untouched while it does.\n+ */\n+static unsigned int mpnic_bdq_desc_unused(struct mpnic_ring *bdq)\n+{\n+\treturn (ALIGN_DOWN(bdq-\u003ehead - 1, MPNIC_BDQ_BATCH_SIZE) - bdq-\u003etail) \u0026\n+\t bdq-\u003esize_mask;\n+}\n+\n+static unsigned int __mpnic_fill_bdq(struct mpnic_ring *bdq)\n+{\n+\tunsigned int i = bdq-\u003etail;\n+\tunsigned int count;\n+\n+\tfor (count = mpnic_bdq_desc_unused(bdq); count; count--) {\n+\t\tstruct page *page;\n+\n+\t\tpage = page_pool_dev_alloc_pages(bdq-\u003epage_pool);\n+\t\tif (!page)\n+\t\t\tbreak;\n+\n+\t\tbdq-\u003erx_buf[i] = page;\n+\t\tmpnic_bd_prep(bdq, i, page);\n+\n+\t\ti++;\n+\t\ti \u0026= bdq-\u003esize_mask;\n+\t}\n+\n+\treturn i;\n+}\n+\n+static void __mpnic_bdq_commit_tail(struct mpnic_ring *bdq, unsigned int tail)\n+{\n+\tif (bdq-\u003etail != tail) {\n+\t\tbdq-\u003etail = tail;\n+\n+\t\twriteq(tail, bdq-\u003edoorbell);\n+\t}\n+}\n+\n+static void mpnic_fill_qt_bdqs(struct mpnic_q_triad *qt)\n+{\n+\tunsigned int ppq_i = __mpnic_fill_bdq(\u0026qt-\u003esub1);\n+\tunsigned int hpq_i = __mpnic_fill_bdq(\u0026qt-\u003esub0);\n+\n+\t/* Force DMA writes to flush before writing to tail(s) */\n+\tdma_wmb();\n+\n+\t/* Flush out the completions we are done with */\n+\tmpnic_commit_cq_head(\u0026qt-\u003ecmpl);\n+\n+\t__mpnic_bdq_commit_tail(\u0026qt-\u003esub0, hpq_i);\n+\t__mpnic_bdq_commit_tail(\u0026qt-\u003esub1, ppq_i);\n+}\n+\n+/* Take one of the references batched on the page at @idx. If the device\n+ * has moved on to a new page, first drop the unused references left on\n+ * the previous one.\n+ */\n+static struct page *\n+mpnic_page_pool_get(struct mpnic_pg_ctxt *pg_ctxt, struct mpnic_ring *ring,\n+\t\t u32 idx)\n+{\n+\tstruct page *page = pg_ctxt-\u003epage;\n+\n+\tif (unlikely(pg_ctxt-\u003eidx != idx)) {\n+\t\tif (pg_ctxt-\u003epagecnt_bias \u0026\u0026\n+\t\t !page_pool_unref_page(page, pg_ctxt-\u003epagecnt_bias))\n+\t\t\tpage_pool_put_unrefed_page(page-\u003epp, page, -1, true);\n+\n+\t\tpage = ring-\u003erx_buf[idx];\n+\t\tpage_pool_fragment_page(page, MPNIC_PAGECNT_BIAS_MAX);\n+\n+\t\tpg_ctxt-\u003epage = page;\n+\t\tpg_ctxt-\u003epagecnt_bias = MPNIC_PAGECNT_BIAS_MAX;\n+\t\tpg_ctxt-\u003eidx = idx;\n+\t}\n+\n+\tpg_ctxt-\u003epagecnt_bias--;\n+\n+\treturn page;\n+}\n+\n+static void mpnic_flush_pg_ctxt(struct mpnic_pg_ctxt *ctxt, bool napi)\n+{\n+\tlong pagecnt_bias = ctxt-\u003epagecnt_bias;\n+\n+\tif (pagecnt_bias) {\n+\t\tstruct page *page = ctxt-\u003epage;\n+\n+\t\tif (!page_pool_unref_page(page, pagecnt_bias))\n+\t\t\tpage_pool_put_unrefed_page(page-\u003epp, page, -1, napi);\n+\t}\n+}\n+\n+static unsigned int mpnic_hdr_pg_start(unsigned int pg_off)\n+{\n+\t/* The headroom of the first header may be larger than\n+\t * MPNIC_RX_HROOM due to alignment. So account for that by just\n+\t * making the page offset 0 if we are starting at the first header.\n+\t */\n+\tif (ALIGN(MPNIC_RX_HROOM, 128) \u003e MPNIC_RX_HROOM \u0026\u0026\n+\t pg_off == ALIGN(MPNIC_RX_HROOM, 128))\n+\t\treturn 0;\n+\n+\treturn pg_off - MPNIC_RX_HROOM;\n+}\n+\n+static unsigned int mpnic_hdr_pg_end(unsigned int pg_off, unsigned int len)\n+{\n+\t/* Determine the end of the buffer by finding the start of the next\n+\t * and then subtracting the headroom from that frame.\n+\t */\n+\tpg_off += len + MPNIC_RX_TROOM + MPNIC_RX_HROOM;\n+\n+\treturn ALIGN(pg_off, 128) - MPNIC_RX_HROOM;\n+}\n+\n+static void\n+mpnic_pkt_prepare(struct mpnic_napi_vector *nv, u64 rcd,\n+\t\t struct mpnic_rcq_state *state, struct mpnic_q_triad *qt)\n+{\n+\tunsigned int pg_off = FIELD_GET(MPNIC_RCD_AL_BUFF_OFF, rcd);\n+\tunsigned int pg_idx = FIELD_GET(MPNIC_RCD_AL_BUFF_ID, rcd);\n+\tunsigned int len = FIELD_GET(MPNIC_RCD_AL_BUFF_LEN, rcd);\n+\tbool fin = FIELD_GET(MPNIC_RCD_AL_PAGE_FIN, rcd);\n+\tunsigned int frame_sz, pg_start, pg_end;\n+\tstruct xdp_buff *buff = \u0026state-\u003epkt;\n+\tstruct page *page;\n+\n+\tpg_start = mpnic_hdr_pg_start(pg_off);\n+\n+\tpage = mpnic_page_pool_get(\u0026state-\u003ehdr, \u0026qt-\u003esub0, pg_idx);\n+\tqt-\u003esub0.head = (pg_idx + 1) \u0026 qt-\u003esub0.size_mask;\n+\n+\t/* Short-cut the end calculation if the page is fully consumed */\n+\tpg_end = fin ? page_size(page) : mpnic_hdr_pg_end(pg_off, len);\n+\tframe_sz = pg_end - pg_start;\n+\n+\tdma_sync_single_range_for_cpu(nv-\u003edev, page_pool_get_dma_addr(page),\n+\t\t\t\t pg_start, frame_sz, DMA_FROM_DEVICE);\n+\n+\txdp_init_buff(buff, frame_sz, \u0026qt-\u003exdp_rxq);\n+\txdp_prepare_buff(buff, page_address(page) + pg_start,\n+\t\t\t pg_off - pg_start, len, true);\n+\tnet_prefetch(buff-\u003edata);\n+\n+\tstate-\u003eadd_frag_failed = false;\n+}\n+\n+static void\n+mpnic_add_rx_frag(struct mpnic_napi_vector *nv, u64 rcd,\n+\t\t struct mpnic_rcq_state *state, struct mpnic_q_triad *qt)\n+{\n+\tunsigned int pg_off = FIELD_GET(MPNIC_RCD_AL_BUFF_OFF, rcd);\n+\tunsigned int pg_idx = FIELD_GET(MPNIC_RCD_AL_BUFF_ID, rcd);\n+\tunsigned int len = FIELD_GET(MPNIC_RCD_AL_BUFF_LEN, rcd);\n+\tbool fin = FIELD_GET(MPNIC_RCD_AL_PAGE_FIN, rcd);\n+\tstruct xdp_buff *buff = \u0026state-\u003epkt;\n+\tunsigned int truesz;\n+\tstruct page *page;\n+\n+\tpage = mpnic_page_pool_get(\u0026state-\u003epayld, \u0026qt-\u003esub1, pg_idx);\n+\tqt-\u003esub1.head = (pg_idx + 1) \u0026 qt-\u003esub1.size_mask;\n+\n+\ttruesz = (fin ? page_size(page) : ALIGN(pg_off + len, 128)) - pg_off;\n+\n+\tdma_sync_single_range_for_cpu(nv-\u003edev, page_pool_get_dma_addr(page),\n+\t\t\t\t pg_off, truesz, DMA_FROM_DEVICE);\n+\n+\tif (!xdp_buff_add_frag(buff, page_to_netmem(page), pg_off, len,\n+\t\t\t truesz)) {\n+\t\tstate-\u003epayld.pagecnt_bias++;\n+\t\tstate-\u003eadd_frag_failed = true;\n+\t}\n+}\n+\n+static void mpnic_put_pkt_buff(struct xdp_buff *buff, bool napi)\n+{\n+\tstruct page *page;\n+\n+\tif (!buff-\u003edata_hard_start)\n+\t\treturn;\n+\n+\tif (unlikely(xdp_buff_has_frags(buff))) {\n+\t\tstruct skb_shared_info *shinfo;\n+\t\tint nr_frags;\n+\n+\t\tshinfo = xdp_get_shared_info_from_buff(buff);\n+\t\tnr_frags = shinfo-\u003enr_frags;\n+\n+\t\twhile (nr_frags--) {\n+\t\t\tpage = skb_frag_page(\u0026shinfo-\u003efrags[nr_frags]);\n+\t\t\tpage_pool_put_full_page(page-\u003epp, page, napi);\n+\t\t}\n+\t}\n+\n+\tpage = virt_to_head_page(buff-\u003edata_hard_start);\n+\tpage_pool_put_full_page(page-\u003epp, page, napi);\n+}\n+\n+static int mpnic_clean_rcq(struct mpnic_napi_vector *nv,\n+\t\t\t struct mpnic_q_triad *qt, int budget)\n+{\n+\tstruct mpnic_ring *rcq = \u0026qt-\u003ecmpl;\n+\tstruct mpnic_rcq_state *state;\n+\tunsigned int packets = 0;\n+\t__le64 *raw_rcd, done;\n+\tu32 head = rcq-\u003ehead;\n+\n+\tdone = (head \u0026 (rcq-\u003esize_mask + 1)) ? 0 : cpu_to_le64(MPNIC_RCD_DONE);\n+\traw_rcd = \u0026rcq-\u003edesc[head \u0026 rcq-\u003esize_mask];\n+\tstate = rcq-\u003estate;\n+\n+\twhile (packets \u003c budget) {\n+\t\tu64 rcd;\n+\n+\t\tif ((*raw_rcd \u0026 cpu_to_le64(MPNIC_RCD_DONE)) != done)\n+\t\t\tbreak;\n+\n+\t\tdma_rmb();\n+\n+\t\trcd = le64_to_cpu(*raw_rcd);\n+\n+\t\tswitch (FIELD_GET(MPNIC_RCD_TYPE, rcd)) {\n+\t\tcase MPNIC_RCD_TYPE_HDR_AL:\n+\t\t\tif (FIELD_GET(MPNIC_RCD_HDR_SUBTYPE, rcd) ==\n+\t\t\t MPNIC_RCD_HDR_SUBTYPE_HDR)\n+\t\t\t\tmpnic_pkt_prepare(nv, rcd, state, qt);\n+\t\t\tbreak;\n+\t\tcase MPNIC_RCD_TYPE_PAY_AL:\n+\t\t\tmpnic_add_rx_frag(nv, rcd, state, qt);\n+\t\t\tbreak;\n+\t\tcase MPNIC_RCD_TYPE_META: {\n+\t\t\tstruct sk_buff *skb = NULL;\n+\n+\t\t\tif (likely(!(rcd \u0026\n+\t\t\t\t MPNIC_RCD_META_UNCORRECTABLE_ERR_MASK) \u0026\u0026\n+\t\t\t\t !state-\u003eadd_frag_failed))\n+\t\t\t\tskb = xdp_build_skb_from_buff(\u0026state-\u003epkt);\n+\n+\t\t\tif (likely(skb))\n+\t\t\t\tnapi_gro_receive(\u0026nv-\u003enapi, skb);\n+\t\t\telse\n+\t\t\t\tmpnic_put_pkt_buff(\u0026state-\u003epkt, true);\n+\n+\t\t\tstate-\u003epkt.data_hard_start = NULL;\n+\t\t\tpackets++;\n+\t\t\tbreak;\n+\t\t}\n+\t\t}\n+\n+\t\traw_rcd++;\n+\t\thead++;\n+\n+\t\tif (unlikely(!(head \u0026 rcq-\u003esize_mask))) {\n+\t\t\tdone ^= cpu_to_le64(MPNIC_RCD_DONE);\n+\t\t\traw_rcd = \u0026rcq-\u003edesc[0];\n+\t\t}\n+\t}\n+\n+\trcq-\u003ehead = head;\n+\n+\t/* Allocate buffers, force dma_wmb(), and then start writing tails */\n+\tmpnic_fill_qt_bdqs(qt);\n+\n+\treturn packets;\n+}\n+\n+static int mpnic_poll(struct napi_struct *napi, int budget)\n+{\n+\tstruct mpnic_napi_vector *nv = container_of(napi,\n+\t\t\t\t\t\t struct mpnic_napi_vector,\n+\t\t\t\t\t\t napi);\n+\tint i, j, work_done = 0;\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++)\n+\t\tmpnic_clean_tcq(nv, \u0026nv-\u003eqt[i], budget);\n+\n+\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, i++)\n+\t\twork_done += mpnic_clean_rcq(nv, \u0026nv-\u003eqt[i], budget);\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++)\n+\t\tmpnic_commit_cq_head(\u0026nv-\u003eqt[i].cmpl);\n+\n+\tif (work_done \u003e= budget)\n+\t\treturn budget;\n+\n+\tif (likely(napi_complete_done(napi, work_done)))\n+\t\tmpnic_nv_irq_rearm(nv);\n+\n+\treturn work_done;\n+}\n+\n+static irqreturn_t mpnic_msix_clean_rings(int __always_unused irq, void *data)\n+{\n+\tstruct mpnic_napi_vector *nv = data;\n+\n+\tnapi_schedule_irqoff(\u0026nv-\u003enapi);\n+\n+\treturn IRQ_HANDLED;\n+}\n+\n+static void mpnic_free_napi_vector(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_napi_vector *nv)\n+{\n+\tint i, j;\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++)\n+\t\tmpn-\u003etx[nv-\u003eqt[i].sub0.q_idx] = NULL;\n+\n+\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, i++)\n+\t\tmpn-\u003erx[nv-\u003eqt[i].cmpl.q_idx] = NULL;\n+\n+\tmpnic_free_irq(nv-\u003empd, nv-\u003ev_idx, nv);\n+\tnetif_napi_del_locked(\u0026nv-\u003enapi);\n+\tmpn-\u003enapi[nv-\u003ev_idx - MPNIC_NON_NAPI_VECTORS] = NULL;\n+\tkfree(nv);\n+}\n+\n+void mpnic_free_napi_vectors(struct mpnic_net *mpn)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++)\n+\t\tif (mpn-\u003enapi[i])\n+\t\t\tmpnic_free_napi_vector(mpn, mpn-\u003enapi[i]);\n+}\n+\n+static void mpnic_ring_init(struct mpnic_ring *ring, u32 __iomem *doorbell,\n+\t\t\t int q_idx)\n+{\n+\tring-\u003edoorbell = doorbell;\n+\tring-\u003eq_idx = q_idx;\n+}\n+\n+static int mpnic_alloc_napi_vector(struct mpnic_dev *mpd,\n+\t\t\t\t struct mpnic_net *mpn, unsigned int idx)\n+{\n+\tu32 __iomem *uc_addr = READ_ONCE(mpd-\u003euc_addr0);\n+\tstruct mpnic_napi_vector *nv;\n+\tint err;\n+\n+\t/* Doorbells are plain pointers into the register window, they have\n+\t * no way of noticing that it went away.\n+\t */\n+\tif (!uc_addr)\n+\t\treturn -EIO;\n+\n+\tnv = kzalloc_flex(*nv, qt, 2);\n+\tif (!nv)\n+\t\treturn -ENOMEM;\n+\n+\tnv-\u003etxt_count = 1;\n+\tnv-\u003erxt_count = 1;\n+\tnv-\u003empd = mpd;\n+\tnv-\u003edev = mpd-\u003edev;\n+\tnv-\u003ev_idx = idx + MPNIC_NON_NAPI_VECTORS;\n+\n+\tmpn-\u003enapi[idx] = nv;\n+\tnetif_napi_add_config_locked(mpn-\u003enetdev, \u0026nv-\u003enapi, mpnic_poll, idx);\n+\tnetif_napi_set_irq_locked(\u0026nv-\u003enapi,\n+\t\t\t\t pci_irq_vector(to_pci_dev(mpd-\u003edev),\n+\t\t\t\t\t\t nv-\u003ev_idx));\n+\n+\tsnprintf(nv-\u003ename, sizeof(nv-\u003ename), \"%s-TxRx-%u\",\n+\t\t mpn-\u003enetdev-\u003ename, idx);\n+\n+\terr = mpnic_request_irq(mpd, nv-\u003ev_idx, mpnic_msix_clean_rings, 0,\n+\t\t\t\tnv-\u003ename, nv);\n+\tif (err)\n+\t\tgoto err_napi_del;\n+\n+\tmpnic_ring_init(\u0026nv-\u003eqt[0].sub0, \u0026uc_addr[MPNIC_TWQ_TAIL(idx, 0)], idx);\n+\tmpnic_ring_init(\u0026nv-\u003eqt[0].cmpl, \u0026uc_addr[MPNIC_TCQ_HEAD(idx)], idx);\n+\tmpn-\u003etx[idx] = \u0026nv-\u003eqt[0].sub0;\n+\n+\tmpnic_ring_init(\u0026nv-\u003eqt[1].sub0, \u0026uc_addr[MPNIC_HPQ_TAIL(idx)], idx);\n+\tmpnic_ring_init(\u0026nv-\u003eqt[1].sub1, \u0026uc_addr[MPNIC_PPQ_TAIL(idx)], idx);\n+\tmpnic_ring_init(\u0026nv-\u003eqt[1].cmpl, \u0026uc_addr[MPNIC_RCQ_HEAD(idx)], idx);\n+\tmpn-\u003erx[idx] = \u0026nv-\u003eqt[1].cmpl;\n+\n+\treturn 0;\n+\n+err_napi_del:\n+\tnetif_napi_del_locked(\u0026nv-\u003enapi);\n+\tmpn-\u003enapi[idx] = NULL;\n+\tkfree(nv);\n+\treturn err;\n+}\n+\n+int mpnic_alloc_napi_vectors(struct mpnic_net *mpn)\n+{\n+\tunsigned int i;\n+\tint err;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\terr = mpnic_alloc_napi_vector(mpn-\u003empd, mpn, i);\n+\t\tif (err)\n+\t\t\tgoto err_free_vectors;\n+\t}\n+\n+\treturn 0;\n+\n+err_free_vectors:\n+\tmpnic_free_napi_vectors(mpn);\n+\n+\treturn err;\n+}\n+\n+static void mpnic_free_ring_resources(struct device *dev,\n+\t\t\t\t struct mpnic_ring *ring)\n+{\n+\tkvfree(ring-\u003ebuffer);\n+\tring-\u003ebuffer = NULL;\n+\n+\t/* If size is not set there are no descriptors present */\n+\tif (!ring-\u003esize)\n+\t\treturn;\n+\n+\tdma_free_coherent(dev, ring-\u003esize, ring-\u003edesc, ring-\u003edma);\n+\tring-\u003esize_mask = 0;\n+\tring-\u003esize = 0;\n+}\n+\n+static int mpnic_alloc_ring_desc(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_ring *ring, u32 count)\n+{\n+\tstruct device *dev = mpn-\u003enetdev-\u003edev.parent;\n+\tsize_t size;\n+\n+\tsize = ALIGN(array_size(sizeof(*ring-\u003edesc), count), 4096);\n+\n+\tring-\u003edesc = dma_alloc_coherent(dev, size, \u0026ring-\u003edma,\n+\t\t\t\t\tGFP_KERNEL | __GFP_NOWARN);\n+\tif (!ring-\u003edesc)\n+\t\treturn -ENOMEM;\n+\n+\tring-\u003esize_mask = count - 1;\n+\tring-\u003esize = size;\n+\n+\treturn 0;\n+}\n+\n+static void mpnic_free_tx_qt_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_q_triad *qt)\n+{\n+\tstruct device *dev = mpn-\u003enetdev-\u003edev.parent;\n+\n+\tmpnic_free_ring_resources(dev, \u0026qt-\u003ecmpl);\n+\tmpnic_free_ring_resources(dev, \u0026qt-\u003esub0);\n+}\n+\n+static int mpnic_alloc_tx_qt_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_q_triad *qt)\n+{\n+\tint err;\n+\n+\terr = mpnic_alloc_ring_desc(mpn, \u0026qt-\u003esub0, mpn-\u003etxq_size);\n+\tif (err)\n+\t\treturn err;\n+\n+\tqt-\u003esub0.tx_buf = kvzalloc_objs(*qt-\u003esub0.tx_buf, mpn-\u003etxq_size,\n+\t\t\t\t\tGFP_KERNEL | __GFP_NOWARN);\n+\tif (!qt-\u003esub0.tx_buf) {\n+\t\terr = -ENOMEM;\n+\t\tgoto err_free_qt;\n+\t}\n+\n+\terr = mpnic_alloc_ring_desc(mpn, \u0026qt-\u003ecmpl, mpn-\u003etxq_size);\n+\tif (err)\n+\t\tgoto err_free_qt;\n+\n+\treturn 0;\n+\n+err_free_qt:\n+\tmpnic_free_tx_qt_resources(mpn, qt);\n+\treturn err;\n+}\n+\n+static int\n+mpnic_alloc_qt_page_pool(struct mpnic_net *mpn, struct mpnic_napi_vector *nv,\n+\t\t\t struct mpnic_q_triad *qt)\n+{\n+\tstruct page_pool_params pp_params = {\n+\t\t.flags\t\t= PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV,\n+\t\t.pool_size\t= min(mpn-\u003ehpq_size + mpn-\u003eppq_size, 32768u),\n+\t\t.nid\t\t= NUMA_NO_NODE,\n+\t\t.dev\t\t= nv-\u003edev,\n+\t\t.dma_dir\t= DMA_FROM_DEVICE,\n+\t\t.max_len\t= PAGE_SIZE,\n+\t\t.napi\t\t= \u0026nv-\u003enapi,\n+\t\t.netdev\t\t= mpn-\u003enetdev,\n+\t\t.queue_idx\t= qt-\u003ecmpl.q_idx,\n+\t};\n+\tstruct page_pool *pp;\n+\n+\tpp = page_pool_create(\u0026pp_params);\n+\tif (IS_ERR(pp))\n+\t\treturn PTR_ERR(pp);\n+\n+\tqt-\u003esub0.page_pool = pp;\n+\tpage_pool_get(pp);\n+\tqt-\u003esub1.page_pool = pp;\n+\n+\treturn 0;\n+}\n+\n+static void mpnic_free_rx_qt_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_q_triad *qt)\n+{\n+\tstruct device *dev = mpn-\u003enetdev-\u003edev.parent;\n+\n+\tmpnic_free_ring_resources(dev, \u0026qt-\u003ecmpl);\n+\tmpnic_free_ring_resources(dev, \u0026qt-\u003esub1);\n+\tmpnic_free_ring_resources(dev, \u0026qt-\u003esub0);\n+\n+\tif (xdp_rxq_info_is_reg(\u0026qt-\u003exdp_rxq)) {\n+\t\txdp_rxq_info_unreg(\u0026qt-\u003exdp_rxq);\n+\t\tpage_pool_destroy(qt-\u003esub1.page_pool);\n+\t\tpage_pool_destroy(qt-\u003esub0.page_pool);\n+\t}\n+}\n+\n+static int mpnic_alloc_rx_qt_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_napi_vector *nv,\n+\t\t\t\t struct mpnic_q_triad *qt)\n+{\n+\tint err;\n+\n+\terr = mpnic_alloc_qt_page_pool(mpn, nv, qt);\n+\tif (err)\n+\t\treturn err;\n+\n+\terr = xdp_rxq_info_reg(\u0026qt-\u003exdp_rxq, mpn-\u003enetdev, qt-\u003ecmpl.q_idx,\n+\t\t\t nv-\u003enapi.napi_id);\n+\tif (err)\n+\t\tgoto err_free_page_pool;\n+\n+\terr = xdp_rxq_info_reg_mem_model(\u0026qt-\u003exdp_rxq, MEM_TYPE_PAGE_POOL,\n+\t\t\t\t\t qt-\u003esub0.page_pool);\n+\tif (err)\n+\t\tgoto err_unreg_rxq;\n+\n+\terr = mpnic_alloc_ring_desc(mpn, \u0026qt-\u003esub0, mpn-\u003ehpq_size);\n+\tif (err)\n+\t\tgoto err_unreg_mm;\n+\n+\tqt-\u003esub0.rx_buf = kvzalloc_objs(*qt-\u003esub0.rx_buf, mpn-\u003ehpq_size,\n+\t\t\t\t\tGFP_KERNEL | __GFP_NOWARN);\n+\tif (!qt-\u003esub0.rx_buf) {\n+\t\terr = -ENOMEM;\n+\t\tgoto err_free_qt;\n+\t}\n+\n+\terr = mpnic_alloc_ring_desc(mpn, \u0026qt-\u003esub1, mpn-\u003eppq_size);\n+\tif (err)\n+\t\tgoto err_free_qt;\n+\n+\tqt-\u003esub1.rx_buf = kvzalloc_objs(*qt-\u003esub1.rx_buf, mpn-\u003eppq_size,\n+\t\t\t\t\tGFP_KERNEL | __GFP_NOWARN);\n+\tif (!qt-\u003esub1.rx_buf) {\n+\t\terr = -ENOMEM;\n+\t\tgoto err_free_qt;\n+\t}\n+\n+\terr = mpnic_alloc_ring_desc(mpn, \u0026qt-\u003ecmpl, mpn-\u003ercq_size);\n+\tif (err)\n+\t\tgoto err_free_qt;\n+\n+\tqt-\u003ecmpl.state = kvzalloc_obj(*qt-\u003ecmpl.state,\n+\t\t\t\t GFP_KERNEL | __GFP_NOWARN);\n+\tif (!qt-\u003ecmpl.state) {\n+\t\terr = -ENOMEM;\n+\t\tgoto err_free_qt;\n+\t}\n+\n+\treturn 0;\n+\n+err_free_qt:\n+\tmpnic_free_rx_qt_resources(mpn, qt);\n+\treturn err;\n+err_unreg_mm:\n+\txdp_rxq_info_unreg_mem_model(\u0026qt-\u003exdp_rxq);\n+err_unreg_rxq:\n+\txdp_rxq_info_unreg(\u0026qt-\u003exdp_rxq);\n+err_free_page_pool:\n+\tpage_pool_destroy(qt-\u003esub1.page_pool);\n+\tpage_pool_destroy(qt-\u003esub0.page_pool);\n+\treturn err;\n+}\n+\n+static void mpnic_free_nv_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_napi_vector *nv)\n+{\n+\tint i, j;\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++)\n+\t\tmpnic_free_tx_qt_resources(mpn, \u0026nv-\u003eqt[i]);\n+\n+\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, i++)\n+\t\tmpnic_free_rx_qt_resources(mpn, \u0026nv-\u003eqt[i]);\n+}\n+\n+static int mpnic_alloc_nv_resources(struct mpnic_net *mpn,\n+\t\t\t\t struct mpnic_napi_vector *nv)\n+{\n+\tint i, j, err;\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++) {\n+\t\terr = mpnic_alloc_tx_qt_resources(mpn, \u0026nv-\u003eqt[i]);\n+\t\tif (err)\n+\t\t\tgoto err_free_qt_resources;\n+\t}\n+\n+\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, i++) {\n+\t\terr = mpnic_alloc_rx_qt_resources(mpn, nv, \u0026nv-\u003eqt[i]);\n+\t\tif (err)\n+\t\t\tgoto err_free_qt_resources;\n+\t}\n+\n+\treturn 0;\n+\n+err_free_qt_resources:\n+\twhile (i--) {\n+\t\tif (i \u003c nv-\u003etxt_count)\n+\t\t\tmpnic_free_tx_qt_resources(mpn, \u0026nv-\u003eqt[i]);\n+\t\telse\n+\t\t\tmpnic_free_rx_qt_resources(mpn, \u0026nv-\u003eqt[i]);\n+\t}\n+\treturn err;\n+}\n+\n+void mpnic_free_resources(struct mpnic_net *mpn)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++)\n+\t\tmpnic_free_nv_resources(mpn, mpn-\u003enapi[i]);\n+}\n+\n+int mpnic_alloc_resources(struct mpnic_net *mpn)\n+{\n+\tint i, err;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\terr = mpnic_alloc_nv_resources(mpn, mpn-\u003enapi[i]);\n+\t\tif (err)\n+\t\t\tgoto err_free_resources;\n+\t}\n+\n+\treturn 0;\n+\n+err_free_resources:\n+\twhile (i--)\n+\t\tmpnic_free_nv_resources(mpn, mpn-\u003enapi[i]);\n+\n+\treturn err;\n+}\n+\n+static void mpnic_set_netif_napi(struct mpnic_napi_vector *nv,\n+\t\t\t\t struct napi_struct *napi)\n+{\n+\tint i, j;\n+\n+\tfor (i = 0; i \u003c nv-\u003etxt_count; i++)\n+\t\tnetif_queue_set_napi(nv-\u003enapi.dev, nv-\u003eqt[i].sub0.q_idx,\n+\t\t\t\t NETDEV_QUEUE_TYPE_TX, napi);\n+\n+\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, i++)\n+\t\tnetif_queue_set_napi(nv-\u003enapi.dev, nv-\u003eqt[i].cmpl.q_idx,\n+\t\t\t\t NETDEV_QUEUE_TYPE_RX, napi);\n+}\n+\n+int mpnic_set_netif_queues(struct mpnic_net *mpn)\n+{\n+\tint i, err;\n+\n+\terr = netif_set_real_num_queues(mpn-\u003enetdev, mpn-\u003enum_tx_queues,\n+\t\t\t\t\tmpn-\u003enum_rx_queues);\n+\tif (err)\n+\t\treturn err;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++)\n+\t\tmpnic_set_netif_napi(mpn-\u003enapi[i], \u0026mpn-\u003enapi[i]-\u003enapi);\n+\n+\treturn 0;\n+}\n+\n+void mpnic_reset_netif_queues(struct mpnic_net *mpn)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++)\n+\t\tmpnic_set_netif_napi(mpn-\u003enapi[i], NULL);\n+}\n+\n+static void mpnic_enable_twq(struct mpnic_dev *mpd, struct mpnic_ring *twq)\n+{\n+\tu32 log_size = fls(twq-\u003esize_mask);\n+\tu32 i = twq-\u003eq_idx;\n+\n+\t/* Reset head/tail */\n+\tmpnic_wr64(mpd, MPNIC_TWQ_CTL(i, 0), MPNIC_TWQ_CTL_RESET);\n+\ttwq-\u003etail = 0;\n+\ttwq-\u003ehead = 0;\n+\ttwq-\u003edeferred_meta = -1;\n+\n+\t/* Store descriptor ring address and size */\n+\tmpnic_wr64(mpd, MPNIC_TWQ_BASE_ADDR(i, 0), twq-\u003edma);\n+\tmpnic_wr64(mpd, MPNIC_TWQ_SIZE(i, 0), log_size \u0026 MPNIC_TWQ_SIZE_SIZE);\n+\n+\tmpnic_wr64(mpd, MPNIC_TWQ_CTL(i, 0), MPNIC_TWQ_CTL_ENABLE);\n+}\n+\n+static void mpnic_enable_tcq(struct mpnic_dev *mpd,\n+\t\t\t struct mpnic_napi_vector *nv,\n+\t\t\t struct mpnic_ring *tcq)\n+{\n+\tu32 log_size = fls(tcq-\u003esize_mask);\n+\tu32 i = tcq-\u003eq_idx;\n+\n+\t/* Reset head/tail */\n+\tmpnic_wr64(mpd, MPNIC_TCQ_CTL(i), MPNIC_TCQ_CTL_RESET);\n+\ttcq-\u003etail = 0;\n+\ttcq-\u003ehead = 0;\n+\n+\t/* Store descriptor ring address and size */\n+\tmpnic_wr64(mpd, MPNIC_TCQ_BASE_ADDR(i), tcq-\u003edma);\n+\tmpnic_wr64(mpd, MPNIC_TCQ_SIZE(i), log_size \u0026 MPNIC_TCQ_SIZE_SIZE);\n+\n+\t/* Store interrupt information for the completion queue */\n+\tmpnic_wr64(mpd, MPNIC_TIM_CTL(i), nv-\u003ev_idx);\n+\tmpnic_wr64(mpd, MPNIC_TIM_INTR_MASK(i), 0);\n+\n+\tmpnic_wr64(mpd, MPNIC_TCQ_CTL(i), MPNIC_TCQ_CTL_ENABLE);\n+}\n+\n+static void mpnic_enable_bdq(struct mpnic_dev *mpd, struct mpnic_ring *hpq,\n+\t\t\t struct mpnic_ring *ppq)\n+{\n+\tu32 hpq_log_size = fls(hpq-\u003esize_mask);\n+\tu32 ppq_log_size = fls(ppq-\u003esize_mask);\n+\tu32 i = hpq-\u003eq_idx;\n+\n+\t/* Reset head/tail */\n+\tmpnic_wr64(mpd, MPNIC_BDQ_CTL(i), MPNIC_BDQ_CTL_RESET);\n+\thpq-\u003etail = 0;\n+\thpq-\u003ehead = 0;\n+\tppq-\u003etail = 0;\n+\tppq-\u003ehead = 0;\n+\n+\t/* Store descriptor ring addresses and sizes */\n+\tmpnic_wr64(mpd, MPNIC_HPQ_BASE_ADDR(i), hpq-\u003edma);\n+\tmpnic_wr64(mpd, MPNIC_HPQ_SIZE(i), hpq_log_size \u0026 MPNIC_HPQ_SIZE_SIZE);\n+\tmpnic_wr64(mpd, MPNIC_PPQ_BASE_ADDR(i), ppq-\u003edma);\n+\tmpnic_wr64(mpd, MPNIC_PPQ_SIZE(i), ppq_log_size \u0026 MPNIC_PPQ_SIZE_SIZE);\n+\n+\tmpnic_wr64(mpd, MPNIC_BDQ_CTL(i),\n+\t\t MPNIC_BDQ_CTL_ENABLE | MPNIC_BDQ_CTL_ENABLE_PPQ);\n+}\n+\n+static void mpnic_set_rde_cfg(struct mpnic_dev *mpd, struct mpnic_ring *rcq)\n+{\n+\tBUILD_BUG_ON(FIELD_MAX(MPNIC_RDE_CFG_MIN_HEAD_ROOM) \u003c MPNIC_RX_HROOM);\n+\tBUILD_BUG_ON(FIELD_MAX(MPNIC_RDE_CFG_MIN_TAIL_ROOM) \u003c MPNIC_RX_TROOM);\n+\n+\tmpnic_wr64(mpd, MPNIC_RDE_CFG(rcq-\u003eq_idx),\n+\t\t FIELD_PREP(MPNIC_RDE_CFG_MIN_HEAD_ROOM, MPNIC_RX_HROOM) |\n+\t\t FIELD_PREP(MPNIC_RDE_CFG_MIN_TAIL_ROOM, MPNIC_RX_TROOM) |\n+\t\t FIELD_PREP(MPNIC_RDE_CFG_MAX_HEADER_BYTES,\n+\t\t\t MPNIC_RX_MAX_HDR));\n+}\n+\n+static void mpnic_enable_rcq(struct mpnic_dev *mpd,\n+\t\t\t struct mpnic_napi_vector *nv,\n+\t\t\t struct mpnic_ring *rcq)\n+{\n+\tu32 log_size = fls(rcq-\u003esize_mask);\n+\tu32 i = rcq-\u003eq_idx;\n+\n+\tmpnic_set_rde_cfg(mpd, rcq);\n+\n+\t/* Reset head/tail */\n+\tmpnic_wr64(mpd, MPNIC_RCQ_CTL(i), MPNIC_RCQ_CTL_RESET);\n+\trcq-\u003ehead = 0;\n+\trcq-\u003etail = 0;\n+\n+\t/* Store descriptor ring address and size */\n+\tmpnic_wr64(mpd, MPNIC_RCQ_BASE_ADDR(i), rcq-\u003edma);\n+\tmpnic_wr64(mpd, MPNIC_RCQ_SIZE(i), log_size \u0026 MPNIC_RCQ_SIZE_SIZE);\n+\n+\t/* Store interrupt information for the completion queue */\n+\tmpnic_wr64(mpd, MPNIC_RIM_CTL(i), nv-\u003ev_idx);\n+\tmpnic_wr64(mpd, MPNIC_RIM_INTR_MASK(i), 0);\n+\n+\tmpnic_wr64(mpd, MPNIC_RCQ_CTL(i), MPNIC_RCQ_CTL_ENABLE);\n+}\n+\n+void mpnic_enable(struct mpnic_net *mpn)\n+{\n+\tstruct mpnic_dev *mpd = mpn-\u003empd;\n+\tint i, j, t;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tstruct mpnic_napi_vector *nv = mpn-\u003enapi[i];\n+\n+\t\tfor (t = 0; t \u003c nv-\u003etxt_count; t++) {\n+\t\t\tmpnic_enable_twq(mpd, \u0026nv-\u003eqt[t].sub0);\n+\t\t\tmpnic_enable_tcq(mpd, nv, \u0026nv-\u003eqt[t].cmpl);\n+\t\t}\n+\n+\t\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, t++) {\n+\t\t\tmpnic_enable_bdq(mpd, \u0026nv-\u003eqt[t].sub0, \u0026nv-\u003eqt[t].sub1);\n+\t\t\tmpnic_enable_rcq(mpd, nv, \u0026nv-\u003eqt[t].cmpl);\n+\t\t}\n+\t}\n+\n+\tmpnic_wrfl(mpd);\n+}\n+\n+static void mpnic_disable_twq(struct mpnic_dev *mpd, struct mpnic_ring *txr)\n+{\n+\tu64 twq_ctl = mpnic_rd64(mpd, MPNIC_TWQ_CTL(txr-\u003eq_idx, 0));\n+\n+\ttwq_ctl \u0026= ~MPNIC_TWQ_CTL_ENABLE;\n+\tmpnic_wr64(mpd, MPNIC_TWQ_CTL(txr-\u003eq_idx, 0), twq_ctl);\n+}\n+\n+static void mpnic_disable_tcq(struct mpnic_dev *mpd, struct mpnic_ring *txr)\n+{\n+\tmpnic_wr64(mpd, MPNIC_TCQ_CTL(txr-\u003eq_idx), 0);\n+\tmpnic_wr64(mpd, MPNIC_TIM_INTR_MASK(txr-\u003eq_idx),\n+\t\t MPNIC_TIM_INTR_MASK_MASK);\n+}\n+\n+static void mpnic_disable_bdq(struct mpnic_dev *mpd, struct mpnic_ring *hpq)\n+{\n+\tu64 bdq_ctl = mpnic_rd64(mpd, MPNIC_BDQ_CTL(hpq-\u003eq_idx));\n+\n+\tbdq_ctl \u0026= ~(MPNIC_BDQ_CTL_ENABLE | MPNIC_BDQ_CTL_ENABLE_PPQ);\n+\tmpnic_wr64(mpd, MPNIC_BDQ_CTL(hpq-\u003eq_idx), bdq_ctl);\n+}\n+\n+static void mpnic_disable_rcq(struct mpnic_dev *mpd, struct mpnic_ring *rcq)\n+{\n+\tmpnic_wr64(mpd, MPNIC_RCQ_CTL(rcq-\u003eq_idx), 0);\n+\tmpnic_wr64(mpd, MPNIC_RIM_INTR_MASK(rcq-\u003eq_idx),\n+\t\t MPNIC_RIM_INTR_MASK_MASK);\n+}\n+\n+void mpnic_disable(struct mpnic_net *mpn)\n+{\n+\tstruct mpnic_dev *mpd = mpn-\u003empd;\n+\tint i, j, t;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tstruct mpnic_napi_vector *nv = mpn-\u003enapi[i];\n+\n+\t\tfor (t = 0; t \u003c nv-\u003etxt_count; t++) {\n+\t\t\tmpnic_disable_twq(mpd, \u0026nv-\u003eqt[t].sub0);\n+\t\t\tmpnic_disable_tcq(mpd, \u0026nv-\u003eqt[t].cmpl);\n+\t\t}\n+\n+\t\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, t++) {\n+\t\t\tmpnic_disable_bdq(mpd, \u0026nv-\u003eqt[t].sub0);\n+\t\t\tmpnic_disable_rcq(mpd, \u0026nv-\u003eqt[t].cmpl);\n+\t\t}\n+\t}\n+\n+\tmpnic_wrfl(mpd);\n+}\n+\n+struct mpnic_idle_regs {\n+\tu32 reg_base;\n+\tu8 reg_cnt;\n+\tchar name[4];\n+};\n+\n+static u32 mpnic_non_idle_queues(struct mpnic_dev *mpd,\n+\t\t\t\t const struct mpnic_idle_regs *regs,\n+\t\t\t\t unsigned int nregs)\n+{\n+\tu32 non_idle_bitmap = 0;\n+\tunsigned int i, j;\n+\n+\tfor (i = 0; i \u003c nregs; i++) {\n+\t\tfor (j = 0; j \u003c regs[i].reg_cnt; j++) {\n+\t\t\tif (mpnic_rd64(mpd, regs[i].reg_base + 2 * j) !=\n+\t\t\t ~0ULL) {\n+\t\t\t\tnon_idle_bitmap |= BIT(i);\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t}\n+\t}\n+\n+\treturn non_idle_bitmap;\n+}\n+\n+static void mpnic_idle_dump(struct mpnic_dev *mpd,\n+\t\t\t const struct mpnic_idle_regs *regs,\n+\t\t\t unsigned int nregs, u32 non_idle_bitmap, int err)\n+{\n+\tunsigned int i, j;\n+\n+\tdev_err(mpd-\u003edev, \"error waiting for queues idle %d\\n\", err);\n+\tfor (i = 0; i \u003c nregs; i++) {\n+\t\tif (!(non_idle_bitmap \u0026 BIT(i)))\n+\t\t\tcontinue;\n+\n+\t\tdev_err(mpd-\u003edev, \"%s block not idle:\\n\", regs[i].name);\n+\t\tfor (j = 0; j \u003c regs[i].reg_cnt; j++)\n+\t\t\tdev_err(mpd-\u003edev, \" 0x%04x: %016llx\\n\",\n+\t\t\t\tregs[i].reg_base + 2 * j,\n+\t\t\t\tmpnic_rd64(mpd, regs[i].reg_base + 2 * j));\n+\t}\n+}\n+\n+void mpnic_wait_all_queues_idle(struct mpnic_dev *mpd)\n+{\n+\tstatic const struct mpnic_idle_regs queues[] = {\n+\t\t{ MPNIC_TWQ_IDLE(0), MPNIC_TWQ_IDLE_CNT, \"TWQ\" },\n+\t\t{ MPNIC_TQS_IDLE(0), MPNIC_TQS_IDLE_CNT, \"TQS\" },\n+\t\t{ MPNIC_TDE_IDLE(0), MPNIC_TDE_IDLE_CNT, \"TDE\" },\n+\t\t{ MPNIC_TCQ_IDLE(0), MPNIC_TCQ_IDLE_CNT, \"TCQ\" },\n+\t\t{ MPNIC_HPQ_IDLE(0), MPNIC_HPQ_IDLE_CNT, \"HPQ\" },\n+\t\t{ MPNIC_PPQ_IDLE(0), MPNIC_PPQ_IDLE_CNT, \"PPQ\" },\n+\t\t{ MPNIC_RCQ_IDLE(0), MPNIC_RCQ_IDLE_CNT, \"RCQ\" },\n+\t};\n+\tu32 non_idle_bitmap;\n+\tint err;\n+\n+\terr = read_poll_timeout(mpnic_non_idle_queues, non_idle_bitmap,\n+\t\t\t\t!non_idle_bitmap, 20, 500000, false, mpd,\n+\t\t\t\tqueues, ARRAY_SIZE(queues));\n+\tif (err)\n+\t\tmpnic_idle_dump(mpd, queues, ARRAY_SIZE(queues),\n+\t\t\t\tnon_idle_bitmap, err);\n+}\n+\n+static void mpnic_clean_bdq(struct mpnic_ring *bdq)\n+{\n+\tunsigned int head = bdq-\u003ehead;\n+\n+\twhile (head != bdq-\u003etail) {\n+\t\tstruct page *page = bdq-\u003erx_buf[head];\n+\n+\t\tpage_pool_put_full_page(page-\u003epp, page, false);\n+\n+\t\thead++;\n+\t\thead \u0026= bdq-\u003esize_mask;\n+\t}\n+\n+\tbdq-\u003ehead = head;\n+}\n+\n+void mpnic_flush(struct mpnic_net *mpn)\n+{\n+\tint i, j, t;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tstruct mpnic_napi_vector *nv = mpn-\u003enapi[i];\n+\n+\t\tfor (t = 0; t \u003c nv-\u003etxt_count; t++) {\n+\t\t\tstruct mpnic_q_triad *qt = \u0026nv-\u003eqt[t];\n+\t\t\tstruct netdev_queue *txq;\n+\n+\t\t\t/* Clean the work queue of unprocessed work */\n+\t\t\tmpnic_clean_twq0(nv, 0, \u0026qt-\u003esub0, true, qt-\u003esub0.tail);\n+\n+\t\t\ttxq = netdev_get_tx_queue(mpn-\u003enetdev, qt-\u003esub0.q_idx);\n+\t\t\tnetdev_tx_reset_queue(txq);\n+\t\t}\n+\n+\t\tfor (j = 0; j \u003c nv-\u003erxt_count; j++, t++) {\n+\t\t\tstruct mpnic_q_triad *qt = \u0026nv-\u003eqt[t];\n+\t\t\tstruct mpnic_rcq_state *state = qt-\u003ecmpl.state;\n+\n+\t\t\t/* Release the partially assembled frame and the\n+\t\t\t * pages the queues are still handing out.\n+\t\t\t */\n+\t\t\tmpnic_put_pkt_buff(\u0026state-\u003epkt, false);\n+\t\t\tmpnic_flush_pg_ctxt(\u0026state-\u003ehdr, false);\n+\t\t\tmpnic_flush_pg_ctxt(\u0026state-\u003epayld, false);\n+\t\t\tmemset(state, 0, sizeof(*state));\n+\n+\t\t\tmpnic_clean_bdq(\u0026qt-\u003esub0);\n+\t\t\tmpnic_clean_bdq(\u0026qt-\u003esub1);\n+\t\t}\n+\t}\n+}\n+\n+void mpnic_fill(struct mpnic_net *mpn)\n+{\n+\tint i, j, t;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tstruct mpnic_napi_vector *nv = mpn-\u003enapi[i];\n+\n+\t\tfor (j = 0, t = nv-\u003etxt_count; j \u003c nv-\u003erxt_count; j++, t++) {\n+\t\t\tstruct mpnic_q_triad *qt = \u0026nv-\u003eqt[t];\n+\t\t\tstruct mpnic_rcq_state *state = qt-\u003ecmpl.state;\n+\n+\t\t\t/* Point the page contexts at an index the device\n+\t\t\t * cannot report, so the first buffer coming out of\n+\t\t\t * either queue is not taken for a page we hold.\n+\t\t\t */\n+\t\t\tstate-\u003ehdr.idx = UINT_MAX;\n+\t\t\tstate-\u003epayld.idx = UINT_MAX;\n+\n+\t\t\tmpnic_fill_qt_bdqs(qt);\n+\t\t}\n+\t}\n+}\n+\n+void mpnic_napi_disable(struct mpnic_net *mpn)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tnapi_disable_locked(\u0026mpn-\u003enapi[i]-\u003enapi);\n+\n+\t\tmpnic_nv_irq_disable(mpn-\u003enapi[i]);\n+\t}\n+}\n+\n+void mpnic_napi_enable(struct mpnic_net *mpn)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++)\n+\t\tnapi_enable_locked(\u0026mpn-\u003enapi[i]-\u003enapi);\n+\n+\t/* Force the first interrupt on each vector to guarantee that any\n+\t * completions posted during bringup are processed. Use the TRIGGER\n+\t * pulse rather than the level triggered global interrupt set, which\n+\t * can jam the mask/pending state machine if it collides with a\n+\t * concurrent unmask.\n+\t */\n+\tfor (i = 0; i \u003c mpn-\u003enum_napi; i++) {\n+\t\tstruct mpnic_napi_vector *nv = mpn-\u003enapi[i];\n+\n+\t\tmpnic_wr64(mpn-\u003empd, MPNIC_TIM_CTL1(nv-\u003eqt[0].cmpl.q_idx),\n+\t\t\t MPNIC_TIM_PARAM_CFG_PRESERVE_MASK |\n+\t\t\t MPNIC_TIM_CTL1_MASK_EN | MPNIC_TIM_CTL1_TRIGGER);\n+\t}\n+\n+\tmpnic_wrfl(mpn-\u003empd);\n+}\ndiff --git a/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h\nnew file mode 100644\nindex 0000000000000..936ad791a3466\n--- /dev/null\n+++ b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h\n@@ -0,0 +1,147 @@\n+/* SPDX-License-Identifier: GPL-2.0 */\n+/* Copyright (c) Meta Platforms, Inc. and affiliates. */\n+\n+#ifndef _MPNIC_TXRX_H_\n+#define _MPNIC_TXRX_H_\n+\n+#include \u003clinux/if_ether.h\u003e\n+#include \u003clinux/netdevice.h\u003e\n+#include \u003clinux/skbuff.h\u003e\n+#include \u003clinux/types.h\u003e\n+#include \u003cnet/netdev_queues.h\u003e\n+#include \u003cnet/xdp.h\u003e\n+\n+#include \"mpnic.h\"\n+\n+struct mpnic_net;\n+\n+/* Space we have to have available in a work queue to take a packet:\n+ *\t1 descriptor per page\n+ *\t+ 1 descriptor for the skb head\n+ *\t+ 1 descriptor for the metadata\n+ *\t+ 7 descriptors to keep the tail out of the head's cacheline\n+ * If we cannot guarantee that we return NETDEV_TX_BUSY.\n+ */\n+#define MPNIC_MAX_SKB_DESC\t\t(MAX_SKB_FRAGS + 9)\n+#define MPNIC_TX_DESC_WAKEUP\t\t(MPNIC_MAX_SKB_DESC * 2)\n+\n+#define MPNIC_MAX_NAPI_VECTORS\t\t1024u\n+\n+/* Number of buffer descriptors the driver posts before ringing the\n+ * doorbell. The device consumes whatever the doorbell points at, this is\n+ * purely to keep the driver from writing the CSR for every descriptor.\n+ */\n+#define MPNIC_BDQ_BATCH_SIZE\t\t64u\n+\n+#define MPNIC_TXQ_SIZE_DEFAULT\t\t1024\n+#define MPNIC_HPQ_SIZE_DEFAULT\t\t256\n+#define MPNIC_PPQ_SIZE_DEFAULT\t\t256\n+#define MPNIC_RCQ_SIZE_DEFAULT\t\t1024\n+\n+/* Room the device has to leave in front of and behind every header so the\n+ * driver can build an skb around it in place. The headroom is padded out\n+ * so that consecutive headers in one page start 128 B aligned.\n+ */\n+#define MPNIC_RX_TROOM \\\n+\tSKB_DATA_ALIGN(sizeof(struct skb_shared_info))\n+#define MPNIC_RX_HROOM \\\n+\t(ALIGN(MPNIC_RX_TROOM + XDP_PACKET_HEADROOM, 128) - MPNIC_RX_TROOM)\n+\n+/* Headers longer than this are split off into the payload queue */\n+#define MPNIC_RX_MAX_HDR\t\t1536\n+\n+/* A page is handed out to many packets, each of which takes one reference.\n+ * Rather than a locked increment per packet the driver takes a batch of\n+ * references up front and returns whatever is left when the page is done.\n+ */\n+#define MPNIC_PAGECNT_BIAS_MAX\t\t(PAGE_SIZE + 1)\n+\n+#define MPNIC_MAX_JUMBO_FRAME_SIZE\t9742\n+\n+/* The page a buffer descriptor queue is currently handing out. Records\n+ * how many of the references taken on it are still unused.\n+ */\n+struct mpnic_pg_ctxt {\n+\tstruct page\t*page;\n+\tlong\t\tpagecnt_bias;\n+\tu32\t\tidx;\n+};\n+\n+struct mpnic_rcq_state {\n+\tstruct xdp_buff pkt;\n+\tstruct mpnic_pg_ctxt hdr;\n+\tstruct mpnic_pg_ctxt payld;\n+\tbool add_frag_failed;\n+};\n+\n+struct mpnic_ring {\n+\tunion {\n+\t\tstruct mpnic_rcq_state *state;\t/* RCQ */\n+\t\tstruct page **rx_buf;\t\t/* BDQ */\n+\t\tvoid **tx_buf;\t\t\t/* TWQ */\n+\t\tvoid *buffer;\t\t\t/* Generic pointer */\n+\t};\n+\n+\tu32 __iomem *doorbell;\t\t/* Pointer to CSR space for ring */\n+\t__le64 *desc;\t\t\t/* Descriptor ring memory */\n+\tu16 size_mask;\t\t\t/* Size of ring in descriptors - 1 */\n+\tu16 q_idx;\t\t\t/* Hardware queue index */\n+\n+\tu32 head, tail;\t\t\t/* Head/Tail of ring */\n+\n+\tunion {\n+\t\t/* BDQ only */\n+\t\tstruct page_pool *page_pool;\n+\n+\t\t/* TWQ only, index of the metadata descriptor of the last\n+\t\t * packet placed in the ring without ringing the doorbell,\n+\t\t * -1 if the doorbell is in sync with the tail.\n+\t\t */\n+\t\ts32 deferred_meta;\n+\t};\n+\n+\t/* Slow path fields follow */\n+\tdma_addr_t dma;\t\t\t/* Phys addr of descriptor memory */\n+\tsize_t size;\t\t\t/* Size of descriptor ring in memory */\n+};\n+\n+/* The device pairs two work queues with one completion queue. On the Rx\n+ * side they are the header and the payload buffer descriptor queues; on\n+ * the Tx side only the first one is used for now, the second one becomes\n+ * the XDP ring.\n+ */\n+struct mpnic_q_triad {\n+\tstruct xdp_rxq_info xdp_rxq;\n+\tstruct mpnic_ring sub0, sub1, cmpl;\n+};\n+\n+struct mpnic_napi_vector {\n+\tstruct napi_struct napi;\n+\tstruct device *dev;\t\t/* Device for DMA unmapping */\n+\tstruct mpnic_dev *mpd;\n+\n+\tu16 v_idx;\n+\tu16 txt_count;\n+\tu16 rxt_count;\n+\n+\tchar name[IFNAMSIZ + 11];\n+\n+\tstruct mpnic_q_triad qt[];\n+};\n+\n+netdev_tx_t mpnic_xmit_frame(struct sk_buff *skb, struct net_device *dev);\n+int mpnic_alloc_napi_vectors(struct mpnic_net *mpn);\n+void mpnic_free_napi_vectors(struct mpnic_net *mpn);\n+int mpnic_alloc_resources(struct mpnic_net *mpn);\n+void mpnic_free_resources(struct mpnic_net *mpn);\n+int mpnic_set_netif_queues(struct mpnic_net *mpn);\n+void mpnic_reset_netif_queues(struct mpnic_net *mpn);\n+void mpnic_napi_enable(struct mpnic_net *mpn);\n+void mpnic_napi_disable(struct mpnic_net *mpn);\n+void mpnic_enable(struct mpnic_net *mpn);\n+void mpnic_disable(struct mpnic_net *mpn);\n+void mpnic_wait_all_queues_idle(struct mpnic_dev *mpd);\n+void mpnic_flush(struct mpnic_net *mpn);\n+void mpnic_fill(struct mpnic_net *mpn);\n+\n+#endif /* _MPNIC_TXRX_H_ */\n"
} |
| 2/1 | 2026/09/25 06:22 | agent | patch-evaluator |
0mModel:core Args: null Results: {
"EnableConfigs": null,
"FocusSymbols": [],
"Reasoning": "The patch introduces a new network driver (MPNIC) for Meta Platforms Network Interface Controllers. This is a vendor-specific PCIe NIC requiring dedicated physical ASIC hardware (PCI Vendor ID 0x1d9b, Device ID 0x0014) not emulated in standard virtualized environments (QEMU/GCE). As such, the code is structurally unreachable during fuzzing.",
"WorthFuzzing": false
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 3b134b7aeaaa9d7b2da52bb108b65ee735fd8e6c
Author: syz-cluster <triage@syzkaller.com>
Date: Fri Sep 25 06:22:52 2026 +0000
syz-cluster: applied patch under review
diff --git a/drivers/net/ethernet/meta/Kconfig b/drivers/net/ethernet/meta/Kconfig
index ca5c7ac2a5bc2..f8b5fe7e31b73 100644
--- a/drivers/net/ethernet/meta/Kconfig
+++ b/drivers/net/ethernet/meta/Kconfig
@@ -35,4 +35,17 @@ config FBNIC
To compile this driver as a module, choose M here. The module
will be called fbnic. MSI-X interrupt support is required.
+config MPNIC
+ tristate "Meta Platforms Network Interface Controller"
+ depends on 64BIT || COMPILE_TEST
+ depends on !S390
+ depends on PCI_MSI
+ select PAGE_POOL
+ help
+ This driver supports the Meta Platforms Network Interface
+ Controller.
+
+ To compile this driver as a module, choose M here. The module
+ will be called mpnic.
+
endif # NET_VENDOR_META
diff --git a/drivers/net/ethernet/meta/Makefile b/drivers/net/ethernet/meta/Makefile
index 88804f3de9633..633973419c21b 100644
--- a/drivers/net/ethernet/meta/Makefile
+++ b/drivers/net/ethernet/meta/Makefile
@@ -4,3 +4,4 @@
#
obj-$(CONFIG_FBNIC) += fbnic/
+obj-$(CONFIG_MPNIC) += mpnic/
diff --git a/drivers/net/ethernet/meta/mpnic/Makefile b/drivers/net/ethernet/meta/mpnic/Makefile
new file mode 100644
index 0000000000000..d5bdbd5bd1c7b
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/Makefile
@@ -0,0 +1,16 @@
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) Meta Platforms, Inc. and affiliates.
+
+#
+# Makefile for the Meta(R) Platforms Network Interface Controller
+#
+
+obj-$(CONFIG_MPNIC) += mpnic.o
+
+mpnic-y := \
+ mpnic_init.o \
+ mpnic_irq.o \
+ mpnic_netdev.o \
+ mpnic_pci.o \
+ mpnic_txrx.o \
+# End of mpnic-y
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic.h b/drivers/net/ethernet/meta/mpnic/mpnic.h
new file mode 100644
index 0000000000000..78359ab6abb12
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic.h
@@ -0,0 +1,66 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#ifndef _MPNIC_H_
+#define _MPNIC_H_
+
+#include <linux/interrupt.h>
+#include <linux/io-64-nonatomic-lo-hi.h>
+#include <linux/types.h>
+
+#include "mpnic_csr.h"
+
+#define MPNIC_DRV_NAME "mpnic"
+
+#define MPNIC_MAX_TXQS 1024u
+#define MPNIC_MAX_RXQS 1024u
+
+/* misc IRQ entries are allocated before the completion queue IRQs */
+enum {
+ MPNIC_FW_MSIX_ENTRY,
+ MPNIC_NON_NAPI_VECTORS
+};
+
+struct mpnic_dev {
+ struct device *dev;
+ struct net_device *netdev;
+
+ u32 __iomem *uc_addr0;
+
+ u16 num_irqs;
+
+ u64 dsn;
+ u32 mps;
+ u32 readrq;
+ u8 relaxed_ord;
+};
+
+u64 mpnic_rd64(struct mpnic_dev *mpd, u32 reg);
+
+int mpnic_dev_init(struct mpnic_dev *mpd);
+
+int mpnic_request_irq(struct mpnic_dev *mpd, int nr, irq_handler_t handler,
+ unsigned long flags, const char *name, void *data);
+void mpnic_free_irq(struct mpnic_dev *mpd, int nr, void *data);
+void mpnic_free_irqs(struct mpnic_dev *mpd);
+int mpnic_alloc_irqs(struct mpnic_dev *mpd);
+
+static inline void mpnic_wr64(struct mpnic_dev *mpd, u32 reg, u64 val)
+{
+ u32 __iomem *csr = READ_ONCE(mpd->uc_addr0);
+
+ if (csr)
+ writeq(val, csr + reg);
+}
+
+static inline void mpnic_wrfl(struct mpnic_dev *mpd)
+{
+ mpnic_rd64(mpd, MPNIC_BDQ_SPARE);
+}
+
+static inline bool mpnic_present(struct mpnic_dev *mpd)
+{
+ return !!READ_ONCE(mpd->uc_addr0);
+}
+
+#endif /* _MPNIC_H_ */
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_csr.h b/drivers/net/ethernet/meta/mpnic/mpnic_csr.h
new file mode 100644
index 0000000000000..423378ca802c5
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_csr.h
@@ -0,0 +1,321 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#ifndef _MPNIC_CSR_H_
+#define _MPNIC_CSR_H_
+
+#include <linux/bits.h>
+
+#define CSR_BIT(nr) BIT_ULL(nr)
+#define CSR_GENMASK(h, l) GENMASK_ULL(h, l)
+
+#define DESC_BIT(nr) BIT_ULL(nr)
+#define DESC_GENMASK(h, l) GENMASK_ULL(h, l)
+
+/* Transmit Work Descriptor Format */
+#define MPNIC_TWD_L2_HLEN DESC_GENMASK(5, 0)
+#define MPNIC_TWD_FLAG_REQ_COMPLETION DESC_BIT(37)
+#define MPNIC_TWD_FLAG_DEST_MAC DESC_BIT(43)
+#define MPNIC_TWD_TYPE DESC_GENMASK(47, 46)
+enum {
+ MPNIC_TWD_TYPE_META = 0,
+ MPNIC_TWD_TYPE_AL = 2,
+ MPNIC_TWD_TYPE_LAST_AL = 3,
+};
+
+#define MPNIC_TWD_ADDR DESC_GENMASK(45, 0)
+#define MPNIC_TWD_LEN DESC_GENMASK(63, 48)
+
+/* Tx Completion Descriptor Format */
+#define MPNIC_TCD_TYPE0_HEAD0 DESC_GENMASK(15, 0)
+#define MPNIC_TCD_DONE DESC_BIT(63)
+
+/* Rx Buffer Descriptor Format */
+#define MPNIC_BD_DESC_ADDR DESC_GENMASK(39, 2)
+#define MPNIC_BD_DESC_ID DESC_GENMASK(57, 40)
+#define MPNIC_BD_DESC_BUF_SZ_LOG2 DESC_GENMASK(62, 58)
+
+/* Rx Completion Queue Descriptors */
+#define MPNIC_RCD_TYPE DESC_GENMASK(62, 61)
+enum {
+ MPNIC_RCD_TYPE_HDR_AL = 0,
+ MPNIC_RCD_TYPE_PAY_AL = 1,
+ MPNIC_RCD_TYPE_META = 3,
+};
+
+#define MPNIC_RCD_DONE DESC_BIT(63)
+
+#define MPNIC_RCD_HDR_SUBTYPE DESC_GENMASK(60, 59)
+enum {
+ MPNIC_RCD_HDR_SUBTYPE_HDR = 2,
+};
+
+/* Address/Length Completion Descriptors */
+#define MPNIC_RCD_AL_BUFF_OFF DESC_GENMASK(15, 0)
+#define MPNIC_RCD_AL_BUFF_ID DESC_GENMASK(33, 16)
+#define MPNIC_RCD_AL_BUFF_LEN DESC_GENMASK(47, 34)
+#define MPNIC_RCD_AL_PAGE_FIN DESC_BIT(53)
+
+/* Metadata Completion Descriptors */
+#define MPNIC_RCD_META_ERR_MAC_EOP DESC_BIT(53)
+#define MPNIC_RCD_META_ERR_TRUNCATED_FRAME DESC_BIT(54)
+#define MPNIC_RCD_META_UNCORRECTABLE_ERR_MASK \
+ (MPNIC_RCD_META_ERR_MAC_EOP | MPNIC_RCD_META_ERR_TRUNCATED_FRAME)
+
+/* Common fields for all DESC_CFG CSRs */
+#define MPNIC_DESC_CFG_NUM_DESCS CSR_GENMASK(2, 0)
+#define MPNIC_DESC_CFG_START_ADDR CSR_GENMASK(19, 8)
+
+/* Register Definitions
+ *
+ * The register file is addressed as an array of le32, so the byte address of
+ * a register is 4 times the index below. Each register is listed with its
+ * name, index and byte address.
+ *
+ * Name Index Address
+ *****************************************************************************/
+
+/* NIC_CORE_TDF */
+#define MPNIC_TWQ_CTL(i, j) (0x0 + 1024 * (i) + 2 * (j))
+ /* 0x0 */
+#define MPNIC_TWQ_CTL_RESET CSR_BIT(0)
+#define MPNIC_TWQ_CTL_ENABLE CSR_BIT(1)
+#define MPNIC_TWQ_TAIL(i, j) (0x4 + 1024 * (i) + 2 * (j))
+ /* 0x10 */
+#define MPNIC_TWQ_SIZE(i, j) (0x10 + 1024 * (i) + 2 * (j))
+ /* 0x40 */
+#define MPNIC_TWQ_SIZE_SIZE CSR_GENMASK(3, 0)
+#define MPNIC_TWQ_BASE_ADDR(i, j) (0x1c + 1024 * (i) + 2 * (j))
+ /* 0x70 */
+
+/* NIC_CORE_TCM */
+#define MPNIC_TCQ_CTL(i) (0x80 + 1024 * (i)) /* 0x200 */
+#define MPNIC_TCQ_CTL_RESET CSR_BIT(0)
+#define MPNIC_TCQ_CTL_ENABLE CSR_BIT(1)
+#define MPNIC_TCQ_BASE_ADDR(i) (0x86 + 1024 * (i)) /* 0x218 */
+#define MPNIC_TCQ_HEAD(i) (0x8e + 1024 * (i)) /* 0x238 */
+#define MPNIC_TCQ_SIZE(i) (0x94 + 1024 * (i)) /* 0x250 */
+#define MPNIC_TCQ_SIZE_SIZE CSR_GENMASK(4, 0)
+
+/* NIC_CORE_TIM */
+#define MPNIC_TIM_CTL1(i) (0xc0 + 1024 * (i)) /* 0x300 */
+#define MPNIC_TIM_CTL1_UPD_IGN_LONG_EVENT_CNT CSR_BIT(48)
+#define MPNIC_TIM_CTL1_UPD_IGN_LONG_TIME_CNT CSR_BIT(49)
+#define MPNIC_TIM_CTL1_UPD_IGN_SHORT_TIME_CNT CSR_BIT(50)
+#define MPNIC_TIM_CTL1_MASK CSR_BIT(51)
+#define MPNIC_TIM_CTL1_MASK_EN CSR_BIT(52)
+#define MPNIC_TIM_CTL1_TRIGGER CSR_BIT(53)
+#define MPNIC_TIM_INTR_MASK(i) (0xc8 + 1024 * (i)) /* 0x320 */
+#define MPNIC_TIM_INTR_MASK_MASK CSR_BIT(0)
+
+/* NIC_CORE_RBP */
+#define MPNIC_BDQ_CTL(i) (0x200 + 1024 * (i)) /* 0x800 */
+#define MPNIC_BDQ_CTL_RESET CSR_BIT(0)
+#define MPNIC_BDQ_CTL_ENABLE CSR_BIT(1)
+#define MPNIC_BDQ_CTL_ENABLE_PPQ CSR_BIT(3)
+#define MPNIC_HPQ_TAIL(i) (0x202 + 1024 * (i)) /* 0x808 */
+#define MPNIC_PPQ_TAIL(i) (0x204 + 1024 * (i)) /* 0x810 */
+#define MPNIC_HPQ_SIZE(i) (0x20a + 1024 * (i)) /* 0x828 */
+#define MPNIC_HPQ_SIZE_SIZE CSR_GENMASK(4, 0)
+#define MPNIC_PPQ_SIZE(i) (0x20c + 1024 * (i)) /* 0x830 */
+#define MPNIC_PPQ_SIZE_SIZE CSR_GENMASK(4, 0)
+#define MPNIC_HPQ_BASE_ADDR(i) (0x216 + 1024 * (i)) /* 0x858 */
+#define MPNIC_PPQ_BASE_ADDR(i) (0x218 + 1024 * (i)) /* 0x860 */
+
+/* NIC_CORE_RCM */
+#define MPNIC_RCQ_CTL(i) (0x280 + 1024 * (i)) /* 0xa00 */
+#define MPNIC_RCQ_CTL_RESET CSR_BIT(0)
+#define MPNIC_RCQ_CTL_ENABLE CSR_BIT(1)
+#define MPNIC_RCQ_BASE_ADDR(i) (0x286 + 1024 * (i)) /* 0xa18 */
+#define MPNIC_RCQ_HEAD(i) (0x28e + 1024 * (i)) /* 0xa38 */
+#define MPNIC_RCQ_SIZE(i) (0x294 + 1024 * (i)) /* 0xa50 */
+#define MPNIC_RCQ_SIZE_SIZE CSR_GENMASK(4, 0)
+
+/* NIC_CORE_RIM */
+#define MPNIC_RIM_INTR_MASK(i) (0x2c8 + 1024 * (i)) /* 0xb20 */
+#define MPNIC_RIM_INTR_MASK_MASK CSR_BIT(0)
+
+/* NIC_CORE_TIM_PRV */
+#define MPNIC_TIM_CTL(i) (0x100100 + 1024 * (i)) /* 0x400400 */
+
+/* NIC_CORE_RDE */
+#define MPNIC_RDE_CFG(i) (0x10021c + 1024 * (i)) /* 0x400870 */
+#define MPNIC_RDE_CFG_MIN_TAIL_ROOM CSR_GENMASK(9, 0)
+#define MPNIC_RDE_CFG_MIN_HEAD_ROOM CSR_GENMASK(18, 10)
+#define MPNIC_RDE_CFG_MAX_HEADER_BYTES CSR_GENMASK(45, 32)
+
+/* NIC_CORE_RIM_PRV */
+#define MPNIC_RIM_CTL(i) (0x100280 + 1024 * (i)) /* 0x400a00 */
+
+/* NIC_CORE_RBP_HP_GLBL */
+#define MPNIC_HPQ_IDLE(i) (0x420000 + 2 * (i)) /* 0x1080000 */
+#define MPNIC_HPQ_IDLE_CNT 16
+#define MPNIC_PPQ_IDLE(i) (0x420060 + 2 * (i)) /* 0x1080180 */
+#define MPNIC_PPQ_IDLE_CNT 16
+#define MPNIC_BDQ_GLBL_CTL0 0x420080 /* 0x1080200 */
+#define MPNIC_BDQ_GLBL_CTL0_MAX_REQ_SIZE CSR_GENMASK(26, 18)
+#define MPNIC_BDQ_GLBL_CTL0_PREFETCH_SPACE_THRESH \
+ CSR_GENMASK(42, 32)
+#define MPNIC_RDE_CTL 0x420082 /* 0x1080208 */
+#define MPNIC_RDE_CTL_HPQ_DROP_THRESHOLD CSR_GENMASK(10, 0)
+#define MPNIC_RDE_CTL_PPQ_DROP_THRESHOLD CSR_GENMASK(21, 11)
+#define MPNIC_RDE_CTL_HPQ_LOCAL_DROP_THRESHOLD CSR_GENMASK(42, 32)
+#define MPNIC_RDE_CTL_PPQ_LOCAL_DROP_THRESHOLD CSR_GENMASK(53, 43)
+#define MPNIC_BDQ_MEM_INIT_REQ 0x42013a /* 0x10804e8 */
+#define MPNIC_BDQ_MEM_INIT_DONE 0x42013c /* 0x10804f0 */
+#define MPNIC_BDQ_SPARE 0x42013e /* 0x10804f8 */
+#define MPNIC_HPQ_DESC_CFG(i) (0x420140 + 2 * (i)) /* 0x1080500 */
+#define MPNIC_PPQ_DESC_CFG(i) (0x420940 + 2 * (i)) /* 0x1082500 */
+
+/* NIC_CORE_RDE_GLBL */
+#define MPNIC_RDE_MEM_INIT_REQ 0x4240e6 /* 0x1090398 */
+#define MPNIC_RDE_MEM_INIT_DONE 0x4240e8 /* 0x10903a0 */
+
+/* NIC_CORE_RCM_GLBL */
+#define MPNIC_RCQ_IDLE(i) (0x42505e + 2 * (i)) /* 0x1094178 */
+#define MPNIC_RCQ_IDLE_CNT 16
+#define MPNIC_RCM_MEM_INIT_REQ 0x42507e /* 0x10941f8 */
+#define MPNIC_RCM_MEM_INIT_DONE 0x425080 /* 0x1094200 */
+
+/* NIC_CORE_RNI_GLBL */
+#define MPNIC_RNI_RBP_CTL 0x427000 /* 0x109c000 */
+#define MPNIC_RNI_RDE_CTL 0x427002 /* 0x109c008 */
+#define MPNIC_RNI_RDE_CTL_MPS CSR_GENMASK(1, 0)
+#define MPNIC_RNI_RDE_CTL_CLS CSR_GENMASK(3, 2)
+#define MPNIC_RNI_RCM_CTL 0x427004 /* 0x109c010 */
+
+/* NIC_CORE_TDF_GLBL */
+#define MPNIC_TWQ_IDLE(i) (0x428042 + 2 * (i)) /* 0x10a0108 */
+#define MPNIC_TWQ_IDLE_CNT 32
+#define MPNIC_TWQ_DEF_PRI_TWD 0x428082 /* 0x10a0208 */
+#define MPNIC_TDF_MEM_INIT_REQ 0x42813a /* 0x10a04e8 */
+#define MPNIC_TDF_MEM_INIT_DONE 0x42813c /* 0x10a04f0 */
+#define MPNIC_TDF_DESC_CFG(i) (0x428140 + 2 * (i)) /* 0x10a0500 */
+
+/* NIC_CORE_TQS_GLBL */
+#define MPNIC_TQS_GLBL_CTL0 0x42a000 /* 0x10a8000 */
+#define MPNIC_TQS_GLBL_CTL0_TWD_ERROR_CHECK_EN CSR_BIT(2)
+#define MPNIC_TQS_GLBL_P0_0 0x42a002 /* 0x10a8008 */
+#define MPNIC_TQS_GLBL_P0_0_TXB_MAX_CRDTS_0 CSR_GENMASK(63, 48)
+#define MPNIC_TQS_GLBL_P0_1 0x42a004 /* 0x10a8010 */
+#define MPNIC_TQS_GLBL_BMC 0x42a012 /* 0x10a8048 */
+#define MPNIC_TQS_GLBL_BMC_TXB_MAX_CRDTS CSR_GENMASK(15, 0)
+#define MPNIC_TQS_SLOWDOWN_CTL 0x42a026 /* 0x10a8098 */
+#define MPNIC_TQS_SLOWDOWN_CTL_ENABLE CSR_BIT(6)
+#define MPNIC_TQS_MTU_CTL0 0x42a030 /* 0x10a80c0 */
+#define MPNIC_TQS_MTU_CTL1 0x42a032 /* 0x10a80c8 */
+#define MPNIC_TQS_IDLE(i) (0x42a040 + 2 * (i)) /* 0x10a8100 */
+#define MPNIC_TQS_IDLE_CNT 32
+#define MPNIC_TQS_SET_P0_MAP0(i) (0x42a082 + 2 * (i)) /* 0x10a8208 */
+#define MPNIC_TQS_SET_P0_MAP1(i) (0x42a092 + 2 * (i)) /* 0x10a8248 */
+#define MPNIC_TQS_GLBL_SHAPING 0x42a108 /* 0x10a8420 */
+#define MPNIC_TQS_GLBL_SHAPING_DISABLE CSR_BIT(0)
+#define MPNIC_TQS_ARB_CTL 0x42a122 /* 0x10a8488 */
+#define MPNIC_TQS_ARB_CTL_SET_CRDT_BUCKET_EN CSR_BIT(9)
+#define MPNIC_TQS_ARB_CTL_SET_IMM_DECR_EN CSR_BIT(8)
+#define MPNIC_TQS_ARB_CTL_GROUP_CRDT_BUCKET_EN CSR_BIT(5)
+#define MPNIC_TQS_ARB_CTL_GROUP_IMM_DECR_EN CSR_BIT(4)
+#define MPNIC_TQS_ARB_CTL_QUEUE_CRDT_BUCKET_EN CSR_BIT(1)
+#define MPNIC_TQS_ARB_CTL_QUEUE_IMM_DECR_EN CSR_BIT(0)
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0 \
+ 0x42a124 /* 0x10a8490 */
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_QUEUE CSR_GENMASK(19, 0)
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_GROUP CSR_GENMASK(59, 32)
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1 \
+ 0x42a126 /* 0x10a8498 */
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_SET CSR_GENMASK(27, 0)
+#define MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_PORT CSR_GENMASK(63, 32)
+#define MPNIC_TQS_SRAM_INIT_CTL 0x42a128 /* 0x10a84a0 */
+#define MPNIC_TQS_SRAM_INIT_CTL_QUANTUM CSR_GENMASK(31, 20)
+#define MPNIC_TQS_SRAM_INIT_CTL_INIT CSR_BIT(32)
+#define MPNIC_TQS_GROUP_INIT_CTL 0x42a12a /* 0x10a84a8 */
+#define MPNIC_TQS_GROUP_INIT_CTL_QUANTUM CSR_GENMASK(51, 32)
+#define MPNIC_TQS_GROUP_INIT_CTL_INIT CSR_BIT(52)
+#define MPNIC_TQS_SET_INIT_CTL 0x42a12c /* 0x10a84b0 */
+#define MPNIC_TQS_SET_INIT_CTL_QUANTUM CSR_GENMASK(51, 32)
+#define MPNIC_TQS_SET_INIT_CTL_INIT CSR_BIT(52)
+#define MPNIC_TQS_PORT_INIT_CTL 0x42a12e /* 0x10a84b8 */
+#define MPNIC_TQS_PORT_INIT_CTL_QUANTUM CSR_GENMASK(55, 32)
+#define MPNIC_TQS_PORT_INIT_CTL_INIT CSR_BIT(56)
+#define MPNIC_TQS_SRAM_STS 0x42a130 /* 0x10a84c0 */
+#define MPNIC_TQS_PORT_CTL(i) (0x42a1e4 + 2 * (i)) /* 0x10a8790 */
+
+/* NIC_CORE_TDE_GLBL */
+#define MPNIC_TDE_IDLE(i) (0x42b000 + 2 * (i)) /* 0x10ac000 */
+#define MPNIC_TDE_IDLE_CNT 32
+#define MPNIC_TDE_MEM_INIT_REQ 0x42b1ee /* 0x10ac7b8 */
+#define MPNIC_TDE_MEM_INIT_DONE 0x42b1f0 /* 0x10ac7c0 */
+
+/* NIC_CORE_TCM_GLBL */
+#define MPNIC_TCQ_IDLE(i) (0x42c09e + 2 * (i)) /* 0x10b0278 */
+#define MPNIC_TCQ_IDLE_CNT 16
+#define MPNIC_TCM_MEM_INIT_REQ 0x42c0be /* 0x10b02f8 */
+#define MPNIC_TCM_MEM_INIT_DONE 0x42c0c0 /* 0x10b0300 */
+
+/* NIC_CORE_TNI_GLBL */
+#define MPNIC_TNI_GLBL_TDF_CTL 0x42e000 /* 0x10b8000 */
+#define MPNIC_TNI_GLBL_TDF_CTL_MRRS CSR_GENMASK(2, 0)
+#define MPNIC_TNI_GLBL_TDF_CTL_CLS CSR_GENMASK(5, 3)
+#define MPNIC_TNI_GLBL_TDE_CTL 0x42e002 /* 0x10b8008 */
+#define MPNIC_TNI_GLBL_TCM_CTL 0x42e004 /* 0x10b8010 */
+
+/* NIC_CORE_TXB */
+#define MPNIC_TXB_PORT_CONFIG 0x600000 /* 0x1800000 */
+#define MPNIC_TXB_PORT_CONFIG_PORT_MODE CSR_GENMASK(15, 13)
+#define MPNIC_TXB_BMC 0x600122 /* 0x1800488 */
+#define MPNIC_TXB_P0(i) (0x600124 + 2 * (i)) /* 0x1800490 */
+#define MPNIC_TXB_P0_CNT 17
+#define MPNIC_TXB_BMC_THRESH 0x60025c /* 0x1800970 */
+#define MPNIC_TXB_P0_THRESH(i) (0x60025e + 2 * (i)) /* 0x1800978 */
+#define MPNIC_TXB_P0_ARB_WEIGHTS(i) (0x600378 + 2 * (i)) /* 0x1800de0 */
+
+/* NIC_CORE_RXB */
+#define MPNIC_RXB_MEM_INIT_REQ 0x620002 /* 0x1880008 */
+#define MPNIC_RXB_MEM_INIT_DONE 0x620004 /* 0x1880010 */
+#define MPNIC_RXB_PORT_CFG(i) (0x620006 + 2 * (i)) /* 0x1880018 */
+#define MPNIC_RXB_PORT_CFG_FCS_STRIP_MODE CSR_GENMASK(22, 22)
+enum {
+ MPNIC_FCS_MODE_KEEP = 0,
+ MPNIC_FCS_MODE_STRIP = 1,
+};
+
+#define MPNIC_RXB_PORT_CLASS_CFG(i) (0x620020 + 2 * (i)) /* 0x1880080 */
+#define MPNIC_RXB_PORT_CLASS_CFG_DEFAULT_L2_ACTION \
+ CSR_GENMASK(0, 0)
+enum {
+ MPNIC_L2_ACTION_DROP = 0,
+ MPNIC_L2_ACTION_PASS = 1,
+};
+
+#define MPNIC_RXB_TC_CRDTS(i) (0x620418 + 2 * (i)) /* 0x1881060 */
+#define MPNIC_RXB_POOL_COMMON_CRDTS(i) (0x620458 + 2 * (i)) /* 0x1881160 */
+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC(i) \
+ (0x620472 + 2 * (i)) /* 0x18811c8 */
+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_THRESH CSR_GENMASK(15, 0)
+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_TC_EN CSR_BIT(29)
+#define MPNIC_RXB_COMMON_CRDT_CTRL_TC_MAX_CRDTS CSR_GENMASK(47, 32)
+#define MPNIC_RXB_HOST_DROP_THRESH(i) (0x6205b0 + 2 * (i)) /* 0x18816c0 */
+#define MPNIC_RXB_BMC_CRDTS(i) (0x6205d0 + 2 * (i)) /* 0x1881740 */
+
+/* NIC_CORE_RPC */
+#define MPNIC_RPC_MEM_INIT_REQ 0x780440 /* 0x1e01100 */
+#define MPNIC_RPC_MEM_INIT_DONE 0x780442 /* 0x1e01108 */
+
+/* NIC_CORE_ROF */
+#define MPNIC_RSC_GLOBAL_CONF 0x7e2002 /* 0x1f88008 */
+#define MPNIC_RSC_GLOBAL_CONF_RSC_DISABLE CSR_BIT(0)
+
+/* NIC_CORE_TOF */
+#define MPNIC_TOF_TCAM_DEST_REMAP 0x7e3022 /* 0x1f8c088 */
+
+/* PEMO_WRAPPER */
+#define MPNIC_OB_ATTR_RO CSR_BIT(1)
+#define MPNIC_OB_ATTR_TDE_H 0x9a000e /* 0x2680038 */
+#define MPNIC_OB_ATTR_TDE_P 0x9a0010 /* 0x2680040 */
+#define MPNIC_OB_ATTR_TDF 0x9a0012 /* 0x2680048 */
+#define MPNIC_OB_ATTR_RBP_HPQ 0x9a0014 /* 0x2680050 */
+#define MPNIC_OB_ATTR_RBP_PPQ 0x9a0016 /* 0x2680058 */
+#define MPNIC_OB_ATTR_RDE_H 0x9a0018 /* 0x2680060 */
+#define MPNIC_OB_ATTR_RDE_P 0x9a001a /* 0x2680068 */
+
+#endif /* _MPNIC_CSR_H_ */
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_init.c b/drivers/net/ethernet/meta/mpnic/mpnic_init.c
new file mode 100644
index 0000000000000..f8ebb19766731
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_init.c
@@ -0,0 +1,553 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bitfield.h>
+#include <linux/bits.h>
+#include <linux/cache.h>
+#include <linux/if_ether.h>
+#include <linux/iopoll.h>
+#include <linux/log2.h>
+#include <linux/sizes.h>
+
+#include "mpnic.h"
+
+#define MPNIC_MEM_INIT_POLL_US 500
+#define MPNIC_MEM_INIT_TO_US 5000
+
+/* BDQ mem init:
+ * bit 1: fifo_wrptr_mem
+ * bit 0: fifo_rdptr_mem
+ */
+#define MPNIC_MEM_INIT_BDQ_VAL 0x3
+
+/* RCM mem init:
+ * bit 3: cq_base_addr
+ * bit 2: cd_fifo_rptr_stats
+ * bit 1: cd_fifo_wptr_stats
+ * bit 0: cq_head_ptr_stats
+ */
+#define MPNIC_MEM_INIT_RCM_VAL 0xf
+
+/* RDE mem init, bits 0-18. Bits 0-12 are the per-queue packet, error and
+ * drop counters, bits 13-18 the two context memories of each of the HPQ,
+ * PPQ and SPQ descriptor prefetchers.
+ */
+#define MPNIC_MEM_INIT_RDE_VAL 0x7ffff
+
+/* RPC mem init, bit 0 covers the whole classifier. */
+#define MPNIC_MEM_INIT_RPC_VAL 0x1
+
+/* TCM mem init:
+ * bit 3: cq_head_ptr_stats
+ * bit 2: cd_fifo_wptr_stats
+ * bit 1: cd_fifo_rptr_stats
+ * bit 0: cq_base_addr
+ */
+#define MPNIC_MEM_INIT_TCM_VAL 0xf
+
+/* TDE mem init:
+ * bit 1: stats mem
+ * bit 0: dma_head_ptr SRAM
+ */
+#define MPNIC_MEM_INIT_TDE_VAL 0x3
+
+/* TDF mem init, bit 0 covers the descriptor fetch SRAM. */
+#define MPNIC_MEM_INIT_TDF_VAL 0x1
+
+/* RXB mem init, bit 0 covers the DMAC TCAM statistics RAM. */
+#define MPNIC_MEM_INIT_RXB_VAL 0x1
+
+/* TQS arbiter SRAM init done, one bit per DWRR level. */
+#define MPNIC_TQS_ARB_INIT_VAL 0xf
+
+/* On-chip SRAM allocated to each queue for descriptor fetch, in units of
+ * descriptors. Valid values are 64, 128, 256, 512 and 1024.
+ *
+ * The partition sizes are chosen for the maximum number of queues the
+ * device supports, so they do not have to be adjusted when the active
+ * queue count changes:
+ *
+ * BDQ: 1 MiB / (8 B/DESC) / (1024 HPQ + 1024 PPQ) = 64 DESC/QUEUE
+ * TWQ: 2 MiB / (8 B/DESC) / (1024 TXQ * 2 TWQ) = 128 DESC/QUEUE
+ */
+#define MPNIC_BDQ_SRAM_DESCS 64u
+#define MPNIC_TDF_SRAM_DESCS 128u
+
+/* A total of 1 MiB worth of Tx credits is available, in units of 128 B.
+ * The BMC gets a guaranteed share of them whether or not the host is
+ * routing anything its way, everything else goes to MAC TC0.
+ */
+#define MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL 800
+#define MPNIC_TXB_P0_MAC_PVT_CRDT_INIT_VAL \
+ (SZ_1M / 128 - 2 * MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL)
+
+/* The recommended lower bound for the TXB threshold is 80, based on a 10K
+ * MTU. Round up by 20% to stay on the defensive side. The same reasoning
+ * applies to the arbitration weight, which has to exceed the full packet
+ * size.
+ */
+#define MPNIC_TXB_INIT_BMC_THRESH 100
+#define MPNIC_TXB_INIT_P0_THRESH 100
+#define MPNIC_TXB_INIT_P0_ARB_WEIGHTS 0x64
+
+/* RXB host drop threshold in units of 128 B beats. Packets targeting a
+ * queue are dropped when the available credits fall below it. 80 beats is
+ * 10 KB, which is also the largest frame the device is configured for.
+ */
+#define MPNIC_RXB_INIT_HOST_DROP_THRESH 0x50
+
+/* A total of 8 MiB of Rx buffer is available. The recommended per-TC pool
+ * for a 800G configuration is 420 KB with all 8 TCs enabled. Only one TC
+ * is in use, so give it 8 * 420 KB (in units of 128 B) and push the rest
+ * to the common pool. Start drawing from the common pool as soon as the
+ * TC0 credits fall below one max sized frame. The BMC keeps a small
+ * reserve of its own whether or not the host talks to it.
+ */
+#define MPNIC_RXB_INIT_POOL_TC_CRDTS_P0 (0xd20 * 8)
+#define MPNIC_RXB_INIT_BMC_CRDTS 0x20
+#define MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH \
+ (SZ_8M / 128 - MPNIC_RXB_INIT_POOL_TC_CRDTS_P0 - \
+ MPNIC_RXB_INIT_BMC_CRDTS)
+#define MPNIC_RXB_INIT_COMMON_CRDT_THRSH 0x50
+
+/* TXB port mode selects the number of active MAC ports for Tx buffer
+ * credit distribution and arbitration. The hardware only supports single,
+ * dual and quad port, encoded as b'001, b'010 and b'100.
+ */
+#define MPNIC_TXB_PORT_MODE_SINGLE 1
+
+/* The MAC traffic classes start at index 8 in the TXB arrays, the BMC
+ * sits above them.
+ */
+#define MPNIC_TXB_TC_IDX_MAC_0 8
+#define MPNIC_TXB_TC_IDX_BMC 16
+
+/* Largest frame the Tx queue scheduler will pass through. Anything above
+ * it gets truncated.
+ */
+#define MPNIC_TQS_MTU_CTL0_MAX 0x2800
+
+/* The unit of the DWRR quantum is 256 B. It has to be large enough for at
+ * least 11 MTUs to be transmitted in one quantum; use 15 for headroom.
+ */
+#define MPNIC_TQS_DWRR_INIT_QUANTUM (15 * MPNIC_TQS_MTU_CTL0_MAX / 256)
+
+/* Lower bound on the credit available to a queue, group, set or port
+ * before the scheduler stops issuing requests for it. 15 MTUs, matching
+ * the DWRR quantum above.
+ */
+#define MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH (15 * MPNIC_TQS_MTU_CTL0_MAX)
+
+/* The 1 MiB Tx buffer is partitioned between the MAC and the BMC in units
+ * of 1 KiB. The BMC portion is fixed at 100 KB.
+ */
+#define MPNIC_TQS_GLBL_TXB_CRDT_BMC 100
+#define MPNIC_TQS_GLBL_TXB_CRDT_MAC (SZ_1M / SZ_1K - \
+ MPNIC_TQS_GLBL_TXB_CRDT_BMC)
+
+struct mpnic_init_poll {
+ u64 exp_val;
+ u32 addr;
+};
+
+struct mpnic_poll_state {
+ int poll_idx;
+ u64 val;
+};
+
+static void mpnic_tdf_glbl_init(struct mpnic_dev *mpd)
+{
+ /* Default metadata descriptor, used for frames the driver did not
+ * prepend one to.
+ */
+ mpnic_wr64(mpd, MPNIC_TWQ_DEF_PRI_TWD,
+ FIELD_PREP(MPNIC_TWD_L2_HLEN, ETH_HLEN) |
+ MPNIC_TWD_FLAG_REQ_COMPLETION);
+}
+
+static void mpnic_txb_init(struct mpnic_dev *mpd)
+{
+ int i;
+
+ mpnic_wr64(mpd, MPNIC_TXB_BMC, MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL);
+
+ /* Zero the private credits of every traffic class, then hand the
+ * unreserved ones to MAC TC0.
+ */
+ for (i = 0; i < MPNIC_TXB_P0_CNT; i++)
+ mpnic_wr64(mpd, MPNIC_TXB_P0(i), 0);
+ mpnic_wr64(mpd, MPNIC_TXB_P0(MPNIC_TXB_TC_IDX_MAC_0),
+ MPNIC_TXB_P0_MAC_PVT_CRDT_INIT_VAL);
+ mpnic_wr64(mpd, MPNIC_TXB_P0(MPNIC_TXB_TC_IDX_BMC),
+ MPNIC_TXB_BMC_PVT_CRDT_INIT_VAL);
+
+ mpnic_wr64(mpd, MPNIC_TXB_BMC_THRESH, MPNIC_TXB_INIT_BMC_THRESH);
+
+ mpnic_wr64(mpd, MPNIC_TXB_P0_THRESH(MPNIC_TXB_TC_IDX_MAC_0),
+ MPNIC_TXB_INIT_P0_THRESH);
+ mpnic_wr64(mpd, MPNIC_TXB_P0_ARB_WEIGHTS(MPNIC_TXB_TC_IDX_MAC_0),
+ MPNIC_TXB_INIT_P0_ARB_WEIGHTS);
+ mpnic_wr64(mpd, MPNIC_TXB_PORT_CONFIG,
+ FIELD_PREP(MPNIC_TXB_PORT_CONFIG_PORT_MODE,
+ MPNIC_TXB_PORT_MODE_SINGLE));
+}
+
+static void mpnic_rxb_init(struct mpnic_dev *mpd)
+{
+ /* Accept all packets that miss dmac tcam, until l2 filtering is
+ * implemented.
+ */
+ mpnic_wr64(mpd, MPNIC_RXB_PORT_CLASS_CFG(0),
+ FIELD_PREP(MPNIC_RXB_PORT_CLASS_CFG_DEFAULT_L2_ACTION,
+ MPNIC_L2_ACTION_PASS));
+ mpnic_wr64(mpd, MPNIC_RXB_PORT_CFG(0),
+ FIELD_PREP(MPNIC_RXB_PORT_CFG_FCS_STRIP_MODE,
+ MPNIC_FCS_MODE_STRIP));
+
+ mpnic_wr64(mpd, MPNIC_RXB_HOST_DROP_THRESH(0),
+ MPNIC_RXB_INIT_HOST_DROP_THRESH);
+
+ mpnic_wr64(mpd, MPNIC_RXB_TC_CRDTS(0),
+ MPNIC_RXB_INIT_POOL_TC_CRDTS_P0);
+ mpnic_wr64(mpd, MPNIC_RXB_COMMON_CRDT_CTRL_TC(0),
+ FIELD_PREP(MPNIC_RXB_COMMON_CRDT_CTRL_TC_THRESH,
+ MPNIC_RXB_INIT_COMMON_CRDT_THRSH) |
+ FIELD_PREP(MPNIC_RXB_COMMON_CRDT_CTRL_TC_MAX_CRDTS,
+ MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH) |
+ MPNIC_RXB_COMMON_CRDT_CTRL_TC_TC_EN);
+
+ /* Only pool 0 is used, it gets all of the common credits */
+ mpnic_wr64(mpd, MPNIC_RXB_POOL_COMMON_CRDTS(0),
+ MPNIC_RXB_INIT_COMMON_CRDT_MAX_THRSH);
+
+ mpnic_wr64(mpd, MPNIC_RXB_BMC_CRDTS(0), MPNIC_RXB_INIT_BMC_CRDTS);
+
+ mpnic_wr64(mpd, MPNIC_RXB_MEM_INIT_REQ, MPNIC_MEM_INIT_RXB_VAL);
+}
+
+static u64 mpnic_desc_cfg(unsigned int sram_descs, unsigned int q_idx)
+{
+ /* The hardware encodes the partition size as 64 * 2^n descriptors,
+ * and its start address in units of 64 descriptors.
+ */
+ return FIELD_PREP(MPNIC_DESC_CFG_NUM_DESCS, __ffs(sram_descs) - 6) |
+ FIELD_PREP(MPNIC_DESC_CFG_START_ADDR,
+ q_idx * (sram_descs >> 6));
+}
+
+static void mpnic_desc_sram_init(struct mpnic_dev *mpd)
+{
+ int i;
+
+ for (i = 0; i < MPNIC_MAX_TXQS * 2; i++)
+ mpnic_wr64(mpd, MPNIC_TDF_DESC_CFG(i),
+ mpnic_desc_cfg(MPNIC_TDF_SRAM_DESCS, i));
+
+ for (i = 0; i < MPNIC_MAX_RXQS; i++) {
+ mpnic_wr64(mpd, MPNIC_HPQ_DESC_CFG(i),
+ mpnic_desc_cfg(MPNIC_BDQ_SRAM_DESCS, i));
+ mpnic_wr64(mpd, MPNIC_PPQ_DESC_CFG(i),
+ mpnic_desc_cfg(MPNIC_BDQ_SRAM_DESCS,
+ MPNIC_MAX_RXQS + i));
+ }
+}
+
+static void mpnic_rxglb_init(struct mpnic_dev *mpd)
+{
+ /* Descriptor prefetch reads are only issued once 32 descriptors
+ * worth of FIFO space is available, and no more than 64 descriptors
+ * are fetched for one queue at a time so that a single queue cannot
+ * monopolize the bus. Both have to be multiples of 16 to keep the
+ * reads 128 B aligned.
+ */
+ mpnic_wr64(mpd, MPNIC_BDQ_GLBL_CTL0,
+ FIELD_PREP(MPNIC_BDQ_GLBL_CTL0_PREFETCH_SPACE_THRESH, 32) |
+ FIELD_PREP(MPNIC_BDQ_GLBL_CTL0_MAX_REQ_SIZE, 64));
+
+ /* Minimum number of descriptors which has to be available before
+ * the descriptor engine considers a queue usable, globally and in
+ * the per-queue prefetch FIFO.
+ */
+ mpnic_wr64(mpd, MPNIC_RDE_CTL,
+ FIELD_PREP(MPNIC_RDE_CTL_HPQ_DROP_THRESHOLD, 16) |
+ FIELD_PREP(MPNIC_RDE_CTL_PPQ_DROP_THRESHOLD, 16) |
+ FIELD_PREP(MPNIC_RDE_CTL_HPQ_LOCAL_DROP_THRESHOLD, 16) |
+ FIELD_PREP(MPNIC_RDE_CTL_PPQ_LOCAL_DROP_THRESHOLD, 16));
+
+ /* Receive side coalescing is not supported yet */
+ mpnic_wr64(mpd, MPNIC_RSC_GLOBAL_CONF,
+ MPNIC_RSC_GLOBAL_CONF_RSC_DISABLE);
+
+ mpnic_wr64(mpd, MPNIC_BDQ_MEM_INIT_REQ, MPNIC_MEM_INIT_BDQ_VAL);
+ mpnic_wr64(mpd, MPNIC_RCM_MEM_INIT_REQ, MPNIC_MEM_INIT_RCM_VAL);
+ mpnic_wr64(mpd, MPNIC_RDE_MEM_INIT_REQ, MPNIC_MEM_INIT_RDE_VAL);
+ mpnic_wr64(mpd, MPNIC_RPC_MEM_INIT_REQ, MPNIC_MEM_INIT_RPC_VAL);
+}
+
+static void mpnic_txglb_init(struct mpnic_dev *mpd)
+{
+ /* Nothing is redirected to the BMC until the Tx offload TCAM gets
+ * programmed with its addresses.
+ */
+ mpnic_wr64(mpd, MPNIC_TOF_TCAM_DEST_REMAP, 0);
+
+ mpnic_wr64(mpd, MPNIC_TCM_MEM_INIT_REQ, MPNIC_MEM_INIT_TCM_VAL);
+ mpnic_wr64(mpd, MPNIC_TDE_MEM_INIT_REQ, MPNIC_MEM_INIT_TDE_VAL);
+ mpnic_wr64(mpd, MPNIC_TDF_MEM_INIT_REQ, MPNIC_MEM_INIT_TDF_VAL);
+}
+
+/* Fill the DWRR arbiter memories. Setting the INIT bit makes the hardware
+ * write the given credit and quantum into every entry at the queue, group,
+ * set and port level, so nothing has to be programmed per queue.
+ */
+static void mpnic_tqs_sram_init(struct mpnic_dev *mpd)
+{
+ mpnic_wr64(mpd, MPNIC_TQS_GROUP_INIT_CTL,
+ FIELD_PREP(MPNIC_TQS_GROUP_INIT_CTL_QUANTUM,
+ MPNIC_TQS_DWRR_INIT_QUANTUM) |
+ MPNIC_TQS_GROUP_INIT_CTL_INIT);
+ mpnic_wr64(mpd, MPNIC_TQS_SET_INIT_CTL,
+ FIELD_PREP(MPNIC_TQS_SET_INIT_CTL_QUANTUM,
+ MPNIC_TQS_DWRR_INIT_QUANTUM) |
+ MPNIC_TQS_SET_INIT_CTL_INIT);
+ mpnic_wr64(mpd, MPNIC_TQS_PORT_INIT_CTL,
+ FIELD_PREP(MPNIC_TQS_PORT_INIT_CTL_QUANTUM,
+ MPNIC_TQS_DWRR_INIT_QUANTUM) |
+ MPNIC_TQS_PORT_INIT_CTL_INIT);
+ mpnic_wr64(mpd, MPNIC_TQS_SRAM_INIT_CTL,
+ FIELD_PREP(MPNIC_TQS_SRAM_INIT_CTL_QUANTUM,
+ MPNIC_TQS_DWRR_INIT_QUANTUM) |
+ MPNIC_TQS_SRAM_INIT_CTL_INIT);
+}
+
+static void mpnic_tqs_init(struct mpnic_dev *mpd)
+{
+ u64 val;
+
+ /* Initialize to the largest frame we support, the scheduler
+ * truncates anything above it. The BMC gets the same limit.
+ */
+ mpnic_wr64(mpd, MPNIC_TQS_MTU_CTL0, MPNIC_TQS_MTU_CTL0_MAX);
+ mpnic_wr64(mpd, MPNIC_TQS_MTU_CTL1, MPNIC_TQS_MTU_CTL0_MAX);
+
+ mpnic_wr64(mpd, MPNIC_TQS_GLBL_CTL0,
+ MPNIC_TQS_GLBL_CTL0_TWD_ERROR_CHECK_EN);
+
+ mpnic_tqs_sram_init(mpd);
+
+ /* Only port 0 is used. A single traffic class is in use as well, so
+ * give all of the Tx buffer credits to TC0.
+ */
+ mpnic_wr64(mpd, MPNIC_TQS_GLBL_P0_0,
+ FIELD_PREP(MPNIC_TQS_GLBL_P0_0_TXB_MAX_CRDTS_0,
+ MPNIC_TQS_GLBL_TXB_CRDT_MAC));
+ mpnic_wr64(mpd, MPNIC_TQS_GLBL_P0_1, 0);
+ mpnic_wr64(mpd, MPNIC_TQS_GLBL_BMC,
+ FIELD_PREP(MPNIC_TQS_GLBL_BMC_TXB_MAX_CRDTS,
+ MPNIC_TQS_GLBL_TXB_CRDT_BMC));
+
+ mpnic_wr64(mpd, MPNIC_TQS_PORT_CTL(0), 0);
+
+ /* Map all sets to port 0 */
+ mpnic_wr64(mpd, MPNIC_TQS_SET_P0_MAP0(0), ~0ULL);
+ mpnic_wr64(mpd, MPNIC_TQS_SET_P0_MAP1(0), ~0ULL);
+
+ mpnic_wr64(mpd, MPNIC_TQS_CEV_MIN_SCHED_THRESH_0,
+ FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_QUEUE,
+ MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH) |
+ FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_0_GROUP,
+ MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH));
+ mpnic_wr64(mpd, MPNIC_TQS_CEV_MIN_SCHED_THRESH_1,
+ FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_SET,
+ MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH) |
+ FIELD_PREP(MPNIC_TQS_CEV_MIN_SCHED_THRESH_1_PORT,
+ MPNIC_TQS_DWRR_CEV_MIN_SCHED_THRESH));
+
+ /* The rate limiters are left uninitialized, so shaping has to stay
+ * off or nothing would ever get scheduled.
+ */
+ mpnic_wr64(mpd, MPNIC_TQS_GLBL_SHAPING, MPNIC_TQS_GLBL_SHAPING_DISABLE);
+
+ /* Enable fairness protection (phantom eligibility). Read modify
+ * write so that the reset default slowdown cycle is preserved.
+ */
+ val = mpnic_rd64(mpd, MPNIC_TQS_SLOWDOWN_CTL);
+ val |= MPNIC_TQS_SLOWDOWN_CTL_ENABLE;
+ mpnic_wr64(mpd, MPNIC_TQS_SLOWDOWN_CTL, val);
+
+ /* Use immediate credit decrement at every DWRR level so that the
+ * credit reflects a grant in the same cycle.
+ */
+ val = mpnic_rd64(mpd, MPNIC_TQS_ARB_CTL);
+ val |= MPNIC_TQS_ARB_CTL_SET_CRDT_BUCKET_EN |
+ MPNIC_TQS_ARB_CTL_SET_IMM_DECR_EN |
+ MPNIC_TQS_ARB_CTL_GROUP_CRDT_BUCKET_EN |
+ MPNIC_TQS_ARB_CTL_GROUP_IMM_DECR_EN |
+ MPNIC_TQS_ARB_CTL_QUEUE_CRDT_BUCKET_EN |
+ MPNIC_TQS_ARB_CTL_QUEUE_IMM_DECR_EN;
+ mpnic_wr64(mpd, MPNIC_TQS_ARB_CTL, val);
+}
+
+/* The MPS and CLS fields sit at the same bit positions in every block, so
+ * one set of masks covers both the RNI and the TNI registers.
+ */
+static void mpnic_mps_init(struct mpnic_dev *mpd, u32 reg, unsigned int mps,
+ unsigned int cls)
+{
+ u64 val = mpnic_rd64(mpd, reg);
+
+ val &= ~(MPNIC_RNI_RDE_CTL_MPS | MPNIC_RNI_RDE_CTL_CLS);
+ val |= FIELD_PREP(MPNIC_RNI_RDE_CTL_MPS, mps) |
+ FIELD_PREP(MPNIC_RNI_RDE_CTL_CLS, cls);
+
+ mpnic_wr64(mpd, reg, val);
+}
+
+/* Likewise for the MRRS and CLS fields, which have their own common
+ * layout.
+ */
+static void mpnic_mrrs_init(struct mpnic_dev *mpd, u32 reg, unsigned int mrrs,
+ unsigned int cls)
+{
+ u64 val = mpnic_rd64(mpd, reg);
+
+ val &= ~(MPNIC_TNI_GLBL_TDF_CTL_MRRS | MPNIC_TNI_GLBL_TDF_CTL_CLS);
+ val |= FIELD_PREP(MPNIC_TNI_GLBL_TDF_CTL_MRRS, mrrs) |
+ FIELD_PREP(MPNIC_TNI_GLBL_TDF_CTL_CLS, cls);
+
+ mpnic_wr64(mpd, reg, val);
+}
+
+/**
+ * mpnic_axi_init - Configure AXI bus parameters from host PCIe capabilities
+ * @mpd: Device to configure
+ *
+ * Programs the max read request size, max payload size and cache line size
+ * of the DMA engines. The hardware encodes all three as a power of 2 index,
+ * MRRS and MPS relative to 128 B and CLS relative to 64 B.
+ *
+ * MAX_OT and MAX_OB are left at their hardware defaults.
+ */
+static void mpnic_axi_init(struct mpnic_dev *mpd)
+{
+ int mps, cls, mrrs;
+
+ mps = clamp(ilog2(mpd->mps) - 7, 0, 3);
+ cls = clamp(ilog2(L1_CACHE_BYTES) - 6, 0, 3);
+
+ mpnic_mps_init(mpd, MPNIC_RNI_RDE_CTL, mps, cls);
+ mpnic_mps_init(mpd, MPNIC_RNI_RCM_CTL, mps, cls);
+ mpnic_mps_init(mpd, MPNIC_TNI_GLBL_TCM_CTL, mps, cls);
+
+ mrrs = clamp(ilog2(mpd->readrq) - 7, 0, 3);
+ mpnic_mrrs_init(mpd, MPNIC_RNI_RBP_CTL, mrrs, cls);
+ mpnic_mrrs_init(mpd, MPNIC_TNI_GLBL_TDF_CTL, mrrs, cls);
+
+ /* TDE supports a wider range of MRRS encodings. */
+ mrrs = clamp(ilog2(mpd->readrq) - 7, 0, 5);
+ mpnic_mrrs_init(mpd, MPNIC_TNI_GLBL_TDE_CTL, mrrs, cls);
+}
+
+/**
+ * mpnic_ro_init - Set relaxed ordering on the outbound TLP attributes
+ * @mpd: Device to configure
+ *
+ * Completions must stay ordered so that they are not observed before the
+ * payload DMA they describe has landed, so RCM and TCM are left alone.
+ */
+static void mpnic_ro_init(struct mpnic_dev *mpd)
+{
+ u64 attr = mpd->relaxed_ord ? MPNIC_OB_ATTR_RO : 0;
+
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_TDE_H, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_TDE_P, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_TDF, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_RBP_HPQ, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_RBP_PPQ, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_RDE_H, attr);
+ mpnic_wr64(mpd, MPNIC_OB_ATTR_RDE_P, attr);
+}
+
+static bool mpnic_init_status_ready(struct mpnic_dev *mpd,
+ const struct mpnic_init_poll *polls,
+ struct mpnic_poll_state *state)
+{
+ u64 val;
+ int i;
+
+ for (i = state->poll_idx; polls[i].addr; i++) {
+ val = mpnic_rd64(mpd, polls[i].addr);
+
+ if ((val & polls[i].exp_val) != polls[i].exp_val) {
+ state->poll_idx = i;
+ state->val = val;
+ return false;
+ }
+ }
+
+ return true;
+}
+
+/**
+ * mpnic_mem_init_poll - Wait for the memory initializations to complete
+ * @mpd: Device to poll
+ *
+ * The blocks initialize their memories in parallel, so walk the status
+ * registers in order and only go back to sleep on the first one which is
+ * not done yet.
+ *
+ * Return: 0 on success, -ETIMEDOUT if not everything completed in time
+ */
+static int mpnic_mem_init_poll(struct mpnic_dev *mpd)
+{
+ static const struct mpnic_init_poll polls[] = {
+ { MPNIC_TQS_ARB_INIT_VAL, MPNIC_TQS_SRAM_STS },
+ { MPNIC_MEM_INIT_BDQ_VAL, MPNIC_BDQ_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_RCM_VAL, MPNIC_RCM_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_RDE_VAL, MPNIC_RDE_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_RPC_VAL, MPNIC_RPC_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_TCM_VAL, MPNIC_TCM_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_TDE_VAL, MPNIC_TDE_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_TDF_VAL, MPNIC_TDF_MEM_INIT_DONE },
+ { MPNIC_MEM_INIT_RXB_VAL, MPNIC_RXB_MEM_INIT_DONE },
+ { 0 },
+ };
+ struct mpnic_poll_state state = {};
+ bool done;
+ int err;
+
+ err = read_poll_timeout(mpnic_init_status_ready, done, done,
+ MPNIC_MEM_INIT_POLL_US, MPNIC_MEM_INIT_TO_US,
+ false, mpd, polls, &state);
+ if (err)
+ dev_err(mpd->dev, "Poll timeout for reg 0x%x: 0x%llx\n",
+ polls[state.poll_idx].addr, state.val);
+
+ return err;
+}
+
+int mpnic_dev_init(struct mpnic_dev *mpd)
+{
+ int err;
+
+ mpnic_tdf_glbl_init(mpd);
+ mpnic_txb_init(mpd);
+ mpnic_rxb_init(mpd);
+ mpnic_desc_sram_init(mpd);
+ mpnic_axi_init(mpd);
+ mpnic_ro_init(mpd);
+ mpnic_rxglb_init(mpd);
+ mpnic_txglb_init(mpd);
+ mpnic_tqs_init(mpd);
+
+ err = mpnic_mem_init_poll(mpd);
+ if (err) {
+ dev_err(mpd->dev, "Device initialization failed: %d\n", err);
+ return err;
+ }
+
+ if (!mpnic_present(mpd))
+ return -EIO;
+
+ return 0;
+}
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_irq.c b/drivers/net/ethernet/meta/mpnic/mpnic_irq.c
new file mode 100644
index 0000000000000..bcc33655cbbea
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_irq.c
@@ -0,0 +1,64 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#include <linux/cpumask.h>
+#include <linux/interrupt.h>
+#include <linux/minmax.h>
+#include <linux/pci.h>
+
+#include "mpnic.h"
+
+int mpnic_request_irq(struct mpnic_dev *mpd, int nr, irq_handler_t handler,
+ unsigned long flags, const char *name, void *data)
+{
+ struct pci_dev *pdev = to_pci_dev(mpd->dev);
+ int irq = pci_irq_vector(pdev, nr);
+
+ if (irq < 0)
+ return irq;
+
+ return request_irq(irq, handler, flags, name, data);
+}
+
+void mpnic_free_irq(struct mpnic_dev *mpd, int nr, void *data)
+{
+ struct pci_dev *pdev = to_pci_dev(mpd->dev);
+ int irq = pci_irq_vector(pdev, nr);
+
+ if (irq < 0)
+ return;
+
+ free_irq(irq, data);
+}
+
+void mpnic_free_irqs(struct mpnic_dev *mpd)
+{
+ struct pci_dev *pdev = to_pci_dev(mpd->dev);
+
+ mpd->num_irqs = 0;
+ pci_free_irq_vectors(pdev);
+}
+
+int mpnic_alloc_irqs(struct mpnic_dev *mpd)
+{
+ unsigned int wanted_irqs = MPNIC_NON_NAPI_VECTORS;
+ struct pci_dev *pdev = to_pci_dev(mpd->dev);
+ int num_irqs;
+
+ wanted_irqs += min_t(unsigned int, num_online_cpus(), MPNIC_MAX_RXQS);
+ num_irqs = pci_alloc_irq_vectors(pdev, MPNIC_NON_NAPI_VECTORS + 1,
+ wanted_irqs, PCI_IRQ_MSIX);
+ if (num_irqs < 0) {
+ dev_err(mpd->dev, "Failed to allocate MSI-X entries: %d\n",
+ num_irqs);
+ return num_irqs;
+ }
+
+ if (num_irqs < wanted_irqs)
+ dev_warn(mpd->dev, "Allocated %d IRQs, expected %u\n",
+ num_irqs, wanted_irqs);
+
+ mpd->num_irqs = num_irqs;
+
+ return 0;
+}
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c
new file mode 100644
index 0000000000000..fd34f48a654ca
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.c
@@ -0,0 +1,181 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#include <linux/etherdevice.h>
+#include <linux/ipv6.h>
+#include <linux/netdevice.h>
+#include <linux/pci.h>
+#include <linux/types.h>
+
+#include "mpnic.h"
+#include "mpnic_netdev.h"
+#include "mpnic_txrx.h"
+
+static int mpnic_open(struct net_device *netdev)
+{
+ struct mpnic_net *mpn = netdev_priv(netdev);
+ int err;
+
+ err = mpnic_alloc_napi_vectors(mpn);
+ if (err)
+ return err;
+
+ err = mpnic_alloc_resources(mpn);
+ if (err)
+ goto err_free_napi_vectors;
+
+ err = mpnic_set_netif_queues(mpn);
+ if (err)
+ goto err_free_resources;
+
+ mpnic_enable(mpn);
+ mpnic_fill(mpn);
+ mpnic_napi_enable(mpn);
+
+ netif_tx_wake_all_queues(netdev);
+ netif_carrier_on(netdev);
+
+ return 0;
+
+err_free_resources:
+ mpnic_free_resources(mpn);
+err_free_napi_vectors:
+ mpnic_free_napi_vectors(mpn);
+ return err;
+}
+
+static int mpnic_stop(struct net_device *netdev)
+{
+ struct mpnic_net *mpn = netdev_priv(netdev);
+
+ netif_carrier_off(netdev);
+
+ mpnic_napi_disable(mpn);
+ netif_tx_disable(netdev);
+
+ mpnic_disable(mpn);
+ mpnic_wait_all_queues_idle(mpn->mpd);
+ mpnic_flush(mpn);
+
+ mpnic_reset_netif_queues(mpn);
+ mpnic_free_resources(mpn);
+ mpnic_free_napi_vectors(mpn);
+
+ return 0;
+}
+
+static const struct net_device_ops mpnic_netdev_ops = {
+ .ndo_open = mpnic_open,
+ .ndo_stop = mpnic_stop,
+ .ndo_validate_addr = eth_validate_addr,
+ .ndo_start_xmit = mpnic_xmit_frame,
+};
+
+/**
+ * mpnic_netdev_free - Free the netdev associated with mpnic
+ * @mpd: Driver specific structure to free netdev from
+ **/
+void mpnic_netdev_free(struct mpnic_dev *mpd)
+{
+ free_netdev(mpd->netdev);
+ mpd->netdev = NULL;
+}
+
+/**
+ * mpnic_netdev_alloc - Allocate a netdev and associate it with mpnic
+ * @mpd: Driver specific structure to associate the netdev with
+ *
+ * Return: NULL on failure.
+ **/
+struct net_device *mpnic_netdev_alloc(struct mpnic_dev *mpd)
+{
+ struct net_device *netdev;
+ struct mpnic_net *mpn;
+ unsigned int queues;
+
+ netdev = alloc_etherdev_mq(sizeof(*mpn), MPNIC_MAX_RXQS);
+ if (!netdev)
+ return NULL;
+
+ SET_NETDEV_DEV(netdev, mpd->dev);
+ mpd->netdev = netdev;
+
+ netdev->netdev_ops = &mpnic_netdev_ops;
+ netdev->request_ops_lock = true;
+
+ mpn = netdev_priv(netdev);
+ mpn->netdev = netdev;
+ mpn->mpd = mpd;
+
+ mpn->txq_size = MPNIC_TXQ_SIZE_DEFAULT;
+ mpn->hpq_size = MPNIC_HPQ_SIZE_DEFAULT;
+ mpn->ppq_size = MPNIC_PPQ_SIZE_DEFAULT;
+ mpn->rcq_size = MPNIC_RCQ_SIZE_DEFAULT;
+
+ queues = min(netif_get_num_default_rss_queues(),
+ mpd->num_irqs - MPNIC_NON_NAPI_VECTORS);
+ mpn->num_tx_queues = queues;
+ mpn->num_rx_queues = queues;
+ mpn->num_napi = queues;
+
+ netdev->features |= NETIF_F_SG;
+ netdev->hw_features |= netdev->features;
+ netdev->vlan_features |= netdev->features;
+
+ netdev->min_mtu = IPV6_MIN_MTU;
+ netdev->max_mtu = MPNIC_MAX_JUMBO_FRAME_SIZE - ETH_HLEN;
+
+ netif_carrier_off(netdev);
+ netif_tx_stop_all_queues(netdev);
+
+ return netdev;
+}
+
+static int mpnic_dsn_to_mac_addr(u64 dsn, char *addr)
+{
+ addr[0] = (dsn >> 56) & 0xFF;
+ addr[1] = (dsn >> 48) & 0xFF;
+ addr[2] = (dsn >> 40) & 0xFF;
+ addr[3] = (dsn >> 16) & 0xFF;
+ addr[4] = (dsn >> 8) & 0xFF;
+ addr[5] = dsn & 0xFF;
+
+ return is_valid_ether_addr(addr) ? 0 : -EINVAL;
+}
+
+/**
+ * mpnic_netdev_register - Assign the MAC address and register the netdev
+ * @netdev: Netdev to register
+ *
+ * The permanent address is derived from the PCIe device serial number, the
+ * same way the firmware and the BMC derive it. A random address would break
+ * provisioning, so refuse to spawn the interface if the serial number does
+ * not yield a valid one.
+ *
+ * Return: non-zero on failure.
+ **/
+int mpnic_netdev_register(struct net_device *netdev)
+{
+ struct mpnic_net *mpn = netdev_priv(netdev);
+ struct mpnic_dev *mpd = mpn->mpd;
+ u8 addr[ETH_ALEN];
+ int err;
+
+ err = mpnic_dsn_to_mac_addr(mpd->dsn, addr);
+ if (err) {
+ dev_err(mpd->dev, "MAC addr %pM invalid\n", addr);
+ return err;
+ }
+
+ ether_addr_copy(netdev->perm_addr, addr);
+ eth_hw_addr_set(netdev, addr);
+
+ /* Abort if MMIO has failed. This has to be the last check before
+ * registration, the register accessors can only detach the device
+ * once it has been registered.
+ */
+ if (!mpnic_present(mpd))
+ return -EIO;
+
+ return register_netdev(netdev);
+}
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h
new file mode 100644
index 0000000000000..ccb0929f9180a
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_netdev.h
@@ -0,0 +1,35 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#ifndef _MPNIC_NETDEV_H_
+#define _MPNIC_NETDEV_H_
+
+#include <linux/types.h>
+
+#include "mpnic.h"
+#include "mpnic_txrx.h"
+
+struct mpnic_net {
+ struct mpnic_ring *tx[MPNIC_MAX_TXQS];
+ struct mpnic_ring *rx[MPNIC_MAX_RXQS];
+
+ struct mpnic_napi_vector *napi[MPNIC_MAX_NAPI_VECTORS];
+
+ struct net_device *netdev;
+ struct mpnic_dev *mpd;
+
+ u32 txq_size;
+ u32 hpq_size;
+ u32 ppq_size;
+ u32 rcq_size;
+
+ u16 num_napi;
+ u16 num_tx_queues;
+ u16 num_rx_queues;
+};
+
+struct net_device *mpnic_netdev_alloc(struct mpnic_dev *mpd);
+void mpnic_netdev_free(struct mpnic_dev *mpd);
+int mpnic_netdev_register(struct net_device *netdev);
+
+#endif /* _MPNIC_NETDEV_H_ */
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_pci.c b/drivers/net/ethernet/meta/mpnic/mpnic_pci.c
new file mode 100644
index 0000000000000..968cd611b8eab
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_pci.c
@@ -0,0 +1,187 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#include <linux/dma-mapping.h>
+#include <linux/err.h>
+#include <linux/module.h>
+#include <linux/netdevice.h>
+#include <linux/pci.h>
+#include <linux/slab.h>
+#include <linux/types.h>
+
+#include "mpnic.h"
+#include "mpnic_netdev.h"
+
+#define PCI_DEVICE_ID_META_MPNIC 0x0014
+
+static void mpnic_mmio_err(struct mpnic_dev *mpd, u32 reg)
+{
+ /* Hardware is giving us all 1's reads, assume it is gone */
+ WRITE_ONCE(mpd->uc_addr0, NULL);
+
+ dev_err(mpd->dev,
+ "Failed read (idx 0x%x AKA addr 0x%x), disabled CSR access, awaiting reset\n",
+ reg, reg << 2);
+
+ /* Tell the stack the device has lost its PCIe link */
+ if (mpd->netdev)
+ netif_device_detach(mpd->netdev);
+}
+
+u64 mpnic_rd64(struct mpnic_dev *mpd, u32 reg)
+{
+ u32 __iomem *csr = READ_ONCE(mpd->uc_addr0);
+ u64 value;
+
+ if (!csr)
+ return ~0ULL;
+
+ value = readq(csr + reg);
+
+ /* If any bits are 0 value should be valid */
+ if (~value)
+ return value;
+
+ /* All ones can be a valid value, so confirm against a register
+ * which never reads that way on a live device.
+ */
+ if (reg != MPNIC_BDQ_SPARE && ~readq(csr + MPNIC_BDQ_SPARE))
+ return value;
+
+ mpnic_mmio_err(mpd, reg);
+
+ return ~0ULL;
+}
+
+static struct mpnic_dev *mpnic_alloc(struct pci_dev *pdev)
+{
+ struct mpnic_dev *mpd;
+
+ mpd = kzalloc_obj(*mpd);
+ if (!mpd)
+ return NULL;
+
+ pci_set_drvdata(pdev, mpd);
+ mpd->dev = &pdev->dev;
+
+ mpd->dsn = pci_get_dsn(pdev);
+ mpd->mps = pcie_get_mps(pdev);
+ mpd->readrq = pcie_get_readrq(pdev);
+ mpd->relaxed_ord = pcie_relaxed_ordering_enabled(pdev);
+
+ return mpd;
+}
+
+/**
+ * mpnic_probe - Device initialization routine
+ * @pdev: PCI device information struct
+ * @ent: entry in mpnic_pci_tbl
+ *
+ * Return: 0 on success, negative on failure
+ **/
+static int mpnic_probe(struct pci_dev *pdev, const struct pci_device_id *ent)
+{
+ struct net_device *netdev;
+ void __iomem *uc_addr0;
+ struct mpnic_dev *mpd;
+ int err;
+
+ if (pdev->error_state != pci_channel_io_normal) {
+ dev_err(&pdev->dev,
+ "PCI device still in an error state. Unable to load...\n");
+ return -EIO;
+ }
+
+ err = pcim_enable_device(pdev);
+ if (err) {
+ dev_err(&pdev->dev, "PCI enable device failed: %d\n", err);
+ return err;
+ }
+
+ err = dma_set_mask_and_coherent(&pdev->dev, DMA_BIT_MASK(46));
+ if (err) {
+ dev_err(&pdev->dev, "DMA configuration failed: %d\n", err);
+ return err;
+ }
+
+ mpd = mpnic_alloc(pdev);
+ if (!mpd)
+ return -ENOMEM;
+
+ uc_addr0 = pcim_iomap_region(pdev, 0, MPNIC_DRV_NAME);
+ if (IS_ERR(uc_addr0)) {
+ err = PTR_ERR(uc_addr0);
+ dev_err(&pdev->dev, "Mapping the register file failed: %d\n",
+ err);
+ goto err_free_mpd;
+ }
+ mpd->uc_addr0 = uc_addr0;
+
+ pci_set_master(pdev);
+ pci_save_state(pdev);
+
+ err = mpnic_alloc_irqs(mpd);
+ if (err)
+ goto err_free_mpd;
+
+ err = mpnic_dev_init(mpd);
+ if (err)
+ goto err_free_irqs;
+
+ netdev = mpnic_netdev_alloc(mpd);
+ if (!netdev) {
+ dev_err(&pdev->dev, "Netdev allocation failed\n");
+ err = -ENOMEM;
+ goto err_free_irqs;
+ }
+
+ err = mpnic_netdev_register(netdev);
+ if (err) {
+ dev_err(&pdev->dev, "Netdev registration failed: %d\n", err);
+ goto err_free_netdev;
+ }
+
+ return 0;
+
+err_free_netdev:
+ mpnic_netdev_free(mpd);
+err_free_irqs:
+ mpnic_free_irqs(mpd);
+err_free_mpd:
+ kfree(mpd);
+
+ return err;
+}
+
+/**
+ * mpnic_remove - Device removal routine
+ * @pdev: PCI device information struct
+ **/
+static void mpnic_remove(struct pci_dev *pdev)
+{
+ struct mpnic_dev *mpd = pci_get_drvdata(pdev);
+
+ unregister_netdev(mpd->netdev);
+ mpnic_netdev_free(mpd);
+ mpnic_free_irqs(mpd);
+ kfree(mpd);
+}
+
+static const struct pci_device_id mpnic_pci_tbl[] = {
+ { PCI_VDEVICE(META, PCI_DEVICE_ID_META_MPNIC) },
+ /* required last entry */
+ {}
+};
+MODULE_DEVICE_TABLE(pci, mpnic_pci_tbl);
+
+static struct pci_driver mpnic_driver = {
+ .name = MPNIC_DRV_NAME,
+ .id_table = mpnic_pci_tbl,
+ .probe = mpnic_probe,
+ .remove = mpnic_remove,
+};
+
+module_pci_driver(mpnic_driver);
+
+MODULE_DESCRIPTION("Meta Platforms Network Interface Controller");
+MODULE_LICENSE("GPL");
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c
new file mode 100644
index 0000000000000..9878ea5a2f8e1
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.c
@@ -0,0 +1,1398 @@
+// SPDX-License-Identifier: GPL-2.0
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#include <linux/bitfield.h>
+#include <linux/dma-mapping.h>
+#include <linux/iopoll.h>
+#include <linux/pci.h>
+#include <linux/slab.h>
+#include <net/page_pool/helpers.h>
+
+#include "mpnic.h"
+#include "mpnic_netdev.h"
+#include "mpnic_txrx.h"
+
+struct mpnic_xmit_cb {
+ u32 bytecount;
+ u8 desc_count;
+};
+
+#define MPNIC_XMIT_CB(__skb) ((struct mpnic_xmit_cb *)((__skb)->cb))
+#define MPNIC_TWD_TYPE_MASK(_type) \
+ cpu_to_le64(FIELD_PREP(MPNIC_TWD_TYPE, MPNIC_TWD_TYPE_##_type))
+
+/* Leave the interrupt moderation counters alone when arming or masking */
+#define MPNIC_TIM_PARAM_CFG_PRESERVE_MASK \
+ (MPNIC_TIM_CTL1_UPD_IGN_LONG_EVENT_CNT | \
+ MPNIC_TIM_CTL1_UPD_IGN_LONG_TIME_CNT | \
+ MPNIC_TIM_CTL1_UPD_IGN_SHORT_TIME_CNT)
+
+static void mpnic_nv_irq_disable(struct mpnic_napi_vector *nv)
+{
+ mpnic_wr64(nv->mpd, MPNIC_TIM_CTL1(nv->qt[0].cmpl.q_idx),
+ MPNIC_TIM_PARAM_CFG_PRESERVE_MASK |
+ MPNIC_TIM_CTL1_MASK_EN | MPNIC_TIM_CTL1_MASK);
+}
+
+static void mpnic_nv_irq_rearm(struct mpnic_napi_vector *nv)
+{
+ /* Rearming a single queue on a given IRQ rearms all the other
+ * queues mapped to the same IRQ.
+ */
+ mpnic_wr64(nv->mpd, MPNIC_TIM_CTL1(nv->qt[0].cmpl.q_idx),
+ MPNIC_TIM_PARAM_CFG_PRESERVE_MASK | MPNIC_TIM_CTL1_MASK_EN);
+}
+
+static unsigned int mpnic_desc_unused(struct mpnic_ring *ring)
+{
+ return (ring->head - ring->tail - 1) & ring->size_mask;
+}
+
+static struct netdev_queue *mpnic_txring_txq(const struct net_device *dev,
+ const struct mpnic_ring *ring)
+{
+ return netdev_get_tx_queue(dev, ring->q_idx);
+}
+
+static void mpnic_tx_doorbell(struct mpnic_ring *ring, __le64 *meta)
+{
+ *meta |= cpu_to_le64(MPNIC_TWD_FLAG_REQ_COMPLETION);
+ ring->deferred_meta = -1;
+
+ /* Force DMA writes to flush before writing to tail */
+ dma_wmb();
+
+ writeq(ring->tail, ring->doorbell);
+}
+
+/* Packets handed to us with xmit_more set are left in the ring without a
+ * doorbell, and without a completion request, in the expectation that the
+ * packet ending the burst will ring for all of them. If that packet gets
+ * dropped instead we have to ring here, otherwise the descriptors sit in
+ * the ring until the next transmit, which may never come.
+ */
+static void mpnic_tx_flush_doorbell(struct mpnic_ring *ring)
+{
+ if (ring->deferred_meta >= 0)
+ mpnic_tx_doorbell(ring, &ring->desc[ring->deferred_meta]);
+}
+
+static void mpnic_unmap_single_twd(struct device *dev, __le64 *twd)
+{
+ u64 raw_twd = le64_to_cpu(*twd);
+
+ dma_unmap_single(dev, FIELD_GET(MPNIC_TWD_ADDR, raw_twd),
+ FIELD_GET(MPNIC_TWD_LEN, raw_twd), DMA_TO_DEVICE);
+}
+
+static void mpnic_unmap_page_twd(struct device *dev, __le64 *twd)
+{
+ u64 raw_twd = le64_to_cpu(*twd);
+
+ dma_unmap_page(dev, FIELD_GET(MPNIC_TWD_ADDR, raw_twd),
+ FIELD_GET(MPNIC_TWD_LEN, raw_twd), DMA_TO_DEVICE);
+}
+
+static bool
+mpnic_tx_map(struct mpnic_ring *ring, struct sk_buff *skb, __le64 *meta)
+{
+ struct device *dev = skb->dev->dev.parent;
+ unsigned int tail = ring->tail, first;
+ unsigned int size, data_len;
+ skb_frag_t *frag;
+ dma_addr_t dma;
+ __le64 *twd;
+
+ tail++;
+ tail &= ring->size_mask;
+ first = tail;
+
+ size = skb_headlen(skb);
+ data_len = skb->data_len;
+
+ if (size > FIELD_MAX(MPNIC_TWD_LEN))
+ goto err_dma;
+
+ dma = dma_map_single(dev, skb->data, size, DMA_TO_DEVICE);
+
+ for (frag = &skb_shinfo(skb)->frags[0];; frag++) {
+ twd = &ring->desc[tail];
+
+ if (dma_mapping_error(dev, dma))
+ goto err_dma;
+
+ *twd = cpu_to_le64(FIELD_PREP(MPNIC_TWD_ADDR, dma) |
+ FIELD_PREP(MPNIC_TWD_LEN, size) |
+ FIELD_PREP(MPNIC_TWD_TYPE,
+ MPNIC_TWD_TYPE_AL));
+
+ tail++;
+ tail &= ring->size_mask;
+
+ if (!data_len)
+ break;
+
+ size = skb_frag_size(frag);
+ data_len -= size;
+
+ if (size > FIELD_MAX(MPNIC_TWD_LEN))
+ goto err_dma;
+
+ dma = skb_frag_dma_map(dev, frag, 0, size, DMA_TO_DEVICE);
+ }
+
+ *twd |= MPNIC_TWD_TYPE_MASK(LAST_AL);
+
+ MPNIC_XMIT_CB(skb)->desc_count = ((twd - meta) + 1) & ring->size_mask;
+
+ skb_tx_timestamp(skb);
+
+ ring->tail = tail;
+
+ /* Verify there is room for another packet */
+ netif_txq_maybe_stop(mpnic_txring_txq(skb->dev, ring),
+ mpnic_desc_unused(ring), MPNIC_MAX_SKB_DESC,
+ MPNIC_TX_DESC_WAKEUP);
+
+ if (__netdev_tx_sent_queue(mpnic_txring_txq(skb->dev, ring),
+ MPNIC_XMIT_CB(skb)->bytecount,
+ netdev_xmit_more()))
+ mpnic_tx_doorbell(ring, meta);
+ else
+ ring->deferred_meta = meta - ring->desc;
+
+ return false;
+err_dma:
+ if (net_ratelimit())
+ netdev_err(skb->dev, "TX DMA map failed\n");
+
+ while (tail != first) {
+ tail--;
+ tail &= ring->size_mask;
+ twd = &ring->desc[tail];
+ if (tail == first)
+ mpnic_unmap_single_twd(dev, twd);
+ else
+ mpnic_unmap_page_twd(dev, twd);
+ }
+
+ return true;
+}
+
+#define MPNIC_MIN_FRAME_LEN 60
+
+static netdev_tx_t mpnic_xmit_frame_ring(struct sk_buff *skb,
+ struct mpnic_ring *ring)
+{
+ __le64 *meta = &ring->desc[ring->tail];
+ u32 tail = ring->tail;
+
+ if (skb_put_padto(skb, MPNIC_MIN_FRAME_LEN))
+ goto err_drop;
+
+ if (!netif_txq_maybe_stop(mpnic_txring_txq(skb->dev, ring),
+ mpnic_desc_unused(ring), MPNIC_MAX_SKB_DESC,
+ MPNIC_TX_DESC_WAKEUP)) {
+ mpnic_tx_flush_doorbell(ring);
+ return NETDEV_TX_BUSY;
+ }
+
+ ring->tx_buf[tail] = skb;
+ *meta = cpu_to_le64(MPNIC_TWD_FLAG_DEST_MAC);
+
+ MPNIC_XMIT_CB(skb)->bytecount = skb->len;
+ MPNIC_XMIT_CB(skb)->desc_count = 0;
+
+ if (mpnic_tx_map(ring, skb, meta))
+ goto err_free;
+
+ return NETDEV_TX_OK;
+
+err_free:
+ dev_kfree_skb_any(skb);
+ ring->tx_buf[tail] = NULL;
+ ring->tail = tail;
+err_drop:
+ mpnic_tx_flush_doorbell(ring);
+
+ return NETDEV_TX_OK;
+}
+
+netdev_tx_t mpnic_xmit_frame(struct sk_buff *skb, struct net_device *dev)
+{
+ struct mpnic_net *mpn = netdev_priv(dev);
+
+ return mpnic_xmit_frame_ring(skb, mpn->tx[skb_get_queue_mapping(skb)]);
+}
+
+static void mpnic_clean_twq0(struct mpnic_napi_vector *nv, int napi_budget,
+ struct mpnic_ring *ring, bool discard,
+ unsigned int hw_head)
+{
+ u64 total_bytes = 0, total_packets = 0;
+ unsigned int head = ring->head;
+ struct netdev_queue *txq;
+ unsigned int clean_desc;
+
+ clean_desc = (hw_head - head) & ring->size_mask;
+
+ while (clean_desc) {
+ struct sk_buff *skb = ring->tx_buf[head];
+ unsigned int desc_cnt;
+
+ desc_cnt = MPNIC_XMIT_CB(skb)->desc_count;
+ if (desc_cnt > clean_desc)
+ break;
+
+ ring->tx_buf[head] = NULL;
+
+ clean_desc -= desc_cnt;
+
+ /* Step over the metadata descriptor */
+ head++;
+ head &= ring->size_mask;
+ desc_cnt--;
+
+ mpnic_unmap_single_twd(nv->dev, &ring->desc[head]);
+ head++;
+ head &= ring->size_mask;
+ desc_cnt--;
+
+ while (desc_cnt--) {
+ mpnic_unmap_page_twd(nv->dev, &ring->desc[head]);
+ head++;
+ head &= ring->size_mask;
+ }
+
+ total_bytes += MPNIC_XMIT_CB(skb)->bytecount;
+ total_packets++;
+
+ napi_consume_skb(skb, napi_budget);
+ }
+
+ if (!total_bytes)
+ return;
+
+ ring->head = head;
+
+ if (discard)
+ return;
+
+ txq = mpnic_txring_txq(nv->napi.dev, ring);
+ netif_txq_completed_wake(txq, total_packets, total_bytes,
+ mpnic_desc_unused(ring),
+ MPNIC_TX_DESC_WAKEUP);
+}
+
+static void mpnic_commit_cq_head(struct mpnic_ring *cmpl)
+{
+ u32 head = cmpl->head;
+
+ /* The tail shadows the last value written to the doorbell, so a
+ * completion queue which has not moved costs no MMIO write.
+ */
+ if (cmpl->tail != head) {
+ cmpl->tail = head;
+ writeq(head & cmpl->size_mask, cmpl->doorbell);
+ }
+}
+
+static void mpnic_clean_tcq(struct mpnic_napi_vector *nv,
+ struct mpnic_q_triad *qt, int napi_budget)
+{
+ struct mpnic_ring *cmpl = &qt->cmpl;
+ __le64 *raw_tcd, done;
+ u32 head = cmpl->head;
+ s32 head0 = -1;
+
+ done = (head & (cmpl->size_mask + 1)) ? 0 : cpu_to_le64(MPNIC_TCD_DONE);
+ raw_tcd = &cmpl->desc[head & cmpl->size_mask];
+
+ /* Walk the completion queue collecting the heads reported by NIC.
+ * Only the first work queue is enabled and no packet asks for a
+ * timestamp, so every completion is a plain head update and the
+ * descriptor type does not have to be decoded.
+ */
+ while ((*raw_tcd & cpu_to_le64(MPNIC_TCD_DONE)) == done) {
+ u64 tcd;
+
+ dma_rmb();
+
+ tcd = le64_to_cpu(*raw_tcd);
+ head0 = FIELD_GET(MPNIC_TCD_TYPE0_HEAD0, tcd);
+
+ raw_tcd++;
+ head++;
+
+ if (unlikely(!(head & cmpl->size_mask))) {
+ done ^= cpu_to_le64(MPNIC_TCD_DONE);
+ raw_tcd = &cmpl->desc[0];
+ }
+ }
+
+ cmpl->head = head;
+
+ if (head0 >= 0)
+ mpnic_clean_twq0(nv, napi_budget, &qt->sub0, false, head0);
+}
+
+static void mpnic_bd_prep(struct mpnic_ring *bdq, u32 idx, struct page *page)
+{
+ dma_addr_t dma = page_pool_get_dma_addr(page);
+
+ bdq->desc[idx] = cpu_to_le64(FIELD_PREP(MPNIC_BD_DESC_ADDR, dma >> 10) |
+ FIELD_PREP(MPNIC_BD_DESC_ID, idx) |
+ FIELD_PREP(MPNIC_BD_DESC_BUF_SZ_LOG2,
+ page_shift(page) - 10));
+}
+
+/* Descriptors are only handed to the device in whole batches, so the slot
+ * the device is working on and everything up to the next batch boundary
+ * stay untouched while it does.
+ */
+static unsigned int mpnic_bdq_desc_unused(struct mpnic_ring *bdq)
+{
+ return (ALIGN_DOWN(bdq->head - 1, MPNIC_BDQ_BATCH_SIZE) - bdq->tail) &
+ bdq->size_mask;
+}
+
+static unsigned int __mpnic_fill_bdq(struct mpnic_ring *bdq)
+{
+ unsigned int i = bdq->tail;
+ unsigned int count;
+
+ for (count = mpnic_bdq_desc_unused(bdq); count; count--) {
+ struct page *page;
+
+ page = page_pool_dev_alloc_pages(bdq->page_pool);
+ if (!page)
+ break;
+
+ bdq->rx_buf[i] = page;
+ mpnic_bd_prep(bdq, i, page);
+
+ i++;
+ i &= bdq->size_mask;
+ }
+
+ return i;
+}
+
+static void __mpnic_bdq_commit_tail(struct mpnic_ring *bdq, unsigned int tail)
+{
+ if (bdq->tail != tail) {
+ bdq->tail = tail;
+
+ writeq(tail, bdq->doorbell);
+ }
+}
+
+static void mpnic_fill_qt_bdqs(struct mpnic_q_triad *qt)
+{
+ unsigned int ppq_i = __mpnic_fill_bdq(&qt->sub1);
+ unsigned int hpq_i = __mpnic_fill_bdq(&qt->sub0);
+
+ /* Force DMA writes to flush before writing to tail(s) */
+ dma_wmb();
+
+ /* Flush out the completions we are done with */
+ mpnic_commit_cq_head(&qt->cmpl);
+
+ __mpnic_bdq_commit_tail(&qt->sub0, hpq_i);
+ __mpnic_bdq_commit_tail(&qt->sub1, ppq_i);
+}
+
+/* Take one of the references batched on the page at @idx. If the device
+ * has moved on to a new page, first drop the unused references left on
+ * the previous one.
+ */
+static struct page *
+mpnic_page_pool_get(struct mpnic_pg_ctxt *pg_ctxt, struct mpnic_ring *ring,
+ u32 idx)
+{
+ struct page *page = pg_ctxt->page;
+
+ if (unlikely(pg_ctxt->idx != idx)) {
+ if (pg_ctxt->pagecnt_bias &&
+ !page_pool_unref_page(page, pg_ctxt->pagecnt_bias))
+ page_pool_put_unrefed_page(page->pp, page, -1, true);
+
+ page = ring->rx_buf[idx];
+ page_pool_fragment_page(page, MPNIC_PAGECNT_BIAS_MAX);
+
+ pg_ctxt->page = page;
+ pg_ctxt->pagecnt_bias = MPNIC_PAGECNT_BIAS_MAX;
+ pg_ctxt->idx = idx;
+ }
+
+ pg_ctxt->pagecnt_bias--;
+
+ return page;
+}
+
+static void mpnic_flush_pg_ctxt(struct mpnic_pg_ctxt *ctxt, bool napi)
+{
+ long pagecnt_bias = ctxt->pagecnt_bias;
+
+ if (pagecnt_bias) {
+ struct page *page = ctxt->page;
+
+ if (!page_pool_unref_page(page, pagecnt_bias))
+ page_pool_put_unrefed_page(page->pp, page, -1, napi);
+ }
+}
+
+static unsigned int mpnic_hdr_pg_start(unsigned int pg_off)
+{
+ /* The headroom of the first header may be larger than
+ * MPNIC_RX_HROOM due to alignment. So account for that by just
+ * making the page offset 0 if we are starting at the first header.
+ */
+ if (ALIGN(MPNIC_RX_HROOM, 128) > MPNIC_RX_HROOM &&
+ pg_off == ALIGN(MPNIC_RX_HROOM, 128))
+ return 0;
+
+ return pg_off - MPNIC_RX_HROOM;
+}
+
+static unsigned int mpnic_hdr_pg_end(unsigned int pg_off, unsigned int len)
+{
+ /* Determine the end of the buffer by finding the start of the next
+ * and then subtracting the headroom from that frame.
+ */
+ pg_off += len + MPNIC_RX_TROOM + MPNIC_RX_HROOM;
+
+ return ALIGN(pg_off, 128) - MPNIC_RX_HROOM;
+}
+
+static void
+mpnic_pkt_prepare(struct mpnic_napi_vector *nv, u64 rcd,
+ struct mpnic_rcq_state *state, struct mpnic_q_triad *qt)
+{
+ unsigned int pg_off = FIELD_GET(MPNIC_RCD_AL_BUFF_OFF, rcd);
+ unsigned int pg_idx = FIELD_GET(MPNIC_RCD_AL_BUFF_ID, rcd);
+ unsigned int len = FIELD_GET(MPNIC_RCD_AL_BUFF_LEN, rcd);
+ bool fin = FIELD_GET(MPNIC_RCD_AL_PAGE_FIN, rcd);
+ unsigned int frame_sz, pg_start, pg_end;
+ struct xdp_buff *buff = &state->pkt;
+ struct page *page;
+
+ pg_start = mpnic_hdr_pg_start(pg_off);
+
+ page = mpnic_page_pool_get(&state->hdr, &qt->sub0, pg_idx);
+ qt->sub0.head = (pg_idx + 1) & qt->sub0.size_mask;
+
+ /* Short-cut the end calculation if the page is fully consumed */
+ pg_end = fin ? page_size(page) : mpnic_hdr_pg_end(pg_off, len);
+ frame_sz = pg_end - pg_start;
+
+ dma_sync_single_range_for_cpu(nv->dev, page_pool_get_dma_addr(page),
+ pg_start, frame_sz, DMA_FROM_DEVICE);
+
+ xdp_init_buff(buff, frame_sz, &qt->xdp_rxq);
+ xdp_prepare_buff(buff, page_address(page) + pg_start,
+ pg_off - pg_start, len, true);
+ net_prefetch(buff->data);
+
+ state->add_frag_failed = false;
+}
+
+static void
+mpnic_add_rx_frag(struct mpnic_napi_vector *nv, u64 rcd,
+ struct mpnic_rcq_state *state, struct mpnic_q_triad *qt)
+{
+ unsigned int pg_off = FIELD_GET(MPNIC_RCD_AL_BUFF_OFF, rcd);
+ unsigned int pg_idx = FIELD_GET(MPNIC_RCD_AL_BUFF_ID, rcd);
+ unsigned int len = FIELD_GET(MPNIC_RCD_AL_BUFF_LEN, rcd);
+ bool fin = FIELD_GET(MPNIC_RCD_AL_PAGE_FIN, rcd);
+ struct xdp_buff *buff = &state->pkt;
+ unsigned int truesz;
+ struct page *page;
+
+ page = mpnic_page_pool_get(&state->payld, &qt->sub1, pg_idx);
+ qt->sub1.head = (pg_idx + 1) & qt->sub1.size_mask;
+
+ truesz = (fin ? page_size(page) : ALIGN(pg_off + len, 128)) - pg_off;
+
+ dma_sync_single_range_for_cpu(nv->dev, page_pool_get_dma_addr(page),
+ pg_off, truesz, DMA_FROM_DEVICE);
+
+ if (!xdp_buff_add_frag(buff, page_to_netmem(page), pg_off, len,
+ truesz)) {
+ state->payld.pagecnt_bias++;
+ state->add_frag_failed = true;
+ }
+}
+
+static void mpnic_put_pkt_buff(struct xdp_buff *buff, bool napi)
+{
+ struct page *page;
+
+ if (!buff->data_hard_start)
+ return;
+
+ if (unlikely(xdp_buff_has_frags(buff))) {
+ struct skb_shared_info *shinfo;
+ int nr_frags;
+
+ shinfo = xdp_get_shared_info_from_buff(buff);
+ nr_frags = shinfo->nr_frags;
+
+ while (nr_frags--) {
+ page = skb_frag_page(&shinfo->frags[nr_frags]);
+ page_pool_put_full_page(page->pp, page, napi);
+ }
+ }
+
+ page = virt_to_head_page(buff->data_hard_start);
+ page_pool_put_full_page(page->pp, page, napi);
+}
+
+static int mpnic_clean_rcq(struct mpnic_napi_vector *nv,
+ struct mpnic_q_triad *qt, int budget)
+{
+ struct mpnic_ring *rcq = &qt->cmpl;
+ struct mpnic_rcq_state *state;
+ unsigned int packets = 0;
+ __le64 *raw_rcd, done;
+ u32 head = rcq->head;
+
+ done = (head & (rcq->size_mask + 1)) ? 0 : cpu_to_le64(MPNIC_RCD_DONE);
+ raw_rcd = &rcq->desc[head & rcq->size_mask];
+ state = rcq->state;
+
+ while (packets < budget) {
+ u64 rcd;
+
+ if ((*raw_rcd & cpu_to_le64(MPNIC_RCD_DONE)) != done)
+ break;
+
+ dma_rmb();
+
+ rcd = le64_to_cpu(*raw_rcd);
+
+ switch (FIELD_GET(MPNIC_RCD_TYPE, rcd)) {
+ case MPNIC_RCD_TYPE_HDR_AL:
+ if (FIELD_GET(MPNIC_RCD_HDR_SUBTYPE, rcd) ==
+ MPNIC_RCD_HDR_SUBTYPE_HDR)
+ mpnic_pkt_prepare(nv, rcd, state, qt);
+ break;
+ case MPNIC_RCD_TYPE_PAY_AL:
+ mpnic_add_rx_frag(nv, rcd, state, qt);
+ break;
+ case MPNIC_RCD_TYPE_META: {
+ struct sk_buff *skb = NULL;
+
+ if (likely(!(rcd &
+ MPNIC_RCD_META_UNCORRECTABLE_ERR_MASK) &&
+ !state->add_frag_failed))
+ skb = xdp_build_skb_from_buff(&state->pkt);
+
+ if (likely(skb))
+ napi_gro_receive(&nv->napi, skb);
+ else
+ mpnic_put_pkt_buff(&state->pkt, true);
+
+ state->pkt.data_hard_start = NULL;
+ packets++;
+ break;
+ }
+ }
+
+ raw_rcd++;
+ head++;
+
+ if (unlikely(!(head & rcq->size_mask))) {
+ done ^= cpu_to_le64(MPNIC_RCD_DONE);
+ raw_rcd = &rcq->desc[0];
+ }
+ }
+
+ rcq->head = head;
+
+ /* Allocate buffers, force dma_wmb(), and then start writing tails */
+ mpnic_fill_qt_bdqs(qt);
+
+ return packets;
+}
+
+static int mpnic_poll(struct napi_struct *napi, int budget)
+{
+ struct mpnic_napi_vector *nv = container_of(napi,
+ struct mpnic_napi_vector,
+ napi);
+ int i, j, work_done = 0;
+
+ for (i = 0; i < nv->txt_count; i++)
+ mpnic_clean_tcq(nv, &nv->qt[i], budget);
+
+ for (j = 0; j < nv->rxt_count; j++, i++)
+ work_done += mpnic_clean_rcq(nv, &nv->qt[i], budget);
+
+ for (i = 0; i < nv->txt_count; i++)
+ mpnic_commit_cq_head(&nv->qt[i].cmpl);
+
+ if (work_done >= budget)
+ return budget;
+
+ if (likely(napi_complete_done(napi, work_done)))
+ mpnic_nv_irq_rearm(nv);
+
+ return work_done;
+}
+
+static irqreturn_t mpnic_msix_clean_rings(int __always_unused irq, void *data)
+{
+ struct mpnic_napi_vector *nv = data;
+
+ napi_schedule_irqoff(&nv->napi);
+
+ return IRQ_HANDLED;
+}
+
+static void mpnic_free_napi_vector(struct mpnic_net *mpn,
+ struct mpnic_napi_vector *nv)
+{
+ int i, j;
+
+ for (i = 0; i < nv->txt_count; i++)
+ mpn->tx[nv->qt[i].sub0.q_idx] = NULL;
+
+ for (j = 0; j < nv->rxt_count; j++, i++)
+ mpn->rx[nv->qt[i].cmpl.q_idx] = NULL;
+
+ mpnic_free_irq(nv->mpd, nv->v_idx, nv);
+ netif_napi_del_locked(&nv->napi);
+ mpn->napi[nv->v_idx - MPNIC_NON_NAPI_VECTORS] = NULL;
+ kfree(nv);
+}
+
+void mpnic_free_napi_vectors(struct mpnic_net *mpn)
+{
+ int i;
+
+ for (i = 0; i < mpn->num_napi; i++)
+ if (mpn->napi[i])
+ mpnic_free_napi_vector(mpn, mpn->napi[i]);
+}
+
+static void mpnic_ring_init(struct mpnic_ring *ring, u32 __iomem *doorbell,
+ int q_idx)
+{
+ ring->doorbell = doorbell;
+ ring->q_idx = q_idx;
+}
+
+static int mpnic_alloc_napi_vector(struct mpnic_dev *mpd,
+ struct mpnic_net *mpn, unsigned int idx)
+{
+ u32 __iomem *uc_addr = READ_ONCE(mpd->uc_addr0);
+ struct mpnic_napi_vector *nv;
+ int err;
+
+ /* Doorbells are plain pointers into the register window, they have
+ * no way of noticing that it went away.
+ */
+ if (!uc_addr)
+ return -EIO;
+
+ nv = kzalloc_flex(*nv, qt, 2);
+ if (!nv)
+ return -ENOMEM;
+
+ nv->txt_count = 1;
+ nv->rxt_count = 1;
+ nv->mpd = mpd;
+ nv->dev = mpd->dev;
+ nv->v_idx = idx + MPNIC_NON_NAPI_VECTORS;
+
+ mpn->napi[idx] = nv;
+ netif_napi_add_config_locked(mpn->netdev, &nv->napi, mpnic_poll, idx);
+ netif_napi_set_irq_locked(&nv->napi,
+ pci_irq_vector(to_pci_dev(mpd->dev),
+ nv->v_idx));
+
+ snprintf(nv->name, sizeof(nv->name), "%s-TxRx-%u",
+ mpn->netdev->name, idx);
+
+ err = mpnic_request_irq(mpd, nv->v_idx, mpnic_msix_clean_rings, 0,
+ nv->name, nv);
+ if (err)
+ goto err_napi_del;
+
+ mpnic_ring_init(&nv->qt[0].sub0, &uc_addr[MPNIC_TWQ_TAIL(idx, 0)], idx);
+ mpnic_ring_init(&nv->qt[0].cmpl, &uc_addr[MPNIC_TCQ_HEAD(idx)], idx);
+ mpn->tx[idx] = &nv->qt[0].sub0;
+
+ mpnic_ring_init(&nv->qt[1].sub0, &uc_addr[MPNIC_HPQ_TAIL(idx)], idx);
+ mpnic_ring_init(&nv->qt[1].sub1, &uc_addr[MPNIC_PPQ_TAIL(idx)], idx);
+ mpnic_ring_init(&nv->qt[1].cmpl, &uc_addr[MPNIC_RCQ_HEAD(idx)], idx);
+ mpn->rx[idx] = &nv->qt[1].cmpl;
+
+ return 0;
+
+err_napi_del:
+ netif_napi_del_locked(&nv->napi);
+ mpn->napi[idx] = NULL;
+ kfree(nv);
+ return err;
+}
+
+int mpnic_alloc_napi_vectors(struct mpnic_net *mpn)
+{
+ unsigned int i;
+ int err;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ err = mpnic_alloc_napi_vector(mpn->mpd, mpn, i);
+ if (err)
+ goto err_free_vectors;
+ }
+
+ return 0;
+
+err_free_vectors:
+ mpnic_free_napi_vectors(mpn);
+
+ return err;
+}
+
+static void mpnic_free_ring_resources(struct device *dev,
+ struct mpnic_ring *ring)
+{
+ kvfree(ring->buffer);
+ ring->buffer = NULL;
+
+ /* If size is not set there are no descriptors present */
+ if (!ring->size)
+ return;
+
+ dma_free_coherent(dev, ring->size, ring->desc, ring->dma);
+ ring->size_mask = 0;
+ ring->size = 0;
+}
+
+static int mpnic_alloc_ring_desc(struct mpnic_net *mpn,
+ struct mpnic_ring *ring, u32 count)
+{
+ struct device *dev = mpn->netdev->dev.parent;
+ size_t size;
+
+ size = ALIGN(array_size(sizeof(*ring->desc), count), 4096);
+
+ ring->desc = dma_alloc_coherent(dev, size, &ring->dma,
+ GFP_KERNEL | __GFP_NOWARN);
+ if (!ring->desc)
+ return -ENOMEM;
+
+ ring->size_mask = count - 1;
+ ring->size = size;
+
+ return 0;
+}
+
+static void mpnic_free_tx_qt_resources(struct mpnic_net *mpn,
+ struct mpnic_q_triad *qt)
+{
+ struct device *dev = mpn->netdev->dev.parent;
+
+ mpnic_free_ring_resources(dev, &qt->cmpl);
+ mpnic_free_ring_resources(dev, &qt->sub0);
+}
+
+static int mpnic_alloc_tx_qt_resources(struct mpnic_net *mpn,
+ struct mpnic_q_triad *qt)
+{
+ int err;
+
+ err = mpnic_alloc_ring_desc(mpn, &qt->sub0, mpn->txq_size);
+ if (err)
+ return err;
+
+ qt->sub0.tx_buf = kvzalloc_objs(*qt->sub0.tx_buf, mpn->txq_size,
+ GFP_KERNEL | __GFP_NOWARN);
+ if (!qt->sub0.tx_buf) {
+ err = -ENOMEM;
+ goto err_free_qt;
+ }
+
+ err = mpnic_alloc_ring_desc(mpn, &qt->cmpl, mpn->txq_size);
+ if (err)
+ goto err_free_qt;
+
+ return 0;
+
+err_free_qt:
+ mpnic_free_tx_qt_resources(mpn, qt);
+ return err;
+}
+
+static int
+mpnic_alloc_qt_page_pool(struct mpnic_net *mpn, struct mpnic_napi_vector *nv,
+ struct mpnic_q_triad *qt)
+{
+ struct page_pool_params pp_params = {
+ .flags = PP_FLAG_DMA_MAP | PP_FLAG_DMA_SYNC_DEV,
+ .pool_size = min(mpn->hpq_size + mpn->ppq_size, 32768u),
+ .nid = NUMA_NO_NODE,
+ .dev = nv->dev,
+ .dma_dir = DMA_FROM_DEVICE,
+ .max_len = PAGE_SIZE,
+ .napi = &nv->napi,
+ .netdev = mpn->netdev,
+ .queue_idx = qt->cmpl.q_idx,
+ };
+ struct page_pool *pp;
+
+ pp = page_pool_create(&pp_params);
+ if (IS_ERR(pp))
+ return PTR_ERR(pp);
+
+ qt->sub0.page_pool = pp;
+ page_pool_get(pp);
+ qt->sub1.page_pool = pp;
+
+ return 0;
+}
+
+static void mpnic_free_rx_qt_resources(struct mpnic_net *mpn,
+ struct mpnic_q_triad *qt)
+{
+ struct device *dev = mpn->netdev->dev.parent;
+
+ mpnic_free_ring_resources(dev, &qt->cmpl);
+ mpnic_free_ring_resources(dev, &qt->sub1);
+ mpnic_free_ring_resources(dev, &qt->sub0);
+
+ if (xdp_rxq_info_is_reg(&qt->xdp_rxq)) {
+ xdp_rxq_info_unreg(&qt->xdp_rxq);
+ page_pool_destroy(qt->sub1.page_pool);
+ page_pool_destroy(qt->sub0.page_pool);
+ }
+}
+
+static int mpnic_alloc_rx_qt_resources(struct mpnic_net *mpn,
+ struct mpnic_napi_vector *nv,
+ struct mpnic_q_triad *qt)
+{
+ int err;
+
+ err = mpnic_alloc_qt_page_pool(mpn, nv, qt);
+ if (err)
+ return err;
+
+ err = xdp_rxq_info_reg(&qt->xdp_rxq, mpn->netdev, qt->cmpl.q_idx,
+ nv->napi.napi_id);
+ if (err)
+ goto err_free_page_pool;
+
+ err = xdp_rxq_info_reg_mem_model(&qt->xdp_rxq, MEM_TYPE_PAGE_POOL,
+ qt->sub0.page_pool);
+ if (err)
+ goto err_unreg_rxq;
+
+ err = mpnic_alloc_ring_desc(mpn, &qt->sub0, mpn->hpq_size);
+ if (err)
+ goto err_unreg_mm;
+
+ qt->sub0.rx_buf = kvzalloc_objs(*qt->sub0.rx_buf, mpn->hpq_size,
+ GFP_KERNEL | __GFP_NOWARN);
+ if (!qt->sub0.rx_buf) {
+ err = -ENOMEM;
+ goto err_free_qt;
+ }
+
+ err = mpnic_alloc_ring_desc(mpn, &qt->sub1, mpn->ppq_size);
+ if (err)
+ goto err_free_qt;
+
+ qt->sub1.rx_buf = kvzalloc_objs(*qt->sub1.rx_buf, mpn->ppq_size,
+ GFP_KERNEL | __GFP_NOWARN);
+ if (!qt->sub1.rx_buf) {
+ err = -ENOMEM;
+ goto err_free_qt;
+ }
+
+ err = mpnic_alloc_ring_desc(mpn, &qt->cmpl, mpn->rcq_size);
+ if (err)
+ goto err_free_qt;
+
+ qt->cmpl.state = kvzalloc_obj(*qt->cmpl.state,
+ GFP_KERNEL | __GFP_NOWARN);
+ if (!qt->cmpl.state) {
+ err = -ENOMEM;
+ goto err_free_qt;
+ }
+
+ return 0;
+
+err_free_qt:
+ mpnic_free_rx_qt_resources(mpn, qt);
+ return err;
+err_unreg_mm:
+ xdp_rxq_info_unreg_mem_model(&qt->xdp_rxq);
+err_unreg_rxq:
+ xdp_rxq_info_unreg(&qt->xdp_rxq);
+err_free_page_pool:
+ page_pool_destroy(qt->sub1.page_pool);
+ page_pool_destroy(qt->sub0.page_pool);
+ return err;
+}
+
+static void mpnic_free_nv_resources(struct mpnic_net *mpn,
+ struct mpnic_napi_vector *nv)
+{
+ int i, j;
+
+ for (i = 0; i < nv->txt_count; i++)
+ mpnic_free_tx_qt_resources(mpn, &nv->qt[i]);
+
+ for (j = 0; j < nv->rxt_count; j++, i++)
+ mpnic_free_rx_qt_resources(mpn, &nv->qt[i]);
+}
+
+static int mpnic_alloc_nv_resources(struct mpnic_net *mpn,
+ struct mpnic_napi_vector *nv)
+{
+ int i, j, err;
+
+ for (i = 0; i < nv->txt_count; i++) {
+ err = mpnic_alloc_tx_qt_resources(mpn, &nv->qt[i]);
+ if (err)
+ goto err_free_qt_resources;
+ }
+
+ for (j = 0; j < nv->rxt_count; j++, i++) {
+ err = mpnic_alloc_rx_qt_resources(mpn, nv, &nv->qt[i]);
+ if (err)
+ goto err_free_qt_resources;
+ }
+
+ return 0;
+
+err_free_qt_resources:
+ while (i--) {
+ if (i < nv->txt_count)
+ mpnic_free_tx_qt_resources(mpn, &nv->qt[i]);
+ else
+ mpnic_free_rx_qt_resources(mpn, &nv->qt[i]);
+ }
+ return err;
+}
+
+void mpnic_free_resources(struct mpnic_net *mpn)
+{
+ int i;
+
+ for (i = 0; i < mpn->num_napi; i++)
+ mpnic_free_nv_resources(mpn, mpn->napi[i]);
+}
+
+int mpnic_alloc_resources(struct mpnic_net *mpn)
+{
+ int i, err;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ err = mpnic_alloc_nv_resources(mpn, mpn->napi[i]);
+ if (err)
+ goto err_free_resources;
+ }
+
+ return 0;
+
+err_free_resources:
+ while (i--)
+ mpnic_free_nv_resources(mpn, mpn->napi[i]);
+
+ return err;
+}
+
+static void mpnic_set_netif_napi(struct mpnic_napi_vector *nv,
+ struct napi_struct *napi)
+{
+ int i, j;
+
+ for (i = 0; i < nv->txt_count; i++)
+ netif_queue_set_napi(nv->napi.dev, nv->qt[i].sub0.q_idx,
+ NETDEV_QUEUE_TYPE_TX, napi);
+
+ for (j = 0; j < nv->rxt_count; j++, i++)
+ netif_queue_set_napi(nv->napi.dev, nv->qt[i].cmpl.q_idx,
+ NETDEV_QUEUE_TYPE_RX, napi);
+}
+
+int mpnic_set_netif_queues(struct mpnic_net *mpn)
+{
+ int i, err;
+
+ err = netif_set_real_num_queues(mpn->netdev, mpn->num_tx_queues,
+ mpn->num_rx_queues);
+ if (err)
+ return err;
+
+ for (i = 0; i < mpn->num_napi; i++)
+ mpnic_set_netif_napi(mpn->napi[i], &mpn->napi[i]->napi);
+
+ return 0;
+}
+
+void mpnic_reset_netif_queues(struct mpnic_net *mpn)
+{
+ int i;
+
+ for (i = 0; i < mpn->num_napi; i++)
+ mpnic_set_netif_napi(mpn->napi[i], NULL);
+}
+
+static void mpnic_enable_twq(struct mpnic_dev *mpd, struct mpnic_ring *twq)
+{
+ u32 log_size = fls(twq->size_mask);
+ u32 i = twq->q_idx;
+
+ /* Reset head/tail */
+ mpnic_wr64(mpd, MPNIC_TWQ_CTL(i, 0), MPNIC_TWQ_CTL_RESET);
+ twq->tail = 0;
+ twq->head = 0;
+ twq->deferred_meta = -1;
+
+ /* Store descriptor ring address and size */
+ mpnic_wr64(mpd, MPNIC_TWQ_BASE_ADDR(i, 0), twq->dma);
+ mpnic_wr64(mpd, MPNIC_TWQ_SIZE(i, 0), log_size & MPNIC_TWQ_SIZE_SIZE);
+
+ mpnic_wr64(mpd, MPNIC_TWQ_CTL(i, 0), MPNIC_TWQ_CTL_ENABLE);
+}
+
+static void mpnic_enable_tcq(struct mpnic_dev *mpd,
+ struct mpnic_napi_vector *nv,
+ struct mpnic_ring *tcq)
+{
+ u32 log_size = fls(tcq->size_mask);
+ u32 i = tcq->q_idx;
+
+ /* Reset head/tail */
+ mpnic_wr64(mpd, MPNIC_TCQ_CTL(i), MPNIC_TCQ_CTL_RESET);
+ tcq->tail = 0;
+ tcq->head = 0;
+
+ /* Store descriptor ring address and size */
+ mpnic_wr64(mpd, MPNIC_TCQ_BASE_ADDR(i), tcq->dma);
+ mpnic_wr64(mpd, MPNIC_TCQ_SIZE(i), log_size & MPNIC_TCQ_SIZE_SIZE);
+
+ /* Store interrupt information for the completion queue */
+ mpnic_wr64(mpd, MPNIC_TIM_CTL(i), nv->v_idx);
+ mpnic_wr64(mpd, MPNIC_TIM_INTR_MASK(i), 0);
+
+ mpnic_wr64(mpd, MPNIC_TCQ_CTL(i), MPNIC_TCQ_CTL_ENABLE);
+}
+
+static void mpnic_enable_bdq(struct mpnic_dev *mpd, struct mpnic_ring *hpq,
+ struct mpnic_ring *ppq)
+{
+ u32 hpq_log_size = fls(hpq->size_mask);
+ u32 ppq_log_size = fls(ppq->size_mask);
+ u32 i = hpq->q_idx;
+
+ /* Reset head/tail */
+ mpnic_wr64(mpd, MPNIC_BDQ_CTL(i), MPNIC_BDQ_CTL_RESET);
+ hpq->tail = 0;
+ hpq->head = 0;
+ ppq->tail = 0;
+ ppq->head = 0;
+
+ /* Store descriptor ring addresses and sizes */
+ mpnic_wr64(mpd, MPNIC_HPQ_BASE_ADDR(i), hpq->dma);
+ mpnic_wr64(mpd, MPNIC_HPQ_SIZE(i), hpq_log_size & MPNIC_HPQ_SIZE_SIZE);
+ mpnic_wr64(mpd, MPNIC_PPQ_BASE_ADDR(i), ppq->dma);
+ mpnic_wr64(mpd, MPNIC_PPQ_SIZE(i), ppq_log_size & MPNIC_PPQ_SIZE_SIZE);
+
+ mpnic_wr64(mpd, MPNIC_BDQ_CTL(i),
+ MPNIC_BDQ_CTL_ENABLE | MPNIC_BDQ_CTL_ENABLE_PPQ);
+}
+
+static void mpnic_set_rde_cfg(struct mpnic_dev *mpd, struct mpnic_ring *rcq)
+{
+ BUILD_BUG_ON(FIELD_MAX(MPNIC_RDE_CFG_MIN_HEAD_ROOM) < MPNIC_RX_HROOM);
+ BUILD_BUG_ON(FIELD_MAX(MPNIC_RDE_CFG_MIN_TAIL_ROOM) < MPNIC_RX_TROOM);
+
+ mpnic_wr64(mpd, MPNIC_RDE_CFG(rcq->q_idx),
+ FIELD_PREP(MPNIC_RDE_CFG_MIN_HEAD_ROOM, MPNIC_RX_HROOM) |
+ FIELD_PREP(MPNIC_RDE_CFG_MIN_TAIL_ROOM, MPNIC_RX_TROOM) |
+ FIELD_PREP(MPNIC_RDE_CFG_MAX_HEADER_BYTES,
+ MPNIC_RX_MAX_HDR));
+}
+
+static void mpnic_enable_rcq(struct mpnic_dev *mpd,
+ struct mpnic_napi_vector *nv,
+ struct mpnic_ring *rcq)
+{
+ u32 log_size = fls(rcq->size_mask);
+ u32 i = rcq->q_idx;
+
+ mpnic_set_rde_cfg(mpd, rcq);
+
+ /* Reset head/tail */
+ mpnic_wr64(mpd, MPNIC_RCQ_CTL(i), MPNIC_RCQ_CTL_RESET);
+ rcq->head = 0;
+ rcq->tail = 0;
+
+ /* Store descriptor ring address and size */
+ mpnic_wr64(mpd, MPNIC_RCQ_BASE_ADDR(i), rcq->dma);
+ mpnic_wr64(mpd, MPNIC_RCQ_SIZE(i), log_size & MPNIC_RCQ_SIZE_SIZE);
+
+ /* Store interrupt information for the completion queue */
+ mpnic_wr64(mpd, MPNIC_RIM_CTL(i), nv->v_idx);
+ mpnic_wr64(mpd, MPNIC_RIM_INTR_MASK(i), 0);
+
+ mpnic_wr64(mpd, MPNIC_RCQ_CTL(i), MPNIC_RCQ_CTL_ENABLE);
+}
+
+void mpnic_enable(struct mpnic_net *mpn)
+{
+ struct mpnic_dev *mpd = mpn->mpd;
+ int i, j, t;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ struct mpnic_napi_vector *nv = mpn->napi[i];
+
+ for (t = 0; t < nv->txt_count; t++) {
+ mpnic_enable_twq(mpd, &nv->qt[t].sub0);
+ mpnic_enable_tcq(mpd, nv, &nv->qt[t].cmpl);
+ }
+
+ for (j = 0; j < nv->rxt_count; j++, t++) {
+ mpnic_enable_bdq(mpd, &nv->qt[t].sub0, &nv->qt[t].sub1);
+ mpnic_enable_rcq(mpd, nv, &nv->qt[t].cmpl);
+ }
+ }
+
+ mpnic_wrfl(mpd);
+}
+
+static void mpnic_disable_twq(struct mpnic_dev *mpd, struct mpnic_ring *txr)
+{
+ u64 twq_ctl = mpnic_rd64(mpd, MPNIC_TWQ_CTL(txr->q_idx, 0));
+
+ twq_ctl &= ~MPNIC_TWQ_CTL_ENABLE;
+ mpnic_wr64(mpd, MPNIC_TWQ_CTL(txr->q_idx, 0), twq_ctl);
+}
+
+static void mpnic_disable_tcq(struct mpnic_dev *mpd, struct mpnic_ring *txr)
+{
+ mpnic_wr64(mpd, MPNIC_TCQ_CTL(txr->q_idx), 0);
+ mpnic_wr64(mpd, MPNIC_TIM_INTR_MASK(txr->q_idx),
+ MPNIC_TIM_INTR_MASK_MASK);
+}
+
+static void mpnic_disable_bdq(struct mpnic_dev *mpd, struct mpnic_ring *hpq)
+{
+ u64 bdq_ctl = mpnic_rd64(mpd, MPNIC_BDQ_CTL(hpq->q_idx));
+
+ bdq_ctl &= ~(MPNIC_BDQ_CTL_ENABLE | MPNIC_BDQ_CTL_ENABLE_PPQ);
+ mpnic_wr64(mpd, MPNIC_BDQ_CTL(hpq->q_idx), bdq_ctl);
+}
+
+static void mpnic_disable_rcq(struct mpnic_dev *mpd, struct mpnic_ring *rcq)
+{
+ mpnic_wr64(mpd, MPNIC_RCQ_CTL(rcq->q_idx), 0);
+ mpnic_wr64(mpd, MPNIC_RIM_INTR_MASK(rcq->q_idx),
+ MPNIC_RIM_INTR_MASK_MASK);
+}
+
+void mpnic_disable(struct mpnic_net *mpn)
+{
+ struct mpnic_dev *mpd = mpn->mpd;
+ int i, j, t;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ struct mpnic_napi_vector *nv = mpn->napi[i];
+
+ for (t = 0; t < nv->txt_count; t++) {
+ mpnic_disable_twq(mpd, &nv->qt[t].sub0);
+ mpnic_disable_tcq(mpd, &nv->qt[t].cmpl);
+ }
+
+ for (j = 0; j < nv->rxt_count; j++, t++) {
+ mpnic_disable_bdq(mpd, &nv->qt[t].sub0);
+ mpnic_disable_rcq(mpd, &nv->qt[t].cmpl);
+ }
+ }
+
+ mpnic_wrfl(mpd);
+}
+
+struct mpnic_idle_regs {
+ u32 reg_base;
+ u8 reg_cnt;
+ char name[4];
+};
+
+static u32 mpnic_non_idle_queues(struct mpnic_dev *mpd,
+ const struct mpnic_idle_regs *regs,
+ unsigned int nregs)
+{
+ u32 non_idle_bitmap = 0;
+ unsigned int i, j;
+
+ for (i = 0; i < nregs; i++) {
+ for (j = 0; j < regs[i].reg_cnt; j++) {
+ if (mpnic_rd64(mpd, regs[i].reg_base + 2 * j) !=
+ ~0ULL) {
+ non_idle_bitmap |= BIT(i);
+ break;
+ }
+ }
+ }
+
+ return non_idle_bitmap;
+}
+
+static void mpnic_idle_dump(struct mpnic_dev *mpd,
+ const struct mpnic_idle_regs *regs,
+ unsigned int nregs, u32 non_idle_bitmap, int err)
+{
+ unsigned int i, j;
+
+ dev_err(mpd->dev, "error waiting for queues idle %d\n", err);
+ for (i = 0; i < nregs; i++) {
+ if (!(non_idle_bitmap & BIT(i)))
+ continue;
+
+ dev_err(mpd->dev, "%s block not idle:\n", regs[i].name);
+ for (j = 0; j < regs[i].reg_cnt; j++)
+ dev_err(mpd->dev, " 0x%04x: %016llx\n",
+ regs[i].reg_base + 2 * j,
+ mpnic_rd64(mpd, regs[i].reg_base + 2 * j));
+ }
+}
+
+void mpnic_wait_all_queues_idle(struct mpnic_dev *mpd)
+{
+ static const struct mpnic_idle_regs queues[] = {
+ { MPNIC_TWQ_IDLE(0), MPNIC_TWQ_IDLE_CNT, "TWQ" },
+ { MPNIC_TQS_IDLE(0), MPNIC_TQS_IDLE_CNT, "TQS" },
+ { MPNIC_TDE_IDLE(0), MPNIC_TDE_IDLE_CNT, "TDE" },
+ { MPNIC_TCQ_IDLE(0), MPNIC_TCQ_IDLE_CNT, "TCQ" },
+ { MPNIC_HPQ_IDLE(0), MPNIC_HPQ_IDLE_CNT, "HPQ" },
+ { MPNIC_PPQ_IDLE(0), MPNIC_PPQ_IDLE_CNT, "PPQ" },
+ { MPNIC_RCQ_IDLE(0), MPNIC_RCQ_IDLE_CNT, "RCQ" },
+ };
+ u32 non_idle_bitmap;
+ int err;
+
+ err = read_poll_timeout(mpnic_non_idle_queues, non_idle_bitmap,
+ !non_idle_bitmap, 20, 500000, false, mpd,
+ queues, ARRAY_SIZE(queues));
+ if (err)
+ mpnic_idle_dump(mpd, queues, ARRAY_SIZE(queues),
+ non_idle_bitmap, err);
+}
+
+static void mpnic_clean_bdq(struct mpnic_ring *bdq)
+{
+ unsigned int head = bdq->head;
+
+ while (head != bdq->tail) {
+ struct page *page = bdq->rx_buf[head];
+
+ page_pool_put_full_page(page->pp, page, false);
+
+ head++;
+ head &= bdq->size_mask;
+ }
+
+ bdq->head = head;
+}
+
+void mpnic_flush(struct mpnic_net *mpn)
+{
+ int i, j, t;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ struct mpnic_napi_vector *nv = mpn->napi[i];
+
+ for (t = 0; t < nv->txt_count; t++) {
+ struct mpnic_q_triad *qt = &nv->qt[t];
+ struct netdev_queue *txq;
+
+ /* Clean the work queue of unprocessed work */
+ mpnic_clean_twq0(nv, 0, &qt->sub0, true, qt->sub0.tail);
+
+ txq = netdev_get_tx_queue(mpn->netdev, qt->sub0.q_idx);
+ netdev_tx_reset_queue(txq);
+ }
+
+ for (j = 0; j < nv->rxt_count; j++, t++) {
+ struct mpnic_q_triad *qt = &nv->qt[t];
+ struct mpnic_rcq_state *state = qt->cmpl.state;
+
+ /* Release the partially assembled frame and the
+ * pages the queues are still handing out.
+ */
+ mpnic_put_pkt_buff(&state->pkt, false);
+ mpnic_flush_pg_ctxt(&state->hdr, false);
+ mpnic_flush_pg_ctxt(&state->payld, false);
+ memset(state, 0, sizeof(*state));
+
+ mpnic_clean_bdq(&qt->sub0);
+ mpnic_clean_bdq(&qt->sub1);
+ }
+ }
+}
+
+void mpnic_fill(struct mpnic_net *mpn)
+{
+ int i, j, t;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ struct mpnic_napi_vector *nv = mpn->napi[i];
+
+ for (j = 0, t = nv->txt_count; j < nv->rxt_count; j++, t++) {
+ struct mpnic_q_triad *qt = &nv->qt[t];
+ struct mpnic_rcq_state *state = qt->cmpl.state;
+
+ /* Point the page contexts at an index the device
+ * cannot report, so the first buffer coming out of
+ * either queue is not taken for a page we hold.
+ */
+ state->hdr.idx = UINT_MAX;
+ state->payld.idx = UINT_MAX;
+
+ mpnic_fill_qt_bdqs(qt);
+ }
+ }
+}
+
+void mpnic_napi_disable(struct mpnic_net *mpn)
+{
+ int i;
+
+ for (i = 0; i < mpn->num_napi; i++) {
+ napi_disable_locked(&mpn->napi[i]->napi);
+
+ mpnic_nv_irq_disable(mpn->napi[i]);
+ }
+}
+
+void mpnic_napi_enable(struct mpnic_net *mpn)
+{
+ int i;
+
+ for (i = 0; i < mpn->num_napi; i++)
+ napi_enable_locked(&mpn->napi[i]->napi);
+
+ /* Force the first interrupt on each vector to guarantee that any
+ * completions posted during bringup are processed. Use the TRIGGER
+ * pulse rather than the level triggered global interrupt set, which
+ * can jam the mask/pending state machine if it collides with a
+ * concurrent unmask.
+ */
+ for (i = 0; i < mpn->num_napi; i++) {
+ struct mpnic_napi_vector *nv = mpn->napi[i];
+
+ mpnic_wr64(mpn->mpd, MPNIC_TIM_CTL1(nv->qt[0].cmpl.q_idx),
+ MPNIC_TIM_PARAM_CFG_PRESERVE_MASK |
+ MPNIC_TIM_CTL1_MASK_EN | MPNIC_TIM_CTL1_TRIGGER);
+ }
+
+ mpnic_wrfl(mpn->mpd);
+}
diff --git a/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h
new file mode 100644
index 0000000000000..936ad791a3466
--- /dev/null
+++ b/drivers/net/ethernet/meta/mpnic/mpnic_txrx.h
@@ -0,0 +1,147 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Copyright (c) Meta Platforms, Inc. and affiliates. */
+
+#ifndef _MPNIC_TXRX_H_
+#define _MPNIC_TXRX_H_
+
+#include <linux/if_ether.h>
+#include <linux/netdevice.h>
+#include <linux/skbuff.h>
+#include <linux/types.h>
+#include <net/netdev_queues.h>
+#include <net/xdp.h>
+
+#include "mpnic.h"
+
+struct mpnic_net;
+
+/* Space we have to have available in a work queue to take a packet:
+ * 1 descriptor per page
+ * + 1 descriptor for the skb head
+ * + 1 descriptor for the metadata
+ * + 7 descriptors to keep the tail out of the head's cacheline
+ * If we cannot guarantee that we return NETDEV_TX_BUSY.
+ */
+#define MPNIC_MAX_SKB_DESC (MAX_SKB_FRAGS + 9)
+#define MPNIC_TX_DESC_WAKEUP (MPNIC_MAX_SKB_DESC * 2)
+
+#define MPNIC_MAX_NAPI_VECTORS 1024u
+
+/* Number of buffer descriptors the driver posts before ringing the
+ * doorbell. The device consumes whatever the doorbell points at, this is
+ * purely to keep the driver from writing the CSR for every descriptor.
+ */
+#define MPNIC_BDQ_BATCH_SIZE 64u
+
+#define MPNIC_TXQ_SIZE_DEFAULT 1024
+#define MPNIC_HPQ_SIZE_DEFAULT 256
+#define MPNIC_PPQ_SIZE_DEFAULT 256
+#define MPNIC_RCQ_SIZE_DEFAULT 1024
+
+/* Room the device has to leave in front of and behind every header so the
+ * driver can build an skb around it in place. The headroom is padded out
+ * so that consecutive headers in one page start 128 B aligned.
+ */
+#define MPNIC_RX_TROOM \
+ SKB_DATA_ALIGN(sizeof(struct skb_shared_info))
+#define MPNIC_RX_HROOM \
+ (ALIGN(MPNIC_RX_TROOM + XDP_PACKET_HEADROOM, 128) - MPNIC_RX_TROOM)
+
+/* Headers longer than this are split off into the payload queue */
+#define MPNIC_RX_MAX_HDR 1536
+
+/* A page is handed out to many packets, each of which takes one reference.
+ * Rather than a locked increment per packet the driver takes a batch of
+ * references up front and returns whatever is left when the page is done.
+ */
+#define MPNIC_PAGECNT_BIAS_MAX (PAGE_SIZE + 1)
+
+#define MPNIC_MAX_JUMBO_FRAME_SIZE 9742
+
+/* The page a buffer descriptor queue is currently handing out. Records
+ * how many of the references taken on it are still unused.
+ */
+struct mpnic_pg_ctxt {
+ struct page *page;
+ long pagecnt_bias;
+ u32 idx;
+};
+
+struct mpnic_rcq_state {
+ struct xdp_buff pkt;
+ struct mpnic_pg_ctxt hdr;
+ struct mpnic_pg_ctxt payld;
+ bool add_frag_failed;
+};
+
+struct mpnic_ring {
+ union {
+ struct mpnic_rcq_state *state; /* RCQ */
+ struct page **rx_buf; /* BDQ */
+ void **tx_buf; /* TWQ */
+ void *buffer; /* Generic pointer */
+ };
+
+ u32 __iomem *doorbell; /* Pointer to CSR space for ring */
+ __le64 *desc; /* Descriptor ring memory */
+ u16 size_mask; /* Size of ring in descriptors - 1 */
+ u16 q_idx; /* Hardware queue index */
+
+ u32 head, tail; /* Head/Tail of ring */
+
+ union {
+ /* BDQ only */
+ struct page_pool *page_pool;
+
+ /* TWQ only, index of the metadata descriptor of the last
+ * packet placed in the ring without ringing the doorbell,
+ * -1 if the doorbell is in sync with the tail.
+ */
+ s32 deferred_meta;
+ };
+
+ /* Slow path fields follow */
+ dma_addr_t dma; /* Phys addr of descriptor memory */
+ size_t size; /* Size of descriptor ring in memory */
+};
+
+/* The device pairs two work queues with one completion queue. On the Rx
+ * side they are the header and the payload buffer descriptor queues; on
+ * the Tx side only the first one is used for now, the second one becomes
+ * the XDP ring.
+ */
+struct mpnic_q_triad {
+ struct xdp_rxq_info xdp_rxq;
+ struct mpnic_ring sub0, sub1, cmpl;
+};
+
+struct mpnic_napi_vector {
+ struct napi_struct napi;
+ struct device *dev; /* Device for DMA unmapping */
+ struct mpnic_dev *mpd;
+
+ u16 v_idx;
+ u16 txt_count;
+ u16 rxt_count;
+
+ char name[IFNAMSIZ + 11];
+
+ struct mpnic_q_triad qt[];
+};
+
+netdev_tx_t mpnic_xmit_frame(struct sk_buff *skb, struct net_device *dev);
+int mpnic_alloc_napi_vectors(struct mpnic_net *mpn);
+void mpnic_free_napi_vectors(struct mpnic_net *mpn);
+int mpnic_alloc_resources(struct mpnic_net *mpn);
+void mpnic_free_resources(struct mpnic_net *mpn);
+int mpnic_set_netif_queues(struct mpnic_net *mpn);
+void mpnic_reset_netif_queues(struct mpnic_net *mpn);
+void mpnic_napi_enable(struct mpnic_net *mpn);
+void mpnic_napi_disable(struct mpnic_net *mpn);
+void mpnic_enable(struct mpnic_net *mpn);
+void mpnic_disable(struct mpnic_net *mpn);
+void mpnic_wait_all_queues_idle(struct mpnic_dev *mpd);
+void mpnic_flush(struct mpnic_net *mpn);
+void mpnic_fill(struct mpnic_net *mpn);
+
+#endif /* _MPNIC_TXRX_H_ */
|
| 3/2 | 2026/09/25 06:22 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44490 output: 21 thoughts: 564 Thoughts: **Analyzing the MPNIC Patch for Fuzzing Suitability** Okay, let's break this down. First, I'm looking at this patch and the associated instructions regarding `WorthFuzzing`. The patch introduces a new network driver, `CONFIG_MPNIC`, which is the "Meta Platforms Network Interface Controller." The file structure shows a typical PCI network driver layout within the `drivers/net/ethernet/meta/mpnic/` directory. The key question is whether this new driver falls under the "NEGATIVE CRITERIA" for fuzzing. Specifically, does it relate to "Structurally unreachable hardware" – that is, hardware like vendor-specific PCIe switches, SmartNICs, or GPU drivers that require physical ASICs or PCIe cards not emulated in standard QEMU? The instructions give us those negative examples, but we need to assess our new driver. Okay, in `mpnic_pci.c`, I see a `PCI_DEVICE_ID_META_MPNIC` and a `pci_driver mpnic_driver`. The vendor ID is `PCI_VENDOR_ID_META`. This strongly suggests that this is a vendor-specific PCI device. We need to ascertain if this is a hardware target. I will need to check if there is an emulation target for this hardware, and what the vendor ID "META" maps to. I should grep for `PCI_VDEVICE(META)` to see if there is any existing driver configuration. Based on the name it seems extremely likely this driver is a "SmartNIC" or similar device, specifically designed for Meta's own infrastructure. This strongly suggests it might be unreachable for standard fuzzing in a QEMU environment. I need to make sure, but my initial assessment is that it might. |
| 4/2 | 2026/09/25 06:23 | tool | grepper |
0mArgs:{
"Expression": "PCI_VENDOR_ID_META"
}Results: {
"Output": "include/linux/pci_ids.h-2618-\ninclude/linux/pci_ids.h:2619:#define PCI_VENDOR_ID_META\t\t0x1d9b\ninclude/linux/pci_ids.h-2620-\n--\nrust/kernel/pci/id.rs=172= fn fmt(\u0026self, f: \u0026mut fmt::Formatter\u003c'_\u003e) -\u003e fmt::Result {\n--\nrust/kernel/pci/id.rs-543- HYGON = bindings::PCI_VENDOR_ID_HYGON, // 0x1d94\nrust/kernel/pci/id.rs:544: META = bindings::PCI_VENDOR_ID_META, // 0x1d9b\nrust/kernel/pci/id.rs-545- FUNGIBLE = bindings::PCI_VENDOR_ID_FUNGIBLE, // 0x1dad\n"
} |
| 5/2 | 2026/09/25 06:23 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45686 output: 112 thoughts: 1318 Thoughts: **Triage Decision: WorthFuzzing = false** Okay, let's break this down systematically. I see a PCI vendor ID of `0x1d9b` and a device ID of `0x0014`. That's Meta's custom NIC, specifically the MPNIC, like a FBNC but different. The immediate question is always whether this is emulated in a standard environment like QEMU or Google Compute Engine (GCE). And, based on my knowledge, no, QEMU *does not* emulate the Meta Platforms Network Interface Controller. Alright, let's hit the negative criteria checklist. Specifically: *"Structurally unreachable hardware: Vendor-specific PCIe switches, SmartNICs..."* That's a perfect fit! This clearly requires a physical ASIC or PCIe card that's *not* emulated in standard environments. Now, the knee-jerk reaction is to check for a soft device, virtual bus or mock device that could allow us to instantiate the driver, but the code in `mpnic_probe` itself strongly suggests the driver will not instantiate if the physical device isn't present: The code calls `pcim_enable_device` (for a PCI device). It maps I/O regions using `pcim_iomap_region`, allocates IRQs, initializes the device using `mpnic_dev_init` which writes to registers, polls memory with `mpnic_mem_init_poll`, and checks a `mpnic_present` which fails if no physical hardware is present to respond to register reads/writes. Crucially, it only binds to a *specific* PCI ID. Looking at the triage instructions, section 2 says to return `WorthFuzzing=false` if the code falls into the "structurally unreachable hardware" category and section 1 confirms this. Section 3 lists the positive criteria but none of them apply here. This is a vendor-specific PCIe NIC driver designed for Meta's hardware, and that hardware is *not* emulated. The code operates on specific PCI IDs. This is a dead end. I've examined the code under `drivers/net/ethernet/meta/` and `drivers/net/ethernet/meta/mpnic/`. No other parts of the kernel are touched. It's all tightly scoped. The `Kconfig`, `Makefile`, and header files are all in these same locations. Nothing else. Therefore, the verdict is clear: `WorthFuzzing = false`. The hardware dependency makes it structurally unreachable. The `FocusSymbols` and `EnableConfigs` lists will be empty. The reasoning is the concise explanation: vendor-specific PCIe NIC driver that is reliant on physical, non-emulated hardware. Okay, I'm confident in the decision. Time to call `set-results`. |
| 6/2 | 2026/09/25 06:23 | tool | set-results |
0mArgs:{
"EnableConfigs": [],
"FocusSymbols": [],
"Reasoning": "The patch introduces a new network driver (MPNIC) for Meta Platforms Network Interface Controllers. This is a vendor-specific PCIe NIC requiring dedicated physical ASIC hardware (PCI Vendor ID 0x1d9b, Device ID 0x0014) not emulated in standard virtualized environments (QEMU/GCE). As such, the code is structurally unreachable during fuzzing.",
"WorthFuzzing": false
}Results: {
"EnableConfigs": null,
"FocusSymbols": [],
"Reasoning": "The patch introduces a new network driver (MPNIC) for Meta Platforms Network Interface Controllers. This is a vendor-specific PCIe NIC requiring dedicated physical ASIC hardware (PCI Vendor ID 0x1d9b, Device ID 0x0014) not emulated in standard virtualized environments (QEMU/GCE). As such, the code is structurally unreachable during fuzzing.",
"WorthFuzzing": false
} |
| 7/1 | 2026/09/25 06:23 | action | If |
0mArgs:{
"WorthFuzzing": false
}Results: null |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|