| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/09/26 21:06 | flow | patch-triage |
3mArgs:null Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_core_mmap",
"vfio_pci_dma_buf_mmap"
],
"KMSANReasoning": "The patch series refactors DMA-BUF naming (allowing kernel exporters to set debug names) and integrates DMA-BUF backing into the VFIO PCI BAR mmap infrastructure, introducing a new VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE feature and granular locking with a new rwsem (dmabuf_lock).\n\nKMSAN vs KASAN Applicability:\n1. Uninitialized Memory / Info-leak Analysis:\n - All heap allocations for newly introduced structures (struct vfio_pci_dma_buf, phys_vec) use kzalloc_obj, ensuring all fields are zero-initialized.\n - The string handling in dma_buf_set_name operates on NULL-terminated, initialized strings produced by kasprintf() and strndup_user().\n - The new ioctl feature (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE) is an input-only command (VFIO_DEVICE_FEATURE_SET) that uses copy_from_user; no kernel data or padding is returned or copied to user space.\n - Local variables in the fault handler (vfio_pci_mmap_huge_fault, vfio_pci_dma_buf_find_pfn) such as pfn are explicitly initialized prior to being used or checked.\n\n2. Bug Detector Suitability:\n - The potential bugs introduced by these changes are concurrency/locking deadlocks (dmabuf_lock vs memory_lock vs DMA-BUF resv lock), object lifetime/racing teardown issues (DMABUF vs VFIO device vs VMA references), and potential use-after-free or invalid PFN insertion.\n - These issues are comprehensively covered by LOCKDEP and KASAN.\n - There are no uninitialized memory reads or info-leaks, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch modifies reachable code in DMA-BUF and VFIO-PCI. It converts VFIO PCI BAR memory mappings to be backed by DMABUFs, introduces a new fault handler (vfio_pci_mmap_huge_fault) and PFN search logic (vfio_pci_dma_buf_find_pfn), adds a new user-facing device feature for buffer revocation (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE), restructures BAR revocation synchronization with lockdep assertions, and refactors dma_buf_set_name and its ioctl handling. Fuzzing is warranted to detect races, regressions, and assertion failures.",
"WorthFuzzing": true
} |
| 1/1 | 2026/09/26 21:06 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 0b43d95e7d80a0760fdb300edd1dd05f46dba412\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Sat Sep 26 21:06:15 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c\nindex 4c9add51f9ef5..e4ec504e92884 100644\n--- a/drivers/dma-buf/dma-buf.c\n+++ b/drivers/dma-buf/dma-buf.c\n@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)\n \treturn events;\n }\n \n-/**\n- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.\n- * It could support changing the name of the dma-buf if the same\n- * piece of memory is used for multiple purpose between different devices.\n- *\n- * @dmabuf: [in] dmabuf buffer that will be renamed.\n- * @buf: [in] A piece of userspace memory that contains the name of\n- * the dma-buf.\n- *\n- * Returns 0 on success. If the dma-buf buffer is already attached to\n- * devices, return -EBUSY.\n- *\n- */\n-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)\n-{\n-\tchar *name = strndup_user(buf, DMA_BUF_NAME_LEN);\n-\n-\tif (IS_ERR(name))\n-\t\treturn PTR_ERR(name);\n-\n-\tspin_lock(\u0026dmabuf-\u003ename_lock);\n-\tkfree(dmabuf-\u003ename);\n-\tdmabuf-\u003ename = name;\n-\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n-\n-\treturn 0;\n-}\n-\n #if IS_ENABLED(CONFIG_SYNC_FILE)\n static long dma_buf_export_sync_file(struct dma_buf *dmabuf,\n \t\t\t\t void __user *user_data)\n@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,\n \t\treturn ret;\n \n \tcase DMA_BUF_SET_NAME_A:\n-\tcase DMA_BUF_SET_NAME_B:\n-\t\treturn dma_buf_set_name(dmabuf, (const char __user *)arg);\n+\tcase DMA_BUF_SET_NAME_B: {\n+\t\tchar *name = strndup_user((const char __user *)arg,\n+\t\t\t\t\t DMA_BUF_NAME_LEN);\n+\n+\t\tif (IS_ERR(name))\n+\t\t\treturn PTR_ERR(name);\n+\n+\t\tret = dma_buf_set_name(dmabuf, name);\n+\t\tif (ret)\n+\t\t\tkfree(name);\n+\n+\t\treturn ret;\n+\t}\n \n #if IS_ENABLED(CONFIG_SYNC_FILE)\n \tcase DMA_BUF_IOCTL_EXPORT_SYNC_FILE:\n@@ -874,6 +857,41 @@ void dma_buf_put(struct dma_buf *dmabuf)\n }\n EXPORT_SYMBOL_NS_GPL(dma_buf_put, \"DMA_BUF\");\n \n+/**\n+ * dma_buf_set_name - Set a dmabuf's name\n+ * Intended to be used by the exporter to set a name for debug\n+ * purposes. This can also change an existing name if the same piece\n+ * of memory is used for multiple purposes over time.\n+ *\n+ * @dmabuf: [in] dmabuf buffer that will be renamed.\n+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or\n+ * similar. This takes ownership of the allocation\n+ * on success, which will be kfree()d when the\n+ * dmabuf is released or a new name assigned.\n+ *\n+ * Returns 0 on success, -EINVAL if the name is NULL, or -E2BIG if the\n+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.\n+ */\n+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\n+{\n+\tif (!name)\n+\t\treturn -EINVAL;\n+\n+\t/* dmabuffs_dname() won't use the string if the length\n+\t * (including terminator) exceeds DMA_BUF_NAME_LEN:\n+\t */\n+\tif (strlen(name) \u003e= DMA_BUF_NAME_LEN)\n+\t\treturn -E2BIG;\n+\n+\tspin_lock(\u0026dmabuf-\u003ename_lock);\n+\tkfree(dmabuf-\u003ename);\n+\tdmabuf-\u003ename = name;\n+\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n+\n+\treturn 0;\n+}\n+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, \"DMA_BUF\");\n+\n static int dma_buf_wrap_sg_table(struct sg_table **sg_table)\n {\n \tstruct scatterlist *to_sg, *from_sg;\ndiff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig\nindex 296bf01e185ec..c6d6fb09af86c 100644\n--- a/drivers/vfio/pci/Kconfig\n+++ b/drivers/vfio/pci/Kconfig\n@@ -6,6 +6,7 @@ config VFIO_PCI_CORE\n \ttristate\n \tselect VFIO_VIRQFD\n \tselect IRQ_BYPASS_MANAGER\n+\tselect DMA_SHARED_BUFFER\n \n config VFIO_PCI_INTX\n \tdef_bool y if !S390\n@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM\n \t To enable s390x KVM vfio-pci extensions, say Y.\n \n config VFIO_PCI_DMABUF\n-\tdef_bool y if VFIO_PCI_CORE \u0026\u0026 PCI_P2PDMA \u0026\u0026 DMA_SHARED_BUFFER\n+\tdef_bool y if PCI_P2PDMA\n+\tdepends on VFIO_PCI_CORE\n \n source \"drivers/vfio/pci/mlx5/Kconfig\"\n \ndiff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile\nindex 6138f1bf241df..881452ea89be0 100644\n--- a/drivers/vfio/pci/Makefile\n+++ b/drivers/vfio/pci/Makefile\n@@ -1,8 +1,7 @@\n # SPDX-License-Identifier: GPL-2.0-only\n \n-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o\n+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o\n vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o\n-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o\n obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o\n \n vfio-pci-y := vfio_pci.o\ndiff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\nindex 86362ec424a50..14622556355eb 100644\n--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\n+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\n@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)\n \tstruct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);\n \tstruct pci_dev *pdev = to_pci_dev(core_vdev-\u003edev);\n \tstruct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);\n+\tint ret;\n \n \thisi_acc_vdev-\u003evf_id = pci_iov_vf_id(pdev) + 1;\n \thisi_acc_vdev-\u003epf_qm = pf_qm;\n@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)\n \tcore_vdev-\u003emigration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;\n \tcore_vdev-\u003emig_ops = \u0026hisi_acc_vfio_pci_migrn_state_ops;\n \n-\treturn vfio_pci_core_init_dev(core_vdev);\n+\tret = vfio_pci_core_init_dev(core_vdev);\n+\tif (ret)\n+\t\treturn ret;\n+\t/*\n+\t * hisi_acc_vfio_pci_mmap() calls down to\n+\t * vfio_pci_core_mmap(), so BAR mappings are still\n+\t * DMABUF-backed. They don't require a zap on revoke, so opt\n+\t * out:\n+\t */\n+\thisi_acc_vdev-\u003ecore_device.zap_bars_on_revoke = false;\n+\n+\treturn 0;\n }\n \n static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {\ndiff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c\nindex 9914f3ac69aef..cef337f4e8f2e 100644\n--- a/drivers/vfio/pci/vfio_pci_config.c\n+++ b/drivers/vfio/pci/vfio_pci_config.c\n@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\tvirt_mem = !!(le16_to_cpu(*virt_cmd) \u0026 PCI_COMMAND_MEMORY);\n \t\tnew_mem = !!(new_cmd \u0026 PCI_COMMAND_MEMORY);\n \n-\t\tif (!new_mem) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t\t} else {\n+\t\tif (!new_mem)\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n+\t\telse\n \t\t\tdown_write(\u0026vdev-\u003ememory_lock);\n-\t\t}\n \n \t\t/*\n \t\t * If the user is writing mem/io enable (new_mem/io) and we\n@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t*virt_cmd |= cpu_to_le16(new_cmd \u0026 mask);\n \n \t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n \t}\n \n@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)\n static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,\n \t\t\t\t\t pci_power_t state)\n {\n-\tif (state \u003e= PCI_D3hot) {\n-\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t} else {\n+\tif (state \u003e= PCI_D3hot)\n+\t\tvfio_pci_lock_revoke_bars(vdev);\n+\telse\n \t\tdown_write(\u0026vdev-\u003ememory_lock);\n-\t}\n \n \tvfio_pci_set_power_state(vdev, state);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t\t\t\t\t \u0026cap);\n \n \t\tif (!ret \u0026\u0026 (cap \u0026 PCI_EXP_DEVCAP_FLR)) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n \t\t\tpci_try_reset_function(vdev-\u003epdev);\n \t\t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\t\tup_write(\u0026vdev-\u003ememory_lock);\n \t\t}\n \t}\n@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t\t\t\t\t\u0026cap);\n \n \t\tif (!ret \u0026\u0026 (cap \u0026 PCI_AF_CAP_FLR) \u0026\u0026 (cap \u0026 PCI_AF_CAP_TP)) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n \t\t\tpci_try_reset_function(vdev-\u003epdev);\n \t\t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\t\tup_write(\u0026vdev-\u003ememory_lock);\n \t\t}\n \t}\ndiff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c\nindex 6757054e9d875..8860185cff49f 100644\n--- a/drivers/vfio/pci/vfio_pci_core.c\n+++ b/drivers/vfio/pci/vfio_pci_core.c\n@@ -13,6 +13,8 @@\n #include \u003clinux/aperture.h\u003e\n #include \u003clinux/debugfs.h\u003e\n #include \u003clinux/device.h\u003e\n+#include \u003clinux/dma-buf.h\u003e\n+#include \u003clinux/dma-resv.h\u003e\n #include \u003clinux/eventfd.h\u003e\n #include \u003clinux/file.h\u003e\n #include \u003clinux/interrupt.h\u003e\n@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,\n \t * The vdev power related flags are protected with 'memory_lock'\n \t * semaphore.\n \t */\n-\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\tvfio_pci_dma_buf_move(vdev, true);\n+\tvfio_pci_lock_revoke_bars(vdev);\n \n \tif (vdev-\u003epm_runtime_engaged) {\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)\n \tdown_write(\u0026vdev-\u003ememory_lock);\n \t__vfio_pci_runtime_pm_exit(vdev);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)\n \t */\n \tdown_write(\u0026vdev-\u003ememory_lock);\n \tif (vdev-\u003epm_wake_eventfd_ctx) {\n-\t\teventfd_signal(vdev-\u003epm_wake_eventfd_ctx);\n+\t\tstruct eventfd_ctx *ctx = vdev-\u003epm_wake_eventfd_ctx;\n+\n+\t\tvdev-\u003epm_wake_eventfd_ctx = NULL;\n \t\t__vfio_pci_runtime_pm_exit(vdev);\n+\t\tif (__vfio_pci_memory_enabled(vdev))\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n+\t\teventfd_signal(ctx);\n+\t\teventfd_ctx_put(ctx);\n \t}\n \tup_write(\u0026vdev-\u003ememory_lock);\n \n@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)\n \t\tvdev-\u003ehas_vga = true;\n \n \tvfio_pci_core_map_bars(vdev);\n+\tvdev-\u003ebars_revoked = false;\n \n \treturn 0;\n \n@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,\n \treturn ret;\n }\n \n+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);\n+\n static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \t\t\t\tvoid __user *arg)\n {\n@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \tif (!vdev-\u003ereset_works)\n \t\treturn -EINVAL;\n \n-\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n+\tdown_write(\u0026vdev-\u003ememory_lock);\n \n \t/*\n \t * This function can be invoked while the power state is non-D0. If\n@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \t * have NoSoftRst-, the reset function can cause the PCI config space\n \t * reset without restoring the original state (saved locally in\n \t * 'vdev-\u003epm_save').\n+\t *\n+\t * The zap is done after making the device accessible in D0,\n+\t * because a DMABUF importer could access the device as part\n+\t * of its revocation cleanup.\n \t */\n \tvfio_pci_set_power_state(vdev, PCI_D0);\n \n-\tvfio_pci_dma_buf_move(vdev, true);\n+\tvfio_pci_revoke_bars(vdev);\n+\n \tret = pci_try_reset_function(vdev-\u003epdev);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n \n \treturn ret;\n@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,\n \t\treturn vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);\n \tcase VFIO_DEVICE_FEATURE_ZPCI_ERROR:\n \t\treturn vfio_pci_zdev_feature_err(device, flags, arg, argsz);\n+\tcase VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:\n+\t\treturn vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);\n \tdefault:\n \t\treturn -ENOTTY;\n \t}\n@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu\n }\n EXPORT_SYMBOL_GPL(vfio_pci_core_write);\n \n-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)\n+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)\n {\n-\tstruct vfio_device *core_vdev = \u0026vdev-\u003evdev;\n-\tloff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);\n-\tloff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);\n-\tloff_t len = end - start;\n+\tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n+\tvfio_pci_dma_buf_move(vdev, true);\n \n-\tunmap_mapping_range(core_vdev-\u003einode-\u003ei_mapping, start, len, true);\n+\t/*\n+\t * If a driver could possibly create BAR mappings in the\n+\t * vdev's address_space, do an additional zap on revoke. See\n+\t * vfio_pci_core_init_dev().\n+\t */\n+\tif (vdev-\u003ezap_bars_on_revoke) {\n+\t\tstruct vfio_device *core_vdev = \u0026vdev-\u003evdev;\n+\t\tloff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);\n+\t\tloff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);\n+\t\tloff_t len = end - start;\n+\n+\t\tunmap_mapping_range(core_vdev-\u003einode-\u003ei_mapping,\n+\t\t\t\t start, len, true);\n+\t}\n }\n \n-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)\n+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)\n {\n \tdown_write(\u0026vdev-\u003ememory_lock);\n-\tvfio_pci_zap_bars(vdev);\n+\tvfio_pci_revoke_bars(vdev);\n+}\n+\n+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)\n+{\n+\tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n+\tvfio_pci_dma_buf_move(vdev, false);\n }\n \n u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)\n@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n-static unsigned long vma_to_pfn(struct vm_area_struct *vma)\n-{\n-\tstruct vfio_pci_core_device *vdev = vma-\u003evm_private_data;\n-\tint index = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n-\tu64 pgoff;\n-\n-\tpgoff = vma-\u003evm_pgoff \u0026\n-\t\t((1U \u003c\u003c (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);\n-\n-\treturn (pci_resource_start(vdev-\u003epdev, index) \u003e\u003e PAGE_SHIFT) + pgoff;\n-}\n-\n vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\n \t\t\t\t struct vm_fault *vmf,\n \t\t\t\t unsigned long pfn,\n@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n \t\t\t\t\t unsigned int order)\n {\n \tstruct vm_area_struct *vma = vmf-\u003evma;\n-\tstruct vfio_pci_core_device *vdev = vma-\u003evm_private_data;\n-\tunsigned long addr = vmf-\u003eaddress \u0026 ~((PAGE_SIZE \u003c\u003c order) - 1);\n-\tunsigned long pgoff = linear_page_delta(vma, addr);\n-\tunsigned long pfn = vma_to_pfn(vma) + pgoff;\n-\tvm_fault_t ret = VM_FAULT_FALLBACK;\n-\n-\tif (is_aligned_for_order(vma, addr, pfn, order)) {\n-\t\tscoped_guard(rwsem_read, \u0026vdev-\u003ememory_lock)\n-\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);\n+\tstruct vfio_pci_dma_buf *priv = vma-\u003evm_private_data;\n+\tstruct vfio_pci_core_device *vdev;\n+\tunsigned long pfn = 0;\n+\tvm_fault_t ret = VM_FAULT_SIGBUS;\n+\n+\t/*\n+\t * The only thing this can rely on is that the DMABUF relating\n+\t * to the VMA's vm_file exists (priv).\n+\t *\n+\t * A DMABUF for a VFIO device fd mmap() holds a reference to\n+\t * the original VFIO device fd, but an explicitly-exported\n+\t * DMABUF does not. The original fd might have closed,\n+\t * meaning this fault can race with\n+\t * vfio_pci_dma_buf_cleanup(), meaning the buffer could have\n+\t * been revoked (in which case priv-\u003evdev might be NULL), and\n+\t * the VFIO device registration might have been dropped.\n+\t *\n+\t * With the goal of taking vdev locks in a world where vdev\n+\t * might not still exist:\n+\t *\n+\t * 1. Take the resv lock on the DMABUF:\n+\t * - If racing cleanup got in first, the buffer is revoked;\n+\t * stop/exit if so.\n+\t * - If we got in first, the buffer is not revoked so vdev is\n+\t * non-NULL, accessible, and cleanup _has not yet put the\n+\t * VFIO device registration_. So, the device refcount must\n+\t * be \u003e0.\n+\t *\n+\t * 2. Take vfio_device registration (refcount guaranteed \u003e0\n+\t * hereafter).\n+\t *\n+\t * 3. Unlock the DMABUF's resv lock:\n+\t * - A racing cleanup can now complete.\n+\t * - But, the device refcount \u003e0, meaning the vfio_device\n+\t * (and vfio_pcie_core device vdev) have not yet been\n+\t * freed. vdev is accessible, even if the DMABUF has been\n+\t * revoked or cleanup has happened, because\n+\t * vfio_unregister_group_dev() can't complete.\n+\t *\n+\t * 4. Take the vdev-\u003ememory_lock then vdev-\u003edmabuf_lock:\n+\t * - Either the DMABUF is usable, or has been cleaned up.\n+\t * - It's not necessary to also take the resv lock, because\n+\t * the status/vdev can't change while dmabuf_lock is held.\n+\t * - Test the DMABUF revocation status again: if it was\n+\t * revoked between 1 and 4, return a SIGBUS. Otherwise,\n+\t * return a PFN.\n+\t *\n+\t * 5. Unlock, done.\n+\t */\n+\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK) {\n+\t\tpr_debug_ratelimited(\"%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\\n\",\n+\t\t\t\t __func__, vmf-\u003eaddress, vma-\u003evm_pgoff);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\treturn VM_FAULT_SIGBUS;\n+\t}\n+\n+\t/* If the buffer isn't revoked, vdev is valid */\n+\tvdev = priv-\u003evdev;\n+\n+\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev)) {\n+\t\t/*\n+\t\t * If vdev != NULL (above), the registration should\n+\t\t * already be \u003e0 and so this try_get should never\n+\t\t * fail.\n+\t\t */\n+\t\tdev_warn_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\t \"%s: Unexpected registration failure\\n\",\n+\t\t\t\t __func__);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\treturn VM_FAULT_SIGBUS;\n+\t}\n+\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\n+\t/* memory_lock for vfio_pci_vmf_insert_pfn() */\n+\tdown_read(\u0026vdev-\u003ememory_lock);\n+\t/* Re-test revocation status under dmabuf_lock */\n+\tdown_read(\u0026vdev-\u003edmabuf_lock);\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_OK) {\n+\t\tint pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,\n+\t\t\t\t\t\t vmf-\u003eaddress,\n+\t\t\t\t\t\t order, \u0026pfn);\n+\n+\t\tif (pres == 0)\n+\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf,\n+\t\t\t\t\t\t pfn, order);\n+\t\telse if (pres == -ERANGE)\n+\t\t\tret = VM_FAULT_FALLBACK;\n \t}\n+\tup_read(\u0026vdev-\u003edmabuf_lock);\n+\tup_read(\u0026vdev-\u003ememory_lock);\n \n \tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n-\t\t\t \"%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\\n\",\n-\t\t\t __func__, order,\n-\t\t\t vma-\u003evm_pgoff \u003e\u003e\n-\t\t\t\t(VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),\n-\t\t\t pgoff, (unsigned int)ret);\n+\t\t\t \"%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\\n\",\n+\t\t\t __func__, order, pfn, vmf-\u003eaddress,\n+\t\t\t vma-\u003evm_pgoff, (unsigned int)ret);\n \n+\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n \treturn ret;\n }\n \n@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {\n #endif\n };\n \n+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)\n+{\n+\tvma-\u003evm_ops = \u0026vfio_pci_mmap_ops;\n+}\n+\n int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)\n {\n \tstruct vfio_pci_core_device *vdev =\n@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma\n \tunsigned int index;\n \tu64 phys_len, req_len, pgoff, req_start;\n \tvoid __iomem *bar_io;\n+\tint ret;\n \n \tindex = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n \n@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma\n \tif (IS_ERR(bar_io))\n \t\treturn PTR_ERR(bar_io);\n \n-\tvma-\u003evm_private_data = vdev;\n+\tret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,\n+\t\t\t\t\t pci_resource_start(pdev, index),\n+\t\t\t\t\t req_len, index);\n+\tif (ret)\n+\t\treturn ret;\n+\n \tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n \tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n \n@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)\n \t\treturn ret;\n \tINIT_LIST_HEAD(\u0026vdev-\u003edmabufs);\n \tinit_rwsem(\u0026vdev-\u003ememory_lock);\n+\tinit_rwsem(\u0026vdev-\u003edmabuf_lock);\n \txa_init(\u0026vdev-\u003ectx);\n \n+\t/*\n+\t * If a driver overrides .mmap, it has to be assumed that it\n+\t * might not use the DMABUF-backed core mmap; this flag\n+\t * enables a zap at revoke time. A driver can opt out by\n+\t * clearing this flag at init, if their .mmap override calls\n+\t * down to vfio_pci_core_mmap().\n+\t */\n+\tif (vdev-\u003evdev.ops-\u003emmap != vfio_pci_core_mmap)\n+\t\tvdev-\u003ezap_bars_on_revoke = true;\n+\n \treturn 0;\n }\n EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);\n@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \t\t}\n \n \t\t/*\n-\t\t * Take the memory write lock for each device and zap BAR\n-\t\t * mappings to prevent the user accessing the device while in\n-\t\t * reset. Locking multiple devices is prone to deadlock,\n+\t\t * Take the memory write lock for each device and\n+\t\t * zap/revoke BAR mappings to prevent the user (or\n+\t\t * peers) accessing the device while in reset.\n+\t\t * Locking multiple devices is prone to deadlock,\n \t\t * runaway and unwind if we hit contention.\n \t\t */\n \t\tif (!down_write_trylock(\u0026vdev-\u003ememory_lock)) {\n@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \t\t\tbreak;\n \t\t}\n \n-\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t\tvfio_pci_zap_bars(vdev);\n+\t\tvfio_pci_revoke_bars(vdev);\n \t}\n \n \tif (!list_entry_is_head(vdev,\n@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \tlist_for_each_entry_from_reverse(vdev, \u0026dev_set-\u003edevice_list,\n \t\t\t\t\t vdev.dev_set_list) {\n \t\tif (vdev-\u003evdev.open_count \u0026\u0026 __vfio_pci_memory_enabled(vdev))\n-\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n \t}\n \ndiff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c\nindex c16f460c01d68..b57bfaefd9fae 100644\n--- a/drivers/vfio/pci/vfio_pci_dmabuf.c\n+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c\n@@ -3,25 +3,14 @@\n */\n #include \u003clinux/dma-buf-mapping.h\u003e\n #include \u003clinux/pci-p2pdma.h\u003e\n+#include \u003clinux/dma-buf.h\u003e\n #include \u003clinux/dma-resv.h\u003e\n \n #include \"vfio_pci_priv.h\"\n \n MODULE_IMPORT_NS(\"DMA_BUF\");\n \n-struct vfio_pci_dma_buf {\n-\tstruct dma_buf *dmabuf;\n-\tstruct vfio_pci_core_device *vdev;\n-\tstruct list_head dmabufs_elm;\n-\tsize_t size;\n-\tstruct phys_vec *phys_vec;\n-\tstruct p2pdma_provider *provider;\n-\tu32 nr_ranges;\n-\tstruct kref kref;\n-\tstruct completion comp;\n-\tu8 revoked : 1;\n-};\n-\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \t\t\t\t struct dma_buf_attachment *attachment)\n {\n@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \tif (!attachment-\u003epeer2peer)\n \t\treturn -EOPNOTSUPP;\n \n-\tif (priv-\u003erevoked)\n+\tif (READ_ONCE(priv-\u003estatus) != VFIO_PCI_DMABUF_OK)\n \t\treturn -ENODEV;\n \n \tif (!dma_buf_attach_revocable(attachment))\n@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \treturn 0;\n }\n \n+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)\n+{\n+\tstruct vfio_pci_dma_buf *priv = dmabuf-\u003epriv;\n+\n+\t/*\n+\t * dma_buf_mmap_internal() has asserted that the VMA is\n+\t * contained within the DMABUF size before calling this.\n+\t *\n+\t * Also, if we observe that the buffer is revoked now then\n+\t * refuse the mmap(). This is a belt-and-braces early failure\n+\t * to ease debugging a revoked buffer being used. Userspace\n+\t * might also race an mmap() against an explicit revocation,\n+\t * or an action causing a revoke; race scenarios are still\n+\t * safe because the fault handler ultimately prevents access\n+\t * to a revoked buffer if it isn't caught here.\n+\t */\n+\tif (READ_ONCE(priv-\u003estatus) != VFIO_PCI_DMABUF_OK)\n+\t\treturn -ENODEV;\n+\t/*\n+\t * Make clear that anything with an offset adjustment is\n+\t * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()\n+\t * maths would underflow; this doesn't happen through the\n+\t * regular DMABUF export path used with this mmap(). A DMABUF\n+\t * implicitly created for BAR mmap could have adjust \u003e 0, but\n+\t * these can't currently be re-opened and mmap()ed again.\n+\t * Catch here in case that assumption ever changes.\n+\t */\n+\tif (priv-\u003evma_pgoff_adjust)\n+\t\treturn -EINVAL;\n+\tif ((vma-\u003evm_flags \u0026 VM_SHARED) == 0)\n+\t\treturn -EINVAL;\n+\n+\tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n+\tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n+\n+\t/* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */\n+\tvm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |\n+\t\t VM_DONTEXPAND | VM_DONTDUMP);\n+\tvma-\u003evm_private_data = priv;\n+\tvfio_pci_set_vma_ops(vma);\n+\n+\treturn 0;\n+}\n+#else\n+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n+\t\t\t\t struct dma_buf_attachment *attachment)\n+{\n+\t/*\n+\t * Explicit export can't occur without the DMABUF feature, but\n+\t * DMABUFs are implicitly created for BAR mappings. An\n+\t * .attach that fails prevents dma_buf_attach().\n+\t */\n+\treturn -EOPNOTSUPP;\n+}\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\n+\n static void vfio_pci_dma_buf_done(struct kref *kref)\n {\n \tstruct vfio_pci_dma_buf *priv =\n@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,\n \n \tdma_resv_assert_held(priv-\u003edmabuf-\u003eresv);\n \n-\tif (priv-\u003erevoked)\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK)\n \t\treturn ERR_PTR(-ENODEV);\n \n \tret = dma_buf_phys_vec_to_sgt(attachment, priv-\u003eprovider,\n@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)\n \t * The refcount prevents both.\n \t */\n \tif (priv-\u003evdev) {\n-\t\tdown_write(\u0026priv-\u003evdev-\u003ememory_lock);\n+\t\tdown_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n \t\tlist_del_init(\u0026priv-\u003edmabufs_elm);\n-\t\tup_write(\u0026priv-\u003evdev-\u003ememory_lock);\n+\t\tup_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n \t\tvfio_device_put_registration(\u0026priv-\u003evdev-\u003evdev);\n \t}\n+\tif (priv-\u003evfile)\n+\t\tfput(priv-\u003evfile);\n \tkfree(priv-\u003ephys_vec);\n \tkfree(priv);\n }\n \n static const struct dma_buf_ops vfio_pci_dmabuf_ops = {\n \t.attach = vfio_pci_dma_buf_attach,\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n+\t.mmap = vfio_pci_dma_buf_mmap,\n+#endif\n \t.map_dma_buf = vfio_pci_dma_buf_map,\n \t.unmap_dma_buf = vfio_pci_dma_buf_unmap,\n \t.release = vfio_pci_dma_buf_release,\n };\n \n+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,\n+\t\t\t struct vfio_pci_dma_buf *priv,\n+\t\t\t struct vm_area_struct *vma,\n+\t\t\t unsigned long fault_addr,\n+\t\t\t unsigned int order,\n+\t\t\t unsigned long *out_pfn)\n+{\n+\t/*\n+\t * Given a VMA (start, end, pgoffs) and a fault address,\n+\t * search the corresponding DMABUF's phys_vec[] to find the\n+\t * range representing the address's offset into the VMA, and\n+\t * its PFN. vdev must be the device that the DMABUF priv was\n+\t * exported from; vdev-\u003edmabuf_lock must be held, and priv\n+\t * must not be revoked.\n+\t *\n+\t * The phys_vec[] ranges represent contiguous spans of VAs\n+\t * upwards from the buffer offset 0; the actual PFNs might be\n+\t * in any order, overlap/alias, etc. Calculate an offset of\n+\t * the desired page given VMA start/pgoff and address, then\n+\t * search upwards from 0 to find which span contains it.\n+\t *\n+\t * On success, a valid PFN for a page sized by 'order' is\n+\t * returned into out_pfn.\n+\t *\n+\t * Failure occurs if:\n+\t * - A hugepage would cross the edge of the VMA,\n+\t * - A hugepage isn't entirely contained within a range\n+\t * (including where it straddles the boundary between\n+\t * ranges),\n+\t * - We find a range, but the final PFN isn't aligned to the\n+\t * requested order.\n+\t *\n+\t * Upon failure, -ERANGE is returned and the caller is\n+\t * expected to try again with a smaller order, which will\n+\t * eventually succeed.\n+\t *\n+\t * It's suboptimal if DMABUFs are created with neighbouring\n+\t * ranges that are physically contiguous, since hugepages\n+\t * can't straddle range boundaries. (The construction of the\n+\t * ranges should merge them in this case.)\n+\t *\n+\t * Finally, vma_pgoff_adjust is used with a DMABUF created for\n+\t * a VFIO BAR mmap: a BAR mapped with vm_pgoff \u003e 0 creates a\n+\t * DMABUF such that byte 0 of the VMA corresponds to byte 0 of\n+\t * the DMABUF and byte 'vm_pgoff \u003c\u003c PAGE_SHIFT' into the BAR.\n+\t * To avoid double-offsetting in this scenario, subtracting\n+\t * vma_pgoff_adjust from this (non-zero) vm_pgoff generates\n+\t * the effective offset. This also removes the VFIO region\n+\t * index encoded in vm_pgoff for VFIO BAR mmaps.\n+\t */\n+\n+\tconst unsigned long pagesize = PAGE_SIZE \u003c\u003c order;\n+\tunsigned long vma_off = (vma-\u003evm_pgoff - priv-\u003evma_pgoff_adjust) \u003c\u003c\n+\t\t\t\t PAGE_SHIFT;\n+\tunsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);\n+\tunsigned long rounded_page_end = rounded_page_addr + pagesize;\n+\tunsigned long fault_offset;\n+\tunsigned long fault_offset_end;\n+\tunsigned long range_start_offset = 0;\n+\tunsigned int i;\n+\tint ret;\n+\n+\tif (unlikely(!vdev))\n+\t\treturn -ENODEV;\n+\n+\t/* This prevents the dmabuf revocation state from changing under us */\n+\tlockdep_assert_held(\u0026vdev-\u003edmabuf_lock);\n+\n+\tif (unlikely(priv-\u003evdev != vdev || priv-\u003estatus != VFIO_PCI_DMABUF_OK))\n+\t\treturn -ENODEV;\n+\n+\tif (rounded_page_addr \u003c vma-\u003evm_start || rounded_page_end \u003e vma-\u003evm_end) {\n+\t\tif (order \u003e 0)\n+\t\t\treturn -ERANGE;\n+\n+\t\t/* A fault address outside of the VMA is absurd. */\n+\t\tdev_warn_ratelimited(\n+\t\t\t\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\"Fault addr 0x%lx outside VMA 0x%lx-0x%lx\\n\",\n+\t\t\tfault_addr, vma-\u003evm_start, vma-\u003evm_end);\n+\t\treturn -EFAULT;\n+\t}\n+\n+\t/*\n+\t * fault_offset[_end] is the span within the DMABUF\n+\t * corresponding to the faulting page:\n+\t */\n+\tif (unlikely(check_add_overflow(rounded_page_addr - vma-\u003evm_start,\n+\t\t\t\t\tvma_off, \u0026fault_offset) ||\n+\t\t check_add_overflow(fault_offset, pagesize,\n+\t\t\t\t\t\u0026fault_offset_end)))\n+\t\treturn -EFAULT;\n+\n+\t/*\n+\t * Iterate over ranges in the buffer, summing their lengths:\n+\t * range_start_offset represents the current range's starting\n+\t * offset in the buffer (from 0 upwards).\n+\t *\n+\t * A failure for order == 0 is unexpected, and triggers a\n+\t * fault/warn.\n+\t */\n+\tret = (order == 0) ? -EFAULT : -ERANGE;\n+\n+\tfor (i = 0; i \u003c priv-\u003enr_ranges; i++) {\n+\t\tsize_t range_len = priv-\u003ephys_vec[i].len;\n+\n+\t\t/* Early exit if range starts after the page end */\n+\t\tif (fault_offset_end \u003c= range_start_offset)\n+\t\t\tbreak;\n+\n+\t\tif (fault_offset \u003e= range_start_offset \u0026\u0026\n+\t\t fault_offset_end \u003c= range_start_offset + range_len) {\n+\t\t\t/*\n+\t\t\t * The faulting page is wholly contained\n+\t\t\t * within the span represented by this range,\n+\t\t\t * so validate PFN alignment for the order.\n+\t\t\t * The if() condition ensures the pfn\n+\t\t\t * arithmetic won't overflow.\n+\t\t\t */\n+\t\t\tunsigned long pfn =\n+\t\t\t\t((fault_offset - range_start_offset) +\n+\t\t\t\t priv-\u003ephys_vec[i].paddr) \u003e\u003e PAGE_SHIFT;\n+\n+\t\t\tif (IS_ALIGNED(pfn, 1 \u003c\u003c order)) {\n+\t\t\t\t*out_pfn = pfn;\n+\t\t\t\tret = 0;\n+\t\t\t}\n+\t\t\t/*\n+\t\t\t * Else order \u003e 0; ERANGE retries with smaller\n+\t\t\t * order\n+\t\t\t */\n+\t\t\tbreak;\n+\t\t}\n+\t\trange_start_offset += range_len;\n+\t}\n+\n+\tif (order == 0 \u0026\u0026 ret != 0)\n+\t\t/*\n+\t\t * The address fell outside of the span represented by\n+\t\t * the (concatenated) ranges. As setup of a mapping\n+\t\t * ensures that the VMA is \u003c= the total size of the\n+\t\t * ranges this should never happen. If it does, warn\n+\t\t * and SIGBUS.\n+\t\t */\n+\t\tdev_warn_ratelimited(\n+\t\t\t\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\"No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\\n\",\n+\t\t\tfault_addr, order, vma-\u003evm_start, vma-\u003evm_end,\n+\t\t\tvma-\u003evm_pgoff, priv-\u003enr_ranges, priv-\u003esize);\n+\n+\treturn ret;\n+}\n+\n+/*\n+ * Create a DMABUF corresponding to priv, add it to vdev-\u003edmabufs list\n+ * for tracking (meaning cleanup or revocation will zap it), and take\n+ * a vfio_device registration.\n+ */\n+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vfio_pci_dma_buf *priv, u32 flags)\n+{\n+\tDEFINE_DMA_BUF_EXPORT_INFO(exp_info);\n+\n+\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev))\n+\t\treturn -ENODEV;\n+\n+\texp_info.ops = \u0026vfio_pci_dmabuf_ops;\n+\texp_info.size = priv-\u003esize;\n+\texp_info.flags = flags;\n+\texp_info.priv = priv;\n+\n+\tpriv-\u003edmabuf = dma_buf_export(\u0026exp_info);\n+\tif (IS_ERR(priv-\u003edmabuf)) {\n+\t\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n+\t\treturn PTR_ERR(priv-\u003edmabuf);\n+\t}\n+\n+\tkref_init(\u0026priv-\u003ekref);\n+\tinit_completion(\u0026priv-\u003ecomp);\n+\n+\t/* dma_buf_put() now frees priv */\n+\tINIT_LIST_HEAD(\u0026priv-\u003edmabufs_elm);\n+\n+\t/*\n+\t * dmabuf_lock synchronises access (R) or updates (W) to the\n+\t * vdev-\u003edmabufs list and to bars_revoked (see below). The\n+\t * revocation state of DMABUF elements in the list is written\n+\t * holding both dmabuf_lock(W) and resv, and tested with\n+\t * either.\n+\t *\n+\t * (memory_lock, if held -\u003e) dmabuf_lock -\u003e resv\n+\t *\n+\t * NOTE: memory_lock is strictly avoided here, to avoid a\n+\t * dependency on memory_lock when mmap_lock is held, when\n+\t * mmap() leads to export. vfio-pci variant drivers are\n+\t * permitted to hold memory_lock across actions that might\n+\t * fault (such as user access); a deadlock could result when\n+\t * that fault path attempts to take mmap_lock (if held by an\n+\t * export waiting for memory_lock).\n+\t *\n+\t * vdev-\u003ebars_revoked tracks the BAR revocation status updated\n+\t * via vfio_pci_dma_buf_move(), so the initial DMABUF state\n+\t * follows the same criteria that later update the DMABUF\n+\t * state (BAR zap, etc.).\n+\t */\n+\tlockdep_assert_not_held(\u0026vdev-\u003ememory_lock);\n+\n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\tpriv-\u003estatus = vdev-\u003ebars_revoked ? VFIO_PCI_DMABUF_REVOKED :\n+\t\tVFIO_PCI_DMABUF_OK;\n+\tlist_add_tail(\u0026priv-\u003edmabufs_elm, \u0026vdev-\u003edmabufs);\n+\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n+\n+\treturn 0;\n+}\n+\n+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vm_area_struct *vma,\n+\t\t\t\t u64 phys_start, u64 req_len,\n+\t\t\t\t unsigned int res_index)\n+{\n+\tstruct vfio_pci_dma_buf *priv;\n+\tunsigned long vma_pgoff = vma-\u003evm_pgoff \u0026 (VFIO_PCI_OFFSET_MASK \u003e\u003e PAGE_SHIFT);\n+\tchar *bufname;\n+\tint ret;\n+\n+\tpriv = kzalloc_obj(*priv);\n+\tif (!priv)\n+\t\treturn -ENOMEM;\n+\n+\tpriv-\u003ephys_vec = kzalloc_obj(*priv-\u003ephys_vec);\n+\tif (!priv-\u003ephys_vec) {\n+\t\tret = -ENOMEM;\n+\t\tgoto err_free_priv;\n+\t}\n+\n+\t/*\n+\t * Debug name: The absolute maximum size of the name\n+\t * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.\n+\t */\n+\tbufname = kasprintf(GFP_KERNEL, \"vfio:%s/%x\",\n+\t\t\t pci_name(vdev-\u003epdev),\n+\t\t\t res_index);\n+\n+\tif (!bufname) {\n+\t\tret = -ENOMEM;\n+\t\tgoto err_free_phys;\n+\t}\n+\n+\t/*\n+\t * The DMABUF begins from the mmap()'s BAR offset, i.e. the\n+\t * start of the VMA corresponds to byte 0 of the DMABUF and\n+\t * byte (vma_pgoff \u003c\u003c PAGE_SHIFT) of the BAR.\n+\t *\n+\t * vfio_pci_dma_buf_find_pfn() reverses this offset using\n+\t * vma_pgoff_adjust, so that ultimately a fault's offset from\n+\t * the start of the _VMA_ has a consistent usage whether the\n+\t * VMA originates from an mmap() of the VFIO device here or a\n+\t * direct DMABUF mmap(). Note vma_pgoff_adjust also includes\n+\t * the encoded VFIO region index, which cancels out the index\n+\t * encoded in vm_pgoff.\n+\t */\n+\tpriv-\u003evdev = vdev;\n+\tpriv-\u003esize = req_len;\n+\tpriv-\u003enr_ranges = 1;\n+\tpriv-\u003evma_pgoff_adjust = vma-\u003evm_pgoff;\n+\n+\t/*\n+\t * The provider can be NULL _iff_ the DMABUF feature isn't\n+\t * supported, because it's only used by DMABUF import and\n+\t * attach is prohibited if the feature isn't present.\n+\t */\n+\tpriv-\u003eprovider = pcim_p2pdma_provider(vdev-\u003epdev, res_index);\n+\tif (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) \u0026\u0026 !priv-\u003eprovider) {\n+\t\tret = -EINVAL;\n+\t\tgoto err_free_name;\n+\t}\n+\n+\tpriv-\u003ephys_vec[0].paddr = phys_start + ((u64)vma_pgoff \u003c\u003c PAGE_SHIFT);\n+\tpriv-\u003ephys_vec[0].len = priv-\u003esize;\n+\n+\tret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);\n+\tif (ret)\n+\t\tgoto err_free_name;\n+\n+\tif (dma_buf_set_name(priv-\u003edmabuf, bufname)) {\n+\t\tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\t \"Failed to set map name '%s'\\n\",\n+\t\t\t\t bufname);\n+\t\tkfree(bufname);\n+\t}\n+\n+\t/*\n+\t * Ownership of the DMABUF file transfers to the VMA so that\n+\t * other users can locate the DMABUF via a VA. Ownership of\n+\t * the original VFIO device file being mmap()ed transfers to\n+\t * priv, and is put when the DMABUF is released. This\n+\t * intentionally does not use get_file()/vma_set_file()\n+\t * because the references are already held, and ownership\n+\t * moves.\n+\t */\n+\tpriv-\u003evfile = vma-\u003evm_file;\n+\tvma-\u003evm_file = priv-\u003edmabuf-\u003efile;\n+\tvma-\u003evm_private_data = priv;\n+\n+\treturn 0;\n+\n+err_free_name:\n+\tkfree(bufname);\n+err_free_phys:\n+\tkfree(priv-\u003ephys_vec);\n+err_free_priv:\n+\tkfree(priv);\n+\treturn ret;\n+}\n+\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n /*\n * This is a temporary \"private interconnect\" between VFIO DMABUF and iommufd.\n * It allows the two co-operating drivers to exchange the physical address of\n@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,\n \t\treturn -EOPNOTSUPP;\n \n \tpriv = attachment-\u003edmabuf-\u003epriv;\n-\tif (priv-\u003erevoked)\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK)\n \t\treturn -ENODEV;\n \n \t/* More than one range to iommufd will require proper DMABUF support */\n@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n {\n \tstruct vfio_device_feature_dma_buf get_dma_buf = {};\n \tstruct vfio_region_dma_range *dma_ranges;\n-\tDEFINE_DMA_BUF_EXPORT_INFO(exp_info);\n \tstruct vfio_pci_dma_buf *priv;\n \tsize_t length;\n \tint ret;\n@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \tkfree(dma_ranges);\n \tdma_ranges = NULL;\n \n-\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev)) {\n-\t\tret = -ENODEV;\n+\tret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);\n+\tif (ret)\n \t\tgoto err_free_phys;\n-\t}\n-\n-\texp_info.ops = \u0026vfio_pci_dmabuf_ops;\n-\texp_info.size = priv-\u003esize;\n-\texp_info.flags = get_dma_buf.open_flags;\n-\texp_info.priv = priv;\n-\n-\tpriv-\u003edmabuf = dma_buf_export(\u0026exp_info);\n-\tif (IS_ERR(priv-\u003edmabuf)) {\n-\t\tret = PTR_ERR(priv-\u003edmabuf);\n-\t\tgoto err_dev_put;\n-\t}\n-\n-\tkref_init(\u0026priv-\u003ekref);\n-\tinit_completion(\u0026priv-\u003ecomp);\n-\n-\t/* dma_buf_put() now frees priv */\n-\tINIT_LIST_HEAD(\u0026priv-\u003edmabufs_elm);\n-\tdown_write(\u0026vdev-\u003ememory_lock);\n-\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\tpriv-\u003erevoked = !__vfio_pci_memory_enabled(vdev);\n-\tlist_add_tail(\u0026priv-\u003edmabufs_elm, \u0026vdev-\u003edmabufs);\n-\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\tup_write(\u0026vdev-\u003ememory_lock);\n-\n \t/*\n \t * dma_buf_fd() consumes the reference, when the file closes the dmabuf\n \t * will be released.\n@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \n \treturn ret;\n \n-err_dev_put:\n-\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n err_free_phys:\n \tkfree(priv-\u003ephys_vec);\n err_free_priv:\n@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \tkfree(dma_ranges);\n \treturn ret;\n }\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\n+\n+/*\n+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives\n+ * the guarantee that all future map/attach attempts will fail no\n+ * matter what, whereas REVOKED can transition back to OK.\n+ */\n+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,\n+\t\t\t\t\tenum vfio_pci_dma_buf_status new_status)\n+{\n+\tbool was_revoked;\n+\n+\t/*\n+\t * Changes to the DMABUF's revocation status are synchronised\n+\t * using dmabuf_lock:\n+\t */\n+\tlockdep_assert_held_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n+\n+\t/* If DEAD, state can no longer change */\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_DEAD ||\n+\t priv-\u003estatus == new_status)\n+\t\treturn;\n+\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\twas_revoked = (priv-\u003estatus == VFIO_PCI_DMABUF_REVOKED);\n+\n+\tif (new_status != VFIO_PCI_DMABUF_OK) {\n+\t\tpriv-\u003estatus = new_status;\n+\n+\t\tif (was_revoked) {\n+\t\t\t/*\n+\t\t\t * A REVOKED buffer is being marked DEAD.\n+\t\t\t * invalidate_mappings/unmap wait happened\n+\t\t\t * when it became REVOKED, don't wait again.\n+\t\t\t */\n+\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\t\treturn;\n+\t\t}\n+\t\tdma_buf_invalidate_mappings(priv-\u003edmabuf);\n+\t\tdma_resv_wait_timeout(priv-\u003edmabuf-\u003eresv,\n+\t\t\t\t DMA_RESV_USAGE_BOOKKEEP, false,\n+\t\t\t\t MAX_SCHEDULE_TIMEOUT);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\tkref_put(\u0026priv-\u003ekref, vfio_pci_dma_buf_done);\n+\t\twait_for_completion(\u0026priv-\u003ecomp);\n+\t\tunmap_mapping_range(priv-\u003edmabuf-\u003efile-\u003ef_mapping,\n+\t\t\t\t 0, 0, true);\n+\t\t/*\n+\t\t * Re-arm the registered kref reference and the\n+\t\t * completion so the post-revoke state matches the\n+\t\t * post-creation state. An un-revoke followed by a\n+\t\t * new mapping needs the kref to be non-zero before\n+\t\t * kref_get(), and vfio_pci_dma_buf_cleanup()\n+\t\t * delegates its drain back through this revoke\n+\t\t * path on a possibly-already-revoked dma-buf.\n+\t\t */\n+\t\tkref_init(\u0026priv-\u003ekref);\n+\t\treinit_completion(\u0026priv-\u003ecomp);\n+\t} else {\n+\t\tpriv-\u003estatus = VFIO_PCI_DMABUF_OK;\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t}\n+}\n \n void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)\n {\n@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)\n \n \tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n \n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tvdev-\u003ebars_revoked = revoked;\n \tlist_for_each_entry_safe(priv, tmp, \u0026vdev-\u003edmabufs, dmabufs_elm) {\n \t\tif (!get_file_active(\u0026priv-\u003edmabuf-\u003efile))\n \t\t\tcontinue;\n-\n-\t\tif (priv-\u003erevoked != revoked) {\n-\t\t\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\t\t\tif (revoked)\n-\t\t\t\tpriv-\u003erevoked = true;\n-\t\t\tdma_buf_invalidate_mappings(priv-\u003edmabuf);\n-\t\t\tdma_resv_wait_timeout(priv-\u003edmabuf-\u003eresv,\n-\t\t\t\t\t DMA_RESV_USAGE_BOOKKEEP, false,\n-\t\t\t\t\t MAX_SCHEDULE_TIMEOUT);\n-\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\t\t\tif (revoked) {\n-\t\t\t\tkref_put(\u0026priv-\u003ekref, vfio_pci_dma_buf_done);\n-\t\t\t\twait_for_completion(\u0026priv-\u003ecomp);\n-\t\t\t\t/*\n-\t\t\t\t * Re-arm the registered kref reference and the\n-\t\t\t\t * completion so the post-revoke state matches the\n-\t\t\t\t * post-creation state. An un-revoke followed by a\n-\t\t\t\t * new mapping needs the kref to be non-zero before\n-\t\t\t\t * kref_get(), and vfio_pci_dma_buf_cleanup()\n-\t\t\t\t * delegates its drain back through this revoke\n-\t\t\t\t * path on a possibly-already-revoked dma-buf.\n-\t\t\t\t */\n-\t\t\t\tkref_init(\u0026priv-\u003ekref);\n-\t\t\t\treinit_completion(\u0026priv-\u003ecomp);\n-\t\t\t} else {\n-\t\t\t\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\t\t\t\tpriv-\u003erevoked = false;\n-\t\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\t\t\t}\n-\t\t}\n+\t\tvfio_pci_dma_buf_set_status(priv, revoked ?\n+\t\t\t\t\t VFIO_PCI_DMABUF_REVOKED :\n+\t\t\t\t\t VFIO_PCI_DMABUF_OK);\n \t\tfput(priv-\u003edmabuf-\u003efile);\n \t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n }\n \n void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n \t */\n \tvfio_pci_dma_buf_move(vdev, true);\n \n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n \tlist_for_each_entry_safe(priv, tmp, \u0026vdev-\u003edmabufs, dmabufs_elm) {\n \t\tif (!get_file_active(\u0026priv-\u003edmabuf-\u003efile))\n \t\t\tcontinue;\n \n \t\tlist_del_init(\u0026priv-\u003edmabufs_elm);\n-\t\tpriv-\u003evdev = NULL;\n+\t\tWRITE_ONCE(priv-\u003evdev, NULL);\n \t\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n \t\tfput(priv-\u003edmabuf-\u003efile);\n \t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n+\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n+int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz)\n+{\n+\tstruct vfio_device_feature_dma_buf_revoke db_revoke;\n+\tstruct vfio_pci_dma_buf *priv;\n+\tstruct dma_buf *dmabuf;\n+\tint ret;\n+\n+\tif (!vdev-\u003epci_ops || !vdev-\u003epci_ops-\u003eget_dmabuf_phys)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tret = vfio_check_feature(flags, argsz,\n+\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n+\t\t\t\t sizeof(db_revoke));\n+\tif (ret != 1)\n+\t\treturn ret;\n+\n+\tif (copy_from_user(\u0026db_revoke, arg, sizeof(db_revoke)))\n+\t\treturn -EFAULT;\n+\n+\tdmabuf = dma_buf_get(db_revoke.dmabuf_fd);\n+\tif (IS_ERR(dmabuf))\n+\t\treturn PTR_ERR(dmabuf);\n+\n+\tpriv = dmabuf-\u003epriv;\n+\t/*\n+\t * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_\n+\t * relates to the VFIO device it was provided with.\n+\t *\n+\t * If the DMABUF relates to this vdev then priv-\u003evdev is\n+\t * stable because this open fd prevents cleanup.\n+\t *\n+\t * If it relates to a different vdev, reading priv-\u003evdev might\n+\t * race with a concurrent cleanup on that device. But if so,\n+\t * it points to a non-matching vdev or NULL and is unusable\n+\t * either way.\n+\t */\n+\tif (dmabuf-\u003eops != \u0026vfio_pci_dmabuf_ops ||\n+\t READ_ONCE(priv-\u003evdev) != vdev) {\n+\t\tret = -ENODEV;\n+\t\tgoto out_put_buf;\n+\t}\n+\n+\t/*\n+\t * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()\n+\t * from getting it and then blocking all devices in the dev_set behind\n+\t * this revoke's drain.\n+\t */\n+\tdown_read(\u0026vdev-\u003ememory_lock);\n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_DEAD) {\n+\t\tret = -EBADFD;\n+\t} else {\n+\t\tvfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);\n+\t\tret = 0;\n+\t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n+\tup_read(\u0026vdev-\u003ememory_lock);\n+\n+out_put_buf:\n+\tdma_buf_put(dmabuf);\n+\n+\treturn ret;\n+}\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\ndiff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h\nindex 4e7162234a2eb..ca12221af5553 100644\n--- a/drivers/vfio/pci/vfio_pci_priv.h\n+++ b/drivers/vfio/pci/vfio_pci_priv.h\n@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {\n \tbool\t\t\ttest_mem;\n };\n \n+enum vfio_pci_dma_buf_status {\n+\tVFIO_PCI_DMABUF_OK = 0,\n+\tVFIO_PCI_DMABUF_REVOKED = 1,\n+\tVFIO_PCI_DMABUF_DEAD = 2,\n+};\n+\n+struct vfio_pci_dma_buf {\n+\tstruct dma_buf *dmabuf;\n+\tstruct vfio_pci_core_device *vdev;\n+\tstruct list_head dmabufs_elm;\n+\tsize_t size;\n+\tstruct phys_vec *phys_vec;\n+\tstruct p2pdma_provider *provider;\n+\tstruct file *vfile;\n+\tu32 nr_ranges;\n+\tstruct kref kref;\n+\tstruct completion comp;\n+\tunsigned long vma_pgoff_adjust;\n+\tenum vfio_pci_dma_buf_status status;\n+};\n+\n bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);\n void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);\n \n@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);\n int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,\n \t\t\t pci_power_t state);\n \n-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);\n+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);\n+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);\n u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);\n void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,\n \t\t\t\t\tu16 cmd);\n@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)\n \treturn (pdev-\u003eclass \u003e\u003e 8) == PCI_CLASS_DISPLAY_VGA;\n }\n \n+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,\n+\t\t\t struct vfio_pci_dma_buf *priv,\n+\t\t\t struct vm_area_struct *vma,\n+\t\t\t unsigned long address,\n+\t\t\t unsigned int order,\n+\t\t\t unsigned long *out_pfn);\n+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vm_area_struct *vma,\n+\t\t\t\t u64 phys_start, u64 req_len,\n+\t\t\t\t unsigned int res_index);\n+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);\n+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);\n+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);\n+\n #ifdef CONFIG_VFIO_PCI_DMABUF\n int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \t\t\t\t struct vfio_device_feature_dma_buf __user *arg,\n \t\t\t\t size_t argsz);\n-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);\n-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);\n+int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz);\n #else\n static inline int\n vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n {\n \treturn -ENOTTY;\n }\n-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n-{\n-}\n-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,\n-\t\t\t\t\t bool revoked)\n+static inline int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz)\n {\n+\treturn -ENOTTY;\n }\n #endif\n \ndiff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h\nindex d15b2b31d3c91..952a2c196ad42 100644\n--- a/include/linux/dma-buf.h\n+++ b/include/linux/dma-buf.h\n@@ -571,6 +571,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);\n struct dma_buf *dma_buf_get(int fd);\n void dma_buf_put(struct dma_buf *dmabuf);\n \n+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);\n+\n struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,\n \t\t\t\t\tenum dma_data_direction);\n void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,\ndiff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h\nindex 9a1674c152aa2..44891fdb7c76e 100644\n--- a/include/linux/vfio_pci_core.h\n+++ b/include/linux/vfio_pci_core.h\n@@ -129,11 +129,13 @@ struct vfio_pci_core_device {\n \tbool\t\t\tdisable_idle_d3:1;\n \tbool\t\t\tnointxmask:1;\n \tbool\t\t\tdisable_vga:1;\n+\tbool\t\t\tzap_bars_on_revoke:1;\n \t/* Flags modified at runtime - dedicated storage unit */\n \tbool\t\t\tneeds_reset;\n \tbool\t\t\tpm_intx_masked;\n \tbool\t\t\tpm_runtime_engaged;\n \tbool\t\t\tsriov_active;\n+\tbool\t\t\tbars_revoked;\n \tstruct pci_saved_state\t*pci_saved_state;\n \tstruct pci_saved_state\t*pm_save;\n \tint\t\t\tioeventfds_nr;\n@@ -148,6 +150,7 @@ struct vfio_pci_core_device {\n \tstruct vfio_pci_core_device\t*sriov_pf_core_dev;\n \tstruct notifier_block\tnb;\n \tstruct rw_semaphore\tmemory_lock;\n+\tstruct rw_semaphore\tdmabuf_lock;\n \tstruct list_head\tdmabufs;\n };\n \ndiff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h\nindex e41437fa17ad0..d3c6057983e09 100644\n--- a/include/uapi/linux/vfio.h\n+++ b/include/uapi/linux/vfio.h\n@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {\n \n #define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13\n \n+/**\n+ * Given a DMABUF fd previously exported from the same device by\n+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that\n+ * access to the corresponding DMABUF is immediately revoked. On\n+ * successful return, the buffer is no longer accessible through any\n+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future\n+ * mmap()s and map/attach requests from any new/existing importer.\n+ *\n+ * Return: 0 on success, -1 and errno is set on failure:\n+ *\n+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.\n+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.\n+ * ENODEV: The DMABUF was not exported from this device.\n+ * EBADFD: The DMABUF is already revoked by this feature.\n+ *\n+ * Additionally, common errors can occur: EFAULT accessing the struct,\n+ * or EINVAL requesting an unsupported feature op.\n+ */\n+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14\n+\n+struct vfio_device_feature_dma_buf_revoke {\n+\t__s32\tdmabuf_fd;\n+};\n+\n /* -------- API for Type1 VFIO IOMMU -------- */\n \n /**\n"
} |
| 2/1 | 2026/09/26 21:06 | agent | patch-evaluator |
2mModel:core Args: null Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_core_mmap",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch modifies reachable code in DMA-BUF and VFIO-PCI. It converts VFIO PCI BAR memory mappings to be backed by DMABUFs, introduces a new fault handler (vfio_pci_mmap_huge_fault) and PFN search logic (vfio_pci_dma_buf_find_pfn), adds a new user-facing device feature for buffer revocation (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE), restructures BAR revocation synchronization with lockdep assertions, and refactors dma_buf_set_name and its ioctl handling. Fuzzing is warranted to detect races, regressions, and assertion failures.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 0b43d95e7d80a0760fdb300edd1dd05f46dba412
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Sep 26 21:06:15 2026 +0000
syz-cluster: applied patch under review
diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
index 4c9add51f9ef5..e4ec504e92884 100644
--- a/drivers/dma-buf/dma-buf.c
+++ b/drivers/dma-buf/dma-buf.c
@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)
return events;
}
-/**
- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.
- * It could support changing the name of the dma-buf if the same
- * piece of memory is used for multiple purpose between different devices.
- *
- * @dmabuf: [in] dmabuf buffer that will be renamed.
- * @buf: [in] A piece of userspace memory that contains the name of
- * the dma-buf.
- *
- * Returns 0 on success. If the dma-buf buffer is already attached to
- * devices, return -EBUSY.
- *
- */
-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)
-{
- char *name = strndup_user(buf, DMA_BUF_NAME_LEN);
-
- if (IS_ERR(name))
- return PTR_ERR(name);
-
- spin_lock(&dmabuf->name_lock);
- kfree(dmabuf->name);
- dmabuf->name = name;
- spin_unlock(&dmabuf->name_lock);
-
- return 0;
-}
-
#if IS_ENABLED(CONFIG_SYNC_FILE)
static long dma_buf_export_sync_file(struct dma_buf *dmabuf,
void __user *user_data)
@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,
return ret;
case DMA_BUF_SET_NAME_A:
- case DMA_BUF_SET_NAME_B:
- return dma_buf_set_name(dmabuf, (const char __user *)arg);
+ case DMA_BUF_SET_NAME_B: {
+ char *name = strndup_user((const char __user *)arg,
+ DMA_BUF_NAME_LEN);
+
+ if (IS_ERR(name))
+ return PTR_ERR(name);
+
+ ret = dma_buf_set_name(dmabuf, name);
+ if (ret)
+ kfree(name);
+
+ return ret;
+ }
#if IS_ENABLED(CONFIG_SYNC_FILE)
case DMA_BUF_IOCTL_EXPORT_SYNC_FILE:
@@ -874,6 +857,41 @@ void dma_buf_put(struct dma_buf *dmabuf)
}
EXPORT_SYMBOL_NS_GPL(dma_buf_put, "DMA_BUF");
+/**
+ * dma_buf_set_name - Set a dmabuf's name
+ * Intended to be used by the exporter to set a name for debug
+ * purposes. This can also change an existing name if the same piece
+ * of memory is used for multiple purposes over time.
+ *
+ * @dmabuf: [in] dmabuf buffer that will be renamed.
+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or
+ * similar. This takes ownership of the allocation
+ * on success, which will be kfree()d when the
+ * dmabuf is released or a new name assigned.
+ *
+ * Returns 0 on success, -EINVAL if the name is NULL, or -E2BIG if the
+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.
+ */
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)
+{
+ if (!name)
+ return -EINVAL;
+
+ /* dmabuffs_dname() won't use the string if the length
+ * (including terminator) exceeds DMA_BUF_NAME_LEN:
+ */
+ if (strlen(name) >= DMA_BUF_NAME_LEN)
+ return -E2BIG;
+
+ spin_lock(&dmabuf->name_lock);
+ kfree(dmabuf->name);
+ dmabuf->name = name;
+ spin_unlock(&dmabuf->name_lock);
+
+ return 0;
+}
+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, "DMA_BUF");
+
static int dma_buf_wrap_sg_table(struct sg_table **sg_table)
{
struct scatterlist *to_sg, *from_sg;
diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig
index 296bf01e185ec..c6d6fb09af86c 100644
--- a/drivers/vfio/pci/Kconfig
+++ b/drivers/vfio/pci/Kconfig
@@ -6,6 +6,7 @@ config VFIO_PCI_CORE
tristate
select VFIO_VIRQFD
select IRQ_BYPASS_MANAGER
+ select DMA_SHARED_BUFFER
config VFIO_PCI_INTX
def_bool y if !S390
@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM
To enable s390x KVM vfio-pci extensions, say Y.
config VFIO_PCI_DMABUF
- def_bool y if VFIO_PCI_CORE && PCI_P2PDMA && DMA_SHARED_BUFFER
+ def_bool y if PCI_P2PDMA
+ depends on VFIO_PCI_CORE
source "drivers/vfio/pci/mlx5/Kconfig"
diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile
index 6138f1bf241df..881452ea89be0 100644
--- a/drivers/vfio/pci/Makefile
+++ b/drivers/vfio/pci/Makefile
@@ -1,8 +1,7 @@
# SPDX-License-Identifier: GPL-2.0-only
-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o
+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o
vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o
-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o
obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o
vfio-pci-y := vfio_pci.o
diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
index 86362ec424a50..14622556355eb 100644
--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);
struct pci_dev *pdev = to_pci_dev(core_vdev->dev);
struct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);
+ int ret;
hisi_acc_vdev->vf_id = pci_iov_vf_id(pdev) + 1;
hisi_acc_vdev->pf_qm = pf_qm;
@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
core_vdev->migration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;
core_vdev->mig_ops = &hisi_acc_vfio_pci_migrn_state_ops;
- return vfio_pci_core_init_dev(core_vdev);
+ ret = vfio_pci_core_init_dev(core_vdev);
+ if (ret)
+ return ret;
+ /*
+ * hisi_acc_vfio_pci_mmap() calls down to
+ * vfio_pci_core_mmap(), so BAR mappings are still
+ * DMABUF-backed. They don't require a zap on revoke, so opt
+ * out:
+ */
+ hisi_acc_vdev->core_device.zap_bars_on_revoke = false;
+
+ return 0;
}
static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index 9914f3ac69aef..cef337f4e8f2e 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
virt_mem = !!(le16_to_cpu(*virt_cmd) & PCI_COMMAND_MEMORY);
new_mem = !!(new_cmd & PCI_COMMAND_MEMORY);
- if (!new_mem) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (!new_mem)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
/*
* If the user is writing mem/io enable (new_mem/io) and we
@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
*virt_cmd |= cpu_to_le16(new_cmd & mask);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state)
{
- if (state >= PCI_D3hot) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (state >= PCI_D3hot)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
vfio_pci_set_power_state(vdev, state);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_EXP_DEVCAP_FLR)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_AF_CAP_FLR) && (cap & PCI_AF_CAP_TP)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d875..8860185cff49f 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -13,6 +13,8 @@
#include <linux/aperture.h>
#include <linux/debugfs.h>
#include <linux/device.h>
+#include <linux/dma-buf.h>
+#include <linux/dma-resv.h>
#include <linux/eventfd.h>
#include <linux/file.h>
#include <linux/interrupt.h>
@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
* The vdev power related flags are protected with 'memory_lock'
* semaphore.
*/
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
if (vdev->pm_runtime_engaged) {
up_write(&vdev->memory_lock);
@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)
down_write(&vdev->memory_lock);
__vfio_pci_runtime_pm_exit(vdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)
*/
down_write(&vdev->memory_lock);
if (vdev->pm_wake_eventfd_ctx) {
- eventfd_signal(vdev->pm_wake_eventfd_ctx);
+ struct eventfd_ctx *ctx = vdev->pm_wake_eventfd_ctx;
+
+ vdev->pm_wake_eventfd_ctx = NULL;
__vfio_pci_runtime_pm_exit(vdev);
+ if (__vfio_pci_memory_enabled(vdev))
+ vfio_pci_unrevoke_bars(vdev);
+ eventfd_signal(ctx);
+ eventfd_ctx_put(ctx);
}
up_write(&vdev->memory_lock);
@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
vdev->has_vga = true;
vfio_pci_core_map_bars(vdev);
+ vdev->bars_revoked = false;
return 0;
@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
return ret;
}
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);
+
static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
void __user *arg)
{
@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
if (!vdev->reset_works)
return -EINVAL;
- vfio_pci_zap_and_down_write_memory_lock(vdev);
+ down_write(&vdev->memory_lock);
/*
* This function can be invoked while the power state is non-D0. If
@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
* have NoSoftRst-, the reset function can cause the PCI config space
* reset without restoring the original state (saved locally in
* 'vdev->pm_save').
+ *
+ * The zap is done after making the device accessible in D0,
+ * because a DMABUF importer could access the device as part
+ * of its revocation cleanup.
*/
vfio_pci_set_power_state(vdev, PCI_D0);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_revoke_bars(vdev);
+
ret = pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
return ret;
@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
return vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);
case VFIO_DEVICE_FEATURE_ZPCI_ERROR:
return vfio_pci_zdev_feature_err(device, flags, arg, argsz);
+ case VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:
+ return vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);
default:
return -ENOTTY;
}
@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu
}
EXPORT_SYMBOL_GPL(vfio_pci_core_write);
-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)
{
- struct vfio_device *core_vdev = &vdev->vdev;
- loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
- loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
- loff_t len = end - start;
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, true);
- unmap_mapping_range(core_vdev->inode->i_mapping, start, len, true);
+ /*
+ * If a driver could possibly create BAR mappings in the
+ * vdev's address_space, do an additional zap on revoke. See
+ * vfio_pci_core_init_dev().
+ */
+ if (vdev->zap_bars_on_revoke) {
+ struct vfio_device *core_vdev = &vdev->vdev;
+ loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
+ loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
+ loff_t len = end - start;
+
+ unmap_mapping_range(core_vdev->inode->i_mapping,
+ start, len, true);
+ }
}
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)
{
down_write(&vdev->memory_lock);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
+}
+
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)
+{
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, false);
}
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)
@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c
up_write(&vdev->memory_lock);
}
-static unsigned long vma_to_pfn(struct vm_area_struct *vma)
-{
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- int index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
- u64 pgoff;
-
- pgoff = vma->vm_pgoff &
- ((1U << (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);
-
- return (pci_resource_start(vdev->pdev, index) >> PAGE_SHIFT) + pgoff;
-}
-
vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,
struct vm_fault *vmf,
unsigned long pfn,
@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,
unsigned int order)
{
struct vm_area_struct *vma = vmf->vma;
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- unsigned long addr = vmf->address & ~((PAGE_SIZE << order) - 1);
- unsigned long pgoff = linear_page_delta(vma, addr);
- unsigned long pfn = vma_to_pfn(vma) + pgoff;
- vm_fault_t ret = VM_FAULT_FALLBACK;
-
- if (is_aligned_for_order(vma, addr, pfn, order)) {
- scoped_guard(rwsem_read, &vdev->memory_lock)
- ret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);
+ struct vfio_pci_dma_buf *priv = vma->vm_private_data;
+ struct vfio_pci_core_device *vdev;
+ unsigned long pfn = 0;
+ vm_fault_t ret = VM_FAULT_SIGBUS;
+
+ /*
+ * The only thing this can rely on is that the DMABUF relating
+ * to the VMA's vm_file exists (priv).
+ *
+ * A DMABUF for a VFIO device fd mmap() holds a reference to
+ * the original VFIO device fd, but an explicitly-exported
+ * DMABUF does not. The original fd might have closed,
+ * meaning this fault can race with
+ * vfio_pci_dma_buf_cleanup(), meaning the buffer could have
+ * been revoked (in which case priv->vdev might be NULL), and
+ * the VFIO device registration might have been dropped.
+ *
+ * With the goal of taking vdev locks in a world where vdev
+ * might not still exist:
+ *
+ * 1. Take the resv lock on the DMABUF:
+ * - If racing cleanup got in first, the buffer is revoked;
+ * stop/exit if so.
+ * - If we got in first, the buffer is not revoked so vdev is
+ * non-NULL, accessible, and cleanup _has not yet put the
+ * VFIO device registration_. So, the device refcount must
+ * be >0.
+ *
+ * 2. Take vfio_device registration (refcount guaranteed >0
+ * hereafter).
+ *
+ * 3. Unlock the DMABUF's resv lock:
+ * - A racing cleanup can now complete.
+ * - But, the device refcount >0, meaning the vfio_device
+ * (and vfio_pcie_core device vdev) have not yet been
+ * freed. vdev is accessible, even if the DMABUF has been
+ * revoked or cleanup has happened, because
+ * vfio_unregister_group_dev() can't complete.
+ *
+ * 4. Take the vdev->memory_lock then vdev->dmabuf_lock:
+ * - Either the DMABUF is usable, or has been cleaned up.
+ * - It's not necessary to also take the resv lock, because
+ * the status/vdev can't change while dmabuf_lock is held.
+ * - Test the DMABUF revocation status again: if it was
+ * revoked between 1 and 4, return a SIGBUS. Otherwise,
+ * return a PFN.
+ *
+ * 5. Unlock, done.
+ */
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+
+ if (priv->status != VFIO_PCI_DMABUF_OK) {
+ pr_debug_ratelimited("%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\n",
+ __func__, vmf->address, vma->vm_pgoff);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+
+ /* If the buffer isn't revoked, vdev is valid */
+ vdev = priv->vdev;
+
+ if (!vfio_device_try_get_registration(&vdev->vdev)) {
+ /*
+ * If vdev != NULL (above), the registration should
+ * already be >0 and so this try_get should never
+ * fail.
+ */
+ dev_warn_ratelimited(&vdev->pdev->dev,
+ "%s: Unexpected registration failure\n",
+ __func__);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+ dma_resv_unlock(priv->dmabuf->resv);
+
+ /* memory_lock for vfio_pci_vmf_insert_pfn() */
+ down_read(&vdev->memory_lock);
+ /* Re-test revocation status under dmabuf_lock */
+ down_read(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_OK) {
+ int pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,
+ vmf->address,
+ order, &pfn);
+
+ if (pres == 0)
+ ret = vfio_pci_vmf_insert_pfn(vdev, vmf,
+ pfn, order);
+ else if (pres == -ERANGE)
+ ret = VM_FAULT_FALLBACK;
}
+ up_read(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
dev_dbg_ratelimited(&vdev->pdev->dev,
- "%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\n",
- __func__, order,
- vma->vm_pgoff >>
- (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),
- pgoff, (unsigned int)ret);
+ "%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\n",
+ __func__, order, pfn, vmf->address,
+ vma->vm_pgoff, (unsigned int)ret);
+ vfio_device_put_registration(&vdev->vdev);
return ret;
}
@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {
#endif
};
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)
+{
+ vma->vm_ops = &vfio_pci_mmap_ops;
+}
+
int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)
{
struct vfio_pci_core_device *vdev =
@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
unsigned int index;
u64 phys_len, req_len, pgoff, req_start;
void __iomem *bar_io;
+ int ret;
index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
if (IS_ERR(bar_io))
return PTR_ERR(bar_io);
- vma->vm_private_data = vdev;
+ ret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,
+ pci_resource_start(pdev, index),
+ req_len, index);
+ if (ret)
+ return ret;
+
vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
return ret;
INIT_LIST_HEAD(&vdev->dmabufs);
init_rwsem(&vdev->memory_lock);
+ init_rwsem(&vdev->dmabuf_lock);
xa_init(&vdev->ctx);
+ /*
+ * If a driver overrides .mmap, it has to be assumed that it
+ * might not use the DMABUF-backed core mmap; this flag
+ * enables a zap at revoke time. A driver can opt out by
+ * clearing this flag at init, if their .mmap override calls
+ * down to vfio_pci_core_mmap().
+ */
+ if (vdev->vdev.ops->mmap != vfio_pci_core_mmap)
+ vdev->zap_bars_on_revoke = true;
+
return 0;
}
EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);
@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
}
/*
- * Take the memory write lock for each device and zap BAR
- * mappings to prevent the user accessing the device while in
- * reset. Locking multiple devices is prone to deadlock,
+ * Take the memory write lock for each device and
+ * zap/revoke BAR mappings to prevent the user (or
+ * peers) accessing the device while in reset.
+ * Locking multiple devices is prone to deadlock,
* runaway and unwind if we hit contention.
*/
if (!down_write_trylock(&vdev->memory_lock)) {
@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
break;
}
- vfio_pci_dma_buf_move(vdev, true);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
}
if (!list_entry_is_head(vdev,
@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
list_for_each_entry_from_reverse(vdev, &dev_set->device_list,
vdev.dev_set_list) {
if (vdev->vdev.open_count && __vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d68..b57bfaefd9fae 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -3,25 +3,14 @@
*/
#include <linux/dma-buf-mapping.h>
#include <linux/pci-p2pdma.h>
+#include <linux/dma-buf.h>
#include <linux/dma-resv.h>
#include "vfio_pci_priv.h"
MODULE_IMPORT_NS("DMA_BUF");
-struct vfio_pci_dma_buf {
- struct dma_buf *dmabuf;
- struct vfio_pci_core_device *vdev;
- struct list_head dmabufs_elm;
- size_t size;
- struct phys_vec *phys_vec;
- struct p2pdma_provider *provider;
- u32 nr_ranges;
- struct kref kref;
- struct completion comp;
- u8 revoked : 1;
-};
-
+#ifdef CONFIG_VFIO_PCI_DMABUF
static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
struct dma_buf_attachment *attachment)
{
@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
if (!attachment->peer2peer)
return -EOPNOTSUPP;
- if (priv->revoked)
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
return -ENODEV;
if (!dma_buf_attach_revocable(attachment))
@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
return 0;
}
+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)
+{
+ struct vfio_pci_dma_buf *priv = dmabuf->priv;
+
+ /*
+ * dma_buf_mmap_internal() has asserted that the VMA is
+ * contained within the DMABUF size before calling this.
+ *
+ * Also, if we observe that the buffer is revoked now then
+ * refuse the mmap(). This is a belt-and-braces early failure
+ * to ease debugging a revoked buffer being used. Userspace
+ * might also race an mmap() against an explicit revocation,
+ * or an action causing a revoke; race scenarios are still
+ * safe because the fault handler ultimately prevents access
+ * to a revoked buffer if it isn't caught here.
+ */
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
+ return -ENODEV;
+ /*
+ * Make clear that anything with an offset adjustment is
+ * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()
+ * maths would underflow; this doesn't happen through the
+ * regular DMABUF export path used with this mmap(). A DMABUF
+ * implicitly created for BAR mmap could have adjust > 0, but
+ * these can't currently be re-opened and mmap()ed again.
+ * Catch here in case that assumption ever changes.
+ */
+ if (priv->vma_pgoff_adjust)
+ return -EINVAL;
+ if ((vma->vm_flags & VM_SHARED) == 0)
+ return -EINVAL;
+
+ vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
+ vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
+
+ /* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */
+ vm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |
+ VM_DONTEXPAND | VM_DONTDUMP);
+ vma->vm_private_data = priv;
+ vfio_pci_set_vma_ops(vma);
+
+ return 0;
+}
+#else
+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
+ struct dma_buf_attachment *attachment)
+{
+ /*
+ * Explicit export can't occur without the DMABUF feature, but
+ * DMABUFs are implicitly created for BAR mappings. An
+ * .attach that fails prevents dma_buf_attach().
+ */
+ return -EOPNOTSUPP;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
static void vfio_pci_dma_buf_done(struct kref *kref)
{
struct vfio_pci_dma_buf *priv =
@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,
dma_resv_assert_held(priv->dmabuf->resv);
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return ERR_PTR(-ENODEV);
ret = dma_buf_phys_vec_to_sgt(attachment, priv->provider,
@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)
* The refcount prevents both.
*/
if (priv->vdev) {
- down_write(&priv->vdev->memory_lock);
+ down_write(&priv->vdev->dmabuf_lock);
list_del_init(&priv->dmabufs_elm);
- up_write(&priv->vdev->memory_lock);
+ up_write(&priv->vdev->dmabuf_lock);
vfio_device_put_registration(&priv->vdev->vdev);
}
+ if (priv->vfile)
+ fput(priv->vfile);
kfree(priv->phys_vec);
kfree(priv);
}
static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
.attach = vfio_pci_dma_buf_attach,
+#ifdef CONFIG_VFIO_PCI_DMABUF
+ .mmap = vfio_pci_dma_buf_mmap,
+#endif
.map_dma_buf = vfio_pci_dma_buf_map,
.unmap_dma_buf = vfio_pci_dma_buf_unmap,
.release = vfio_pci_dma_buf_release,
};
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long fault_addr,
+ unsigned int order,
+ unsigned long *out_pfn)
+{
+ /*
+ * Given a VMA (start, end, pgoffs) and a fault address,
+ * search the corresponding DMABUF's phys_vec[] to find the
+ * range representing the address's offset into the VMA, and
+ * its PFN. vdev must be the device that the DMABUF priv was
+ * exported from; vdev->dmabuf_lock must be held, and priv
+ * must not be revoked.
+ *
+ * The phys_vec[] ranges represent contiguous spans of VAs
+ * upwards from the buffer offset 0; the actual PFNs might be
+ * in any order, overlap/alias, etc. Calculate an offset of
+ * the desired page given VMA start/pgoff and address, then
+ * search upwards from 0 to find which span contains it.
+ *
+ * On success, a valid PFN for a page sized by 'order' is
+ * returned into out_pfn.
+ *
+ * Failure occurs if:
+ * - A hugepage would cross the edge of the VMA,
+ * - A hugepage isn't entirely contained within a range
+ * (including where it straddles the boundary between
+ * ranges),
+ * - We find a range, but the final PFN isn't aligned to the
+ * requested order.
+ *
+ * Upon failure, -ERANGE is returned and the caller is
+ * expected to try again with a smaller order, which will
+ * eventually succeed.
+ *
+ * It's suboptimal if DMABUFs are created with neighbouring
+ * ranges that are physically contiguous, since hugepages
+ * can't straddle range boundaries. (The construction of the
+ * ranges should merge them in this case.)
+ *
+ * Finally, vma_pgoff_adjust is used with a DMABUF created for
+ * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
+ * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
+ * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
+ * To avoid double-offsetting in this scenario, subtracting
+ * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
+ * the effective offset. This also removes the VFIO region
+ * index encoded in vm_pgoff for VFIO BAR mmaps.
+ */
+
+ const unsigned long pagesize = PAGE_SIZE << order;
+ unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) <<
+ PAGE_SHIFT;
+ unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
+ unsigned long rounded_page_end = rounded_page_addr + pagesize;
+ unsigned long fault_offset;
+ unsigned long fault_offset_end;
+ unsigned long range_start_offset = 0;
+ unsigned int i;
+ int ret;
+
+ if (unlikely(!vdev))
+ return -ENODEV;
+
+ /* This prevents the dmabuf revocation state from changing under us */
+ lockdep_assert_held(&vdev->dmabuf_lock);
+
+ if (unlikely(priv->vdev != vdev || priv->status != VFIO_PCI_DMABUF_OK))
+ return -ENODEV;
+
+ if (rounded_page_addr < vma->vm_start || rounded_page_end > vma->vm_end) {
+ if (order > 0)
+ return -ERANGE;
+
+ /* A fault address outside of the VMA is absurd. */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "Fault addr 0x%lx outside VMA 0x%lx-0x%lx\n",
+ fault_addr, vma->vm_start, vma->vm_end);
+ return -EFAULT;
+ }
+
+ /*
+ * fault_offset[_end] is the span within the DMABUF
+ * corresponding to the faulting page:
+ */
+ if (unlikely(check_add_overflow(rounded_page_addr - vma->vm_start,
+ vma_off, &fault_offset) ||
+ check_add_overflow(fault_offset, pagesize,
+ &fault_offset_end)))
+ return -EFAULT;
+
+ /*
+ * Iterate over ranges in the buffer, summing their lengths:
+ * range_start_offset represents the current range's starting
+ * offset in the buffer (from 0 upwards).
+ *
+ * A failure for order == 0 is unexpected, and triggers a
+ * fault/warn.
+ */
+ ret = (order == 0) ? -EFAULT : -ERANGE;
+
+ for (i = 0; i < priv->nr_ranges; i++) {
+ size_t range_len = priv->phys_vec[i].len;
+
+ /* Early exit if range starts after the page end */
+ if (fault_offset_end <= range_start_offset)
+ break;
+
+ if (fault_offset >= range_start_offset &&
+ fault_offset_end <= range_start_offset + range_len) {
+ /*
+ * The faulting page is wholly contained
+ * within the span represented by this range,
+ * so validate PFN alignment for the order.
+ * The if() condition ensures the pfn
+ * arithmetic won't overflow.
+ */
+ unsigned long pfn =
+ ((fault_offset - range_start_offset) +
+ priv->phys_vec[i].paddr) >> PAGE_SHIFT;
+
+ if (IS_ALIGNED(pfn, 1 << order)) {
+ *out_pfn = pfn;
+ ret = 0;
+ }
+ /*
+ * Else order > 0; ERANGE retries with smaller
+ * order
+ */
+ break;
+ }
+ range_start_offset += range_len;
+ }
+
+ if (order == 0 && ret != 0)
+ /*
+ * The address fell outside of the span represented by
+ * the (concatenated) ranges. As setup of a mapping
+ * ensures that the VMA is <= the total size of the
+ * ranges this should never happen. If it does, warn
+ * and SIGBUS.
+ */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\n",
+ fault_addr, order, vma->vm_start, vma->vm_end,
+ vma->vm_pgoff, priv->nr_ranges, priv->size);
+
+ return ret;
+}
+
+/*
+ * Create a DMABUF corresponding to priv, add it to vdev->dmabufs list
+ * for tracking (meaning cleanup or revocation will zap it), and take
+ * a vfio_device registration.
+ */
+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv, u32 flags)
+{
+ DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
+
+ if (!vfio_device_try_get_registration(&vdev->vdev))
+ return -ENODEV;
+
+ exp_info.ops = &vfio_pci_dmabuf_ops;
+ exp_info.size = priv->size;
+ exp_info.flags = flags;
+ exp_info.priv = priv;
+
+ priv->dmabuf = dma_buf_export(&exp_info);
+ if (IS_ERR(priv->dmabuf)) {
+ vfio_device_put_registration(&vdev->vdev);
+ return PTR_ERR(priv->dmabuf);
+ }
+
+ kref_init(&priv->kref);
+ init_completion(&priv->comp);
+
+ /* dma_buf_put() now frees priv */
+ INIT_LIST_HEAD(&priv->dmabufs_elm);
+
+ /*
+ * dmabuf_lock synchronises access (R) or updates (W) to the
+ * vdev->dmabufs list and to bars_revoked (see below). The
+ * revocation state of DMABUF elements in the list is written
+ * holding both dmabuf_lock(W) and resv, and tested with
+ * either.
+ *
+ * (memory_lock, if held ->) dmabuf_lock -> resv
+ *
+ * NOTE: memory_lock is strictly avoided here, to avoid a
+ * dependency on memory_lock when mmap_lock is held, when
+ * mmap() leads to export. vfio-pci variant drivers are
+ * permitted to hold memory_lock across actions that might
+ * fault (such as user access); a deadlock could result when
+ * that fault path attempts to take mmap_lock (if held by an
+ * export waiting for memory_lock).
+ *
+ * vdev->bars_revoked tracks the BAR revocation status updated
+ * via vfio_pci_dma_buf_move(), so the initial DMABUF state
+ * follows the same criteria that later update the DMABUF
+ * state (BAR zap, etc.).
+ */
+ lockdep_assert_not_held(&vdev->memory_lock);
+
+ down_write(&vdev->dmabuf_lock);
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ priv->status = vdev->bars_revoked ? VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK;
+ list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
+ dma_resv_unlock(priv->dmabuf->resv);
+ up_write(&vdev->dmabuf_lock);
+
+ return 0;
+}
+
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index)
+{
+ struct vfio_pci_dma_buf *priv;
+ unsigned long vma_pgoff = vma->vm_pgoff & (VFIO_PCI_OFFSET_MASK >> PAGE_SHIFT);
+ char *bufname;
+ int ret;
+
+ priv = kzalloc_obj(*priv);
+ if (!priv)
+ return -ENOMEM;
+
+ priv->phys_vec = kzalloc_obj(*priv->phys_vec);
+ if (!priv->phys_vec) {
+ ret = -ENOMEM;
+ goto err_free_priv;
+ }
+
+ /*
+ * Debug name: The absolute maximum size of the name
+ * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.
+ */
+ bufname = kasprintf(GFP_KERNEL, "vfio:%s/%x",
+ pci_name(vdev->pdev),
+ res_index);
+
+ if (!bufname) {
+ ret = -ENOMEM;
+ goto err_free_phys;
+ }
+
+ /*
+ * The DMABUF begins from the mmap()'s BAR offset, i.e. the
+ * start of the VMA corresponds to byte 0 of the DMABUF and
+ * byte (vma_pgoff << PAGE_SHIFT) of the BAR.
+ *
+ * vfio_pci_dma_buf_find_pfn() reverses this offset using
+ * vma_pgoff_adjust, so that ultimately a fault's offset from
+ * the start of the _VMA_ has a consistent usage whether the
+ * VMA originates from an mmap() of the VFIO device here or a
+ * direct DMABUF mmap(). Note vma_pgoff_adjust also includes
+ * the encoded VFIO region index, which cancels out the index
+ * encoded in vm_pgoff.
+ */
+ priv->vdev = vdev;
+ priv->size = req_len;
+ priv->nr_ranges = 1;
+ priv->vma_pgoff_adjust = vma->vm_pgoff;
+
+ /*
+ * The provider can be NULL _iff_ the DMABUF feature isn't
+ * supported, because it's only used by DMABUF import and
+ * attach is prohibited if the feature isn't present.
+ */
+ priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
+ if (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) && !priv->provider) {
+ ret = -EINVAL;
+ goto err_free_name;
+ }
+
+ priv->phys_vec[0].paddr = phys_start + ((u64)vma_pgoff << PAGE_SHIFT);
+ priv->phys_vec[0].len = priv->size;
+
+ ret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);
+ if (ret)
+ goto err_free_name;
+
+ if (dma_buf_set_name(priv->dmabuf, bufname)) {
+ dev_dbg_ratelimited(&vdev->pdev->dev,
+ "Failed to set map name '%s'\n",
+ bufname);
+ kfree(bufname);
+ }
+
+ /*
+ * Ownership of the DMABUF file transfers to the VMA so that
+ * other users can locate the DMABUF via a VA. Ownership of
+ * the original VFIO device file being mmap()ed transfers to
+ * priv, and is put when the DMABUF is released. This
+ * intentionally does not use get_file()/vma_set_file()
+ * because the references are already held, and ownership
+ * moves.
+ */
+ priv->vfile = vma->vm_file;
+ vma->vm_file = priv->dmabuf->file;
+ vma->vm_private_data = priv;
+
+ return 0;
+
+err_free_name:
+ kfree(bufname);
+err_free_phys:
+ kfree(priv->phys_vec);
+err_free_priv:
+ kfree(priv);
+ return ret;
+}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
/*
* This is a temporary "private interconnect" between VFIO DMABUF and iommufd.
* It allows the two co-operating drivers to exchange the physical address of
@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
return -EOPNOTSUPP;
priv = attachment->dmabuf->priv;
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return -ENODEV;
/* More than one range to iommufd will require proper DMABUF support */
@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
struct vfio_device_feature_dma_buf get_dma_buf = {};
struct vfio_region_dma_range *dma_ranges;
- DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
struct vfio_pci_dma_buf *priv;
size_t length;
int ret;
@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
dma_ranges = NULL;
- if (!vfio_device_try_get_registration(&vdev->vdev)) {
- ret = -ENODEV;
+ ret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);
+ if (ret)
goto err_free_phys;
- }
-
- exp_info.ops = &vfio_pci_dmabuf_ops;
- exp_info.size = priv->size;
- exp_info.flags = get_dma_buf.open_flags;
- exp_info.priv = priv;
-
- priv->dmabuf = dma_buf_export(&exp_info);
- if (IS_ERR(priv->dmabuf)) {
- ret = PTR_ERR(priv->dmabuf);
- goto err_dev_put;
- }
-
- kref_init(&priv->kref);
- init_completion(&priv->comp);
-
- /* dma_buf_put() now frees priv */
- INIT_LIST_HEAD(&priv->dmabufs_elm);
- down_write(&vdev->memory_lock);
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = !__vfio_pci_memory_enabled(vdev);
- list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
- dma_resv_unlock(priv->dmabuf->resv);
- up_write(&vdev->memory_lock);
-
/*
* dma_buf_fd() consumes the reference, when the file closes the dmabuf
* will be released.
@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
return ret;
-err_dev_put:
- vfio_device_put_registration(&vdev->vdev);
err_free_phys:
kfree(priv->phys_vec);
err_free_priv:
@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
return ret;
}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
+/*
+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives
+ * the guarantee that all future map/attach attempts will fail no
+ * matter what, whereas REVOKED can transition back to OK.
+ */
+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,
+ enum vfio_pci_dma_buf_status new_status)
+{
+ bool was_revoked;
+
+ /*
+ * Changes to the DMABUF's revocation status are synchronised
+ * using dmabuf_lock:
+ */
+ lockdep_assert_held_write(&priv->vdev->dmabuf_lock);
+
+ /* If DEAD, state can no longer change */
+ if (priv->status == VFIO_PCI_DMABUF_DEAD ||
+ priv->status == new_status)
+ return;
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ was_revoked = (priv->status == VFIO_PCI_DMABUF_REVOKED);
+
+ if (new_status != VFIO_PCI_DMABUF_OK) {
+ priv->status = new_status;
+
+ if (was_revoked) {
+ /*
+ * A REVOKED buffer is being marked DEAD.
+ * invalidate_mappings/unmap wait happened
+ * when it became REVOKED, don't wait again.
+ */
+ dma_resv_unlock(priv->dmabuf->resv);
+ return;
+ }
+ dma_buf_invalidate_mappings(priv->dmabuf);
+ dma_resv_wait_timeout(priv->dmabuf->resv,
+ DMA_RESV_USAGE_BOOKKEEP, false,
+ MAX_SCHEDULE_TIMEOUT);
+ dma_resv_unlock(priv->dmabuf->resv);
+ kref_put(&priv->kref, vfio_pci_dma_buf_done);
+ wait_for_completion(&priv->comp);
+ unmap_mapping_range(priv->dmabuf->file->f_mapping,
+ 0, 0, true);
+ /*
+ * Re-arm the registered kref reference and the
+ * completion so the post-revoke state matches the
+ * post-creation state. An un-revoke followed by a
+ * new mapping needs the kref to be non-zero before
+ * kref_get(), and vfio_pci_dma_buf_cleanup()
+ * delegates its drain back through this revoke
+ * path on a possibly-already-revoked dma-buf.
+ */
+ kref_init(&priv->kref);
+ reinit_completion(&priv->comp);
+ } else {
+ priv->status = VFIO_PCI_DMABUF_OK;
+ dma_resv_unlock(priv->dmabuf->resv);
+ }
+}
void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
{
@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
lockdep_assert_held_write(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ vdev->bars_revoked = revoked;
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
-
- if (priv->revoked != revoked) {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- if (revoked)
- priv->revoked = true;
- dma_buf_invalidate_mappings(priv->dmabuf);
- dma_resv_wait_timeout(priv->dmabuf->resv,
- DMA_RESV_USAGE_BOOKKEEP, false,
- MAX_SCHEDULE_TIMEOUT);
- dma_resv_unlock(priv->dmabuf->resv);
- if (revoked) {
- kref_put(&priv->kref, vfio_pci_dma_buf_done);
- wait_for_completion(&priv->comp);
- /*
- * Re-arm the registered kref reference and the
- * completion so the post-revoke state matches the
- * post-creation state. An un-revoke followed by a
- * new mapping needs the kref to be non-zero before
- * kref_get(), and vfio_pci_dma_buf_cleanup()
- * delegates its drain back through this revoke
- * path on a possibly-already-revoked dma-buf.
- */
- kref_init(&priv->kref);
- reinit_completion(&priv->comp);
- } else {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = false;
- dma_resv_unlock(priv->dmabuf->resv);
- }
- }
+ vfio_pci_dma_buf_set_status(priv, revoked ?
+ VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
}
void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
*/
vfio_pci_dma_buf_move(vdev, true);
+ down_write(&vdev->dmabuf_lock);
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
list_del_init(&priv->dmabufs_elm);
- priv->vdev = NULL;
+ WRITE_ONCE(priv->vdev, NULL);
vfio_device_put_registration(&vdev->vdev);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
up_write(&vdev->memory_lock);
}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
+{
+ struct vfio_device_feature_dma_buf_revoke db_revoke;
+ struct vfio_pci_dma_buf *priv;
+ struct dma_buf *dmabuf;
+ int ret;
+
+ if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
+ return -EOPNOTSUPP;
+
+ ret = vfio_check_feature(flags, argsz,
+ VFIO_DEVICE_FEATURE_SET,
+ sizeof(db_revoke));
+ if (ret != 1)
+ return ret;
+
+ if (copy_from_user(&db_revoke, arg, sizeof(db_revoke)))
+ return -EFAULT;
+
+ dmabuf = dma_buf_get(db_revoke.dmabuf_fd);
+ if (IS_ERR(dmabuf))
+ return PTR_ERR(dmabuf);
+
+ priv = dmabuf->priv;
+ /*
+ * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_
+ * relates to the VFIO device it was provided with.
+ *
+ * If the DMABUF relates to this vdev then priv->vdev is
+ * stable because this open fd prevents cleanup.
+ *
+ * If it relates to a different vdev, reading priv->vdev might
+ * race with a concurrent cleanup on that device. But if so,
+ * it points to a non-matching vdev or NULL and is unusable
+ * either way.
+ */
+ if (dmabuf->ops != &vfio_pci_dmabuf_ops ||
+ READ_ONCE(priv->vdev) != vdev) {
+ ret = -ENODEV;
+ goto out_put_buf;
+ }
+
+ /*
+ * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()
+ * from getting it and then blocking all devices in the dev_set behind
+ * this revoke's drain.
+ */
+ down_read(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_DEAD) {
+ ret = -EBADFD;
+ } else {
+ vfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);
+ ret = 0;
+ }
+ up_write(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
+
+out_put_buf:
+ dma_buf_put(dmabuf);
+
+ return ret;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 4e7162234a2eb..ca12221af5553 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {
bool test_mem;
};
+enum vfio_pci_dma_buf_status {
+ VFIO_PCI_DMABUF_OK = 0,
+ VFIO_PCI_DMABUF_REVOKED = 1,
+ VFIO_PCI_DMABUF_DEAD = 2,
+};
+
+struct vfio_pci_dma_buf {
+ struct dma_buf *dmabuf;
+ struct vfio_pci_core_device *vdev;
+ struct list_head dmabufs_elm;
+ size_t size;
+ struct phys_vec *phys_vec;
+ struct p2pdma_provider *provider;
+ struct file *vfile;
+ u32 nr_ranges;
+ struct kref kref;
+ struct completion comp;
+ unsigned long vma_pgoff_adjust;
+ enum vfio_pci_dma_buf_status status;
+};
+
bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);
void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);
@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);
int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state);
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
u16 cmd);
@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)
return (pdev->class >> 8) == PCI_CLASS_DISPLAY_VGA;
}
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long address,
+ unsigned int order,
+ unsigned long *out_pfn);
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index);
+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);
+
#ifdef CONFIG_VFIO_PCI_DMABUF
int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
struct vfio_device_feature_dma_buf __user *arg,
size_t argsz);
-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz);
#else
static inline int
vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
return -ENOTTY;
}
-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
-{
-}
-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,
- bool revoked)
+static inline int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
{
+ return -ENOTTY;
}
#endif
diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
index d15b2b31d3c91..952a2c196ad42 100644
--- a/include/linux/dma-buf.h
+++ b/include/linux/dma-buf.h
@@ -571,6 +571,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);
struct dma_buf *dma_buf_get(int fd);
void dma_buf_put(struct dma_buf *dmabuf);
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);
+
struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,
enum dma_data_direction);
void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa2..44891fdb7c76e 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -129,11 +129,13 @@ struct vfio_pci_core_device {
bool disable_idle_d3:1;
bool nointxmask:1;
bool disable_vga:1;
+ bool zap_bars_on_revoke:1;
/* Flags modified at runtime - dedicated storage unit */
bool needs_reset;
bool pm_intx_masked;
bool pm_runtime_engaged;
bool sriov_active;
+ bool bars_revoked;
struct pci_saved_state *pci_saved_state;
struct pci_saved_state *pm_save;
int ioeventfds_nr;
@@ -148,6 +150,7 @@ struct vfio_pci_core_device {
struct vfio_pci_core_device *sriov_pf_core_dev;
struct notifier_block nb;
struct rw_semaphore memory_lock;
+ struct rw_semaphore dmabuf_lock;
struct list_head dmabufs;
};
diff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h
index e41437fa17ad0..d3c6057983e09 100644
--- a/include/uapi/linux/vfio.h
+++ b/include/uapi/linux/vfio.h
@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {
#define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
+/**
+ * Given a DMABUF fd previously exported from the same device by
+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that
+ * access to the corresponding DMABUF is immediately revoked. On
+ * successful return, the buffer is no longer accessible through any
+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future
+ * mmap()s and map/attach requests from any new/existing importer.
+ *
+ * Return: 0 on success, -1 and errno is set on failure:
+ *
+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.
+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.
+ * ENODEV: The DMABUF was not exported from this device.
+ * EBADFD: The DMABUF is already revoked by this feature.
+ *
+ * Additionally, common errors can occur: EFAULT accessing the struct,
+ * or EINVAL requesting an unsupported feature op.
+ */
+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14
+
+struct vfio_device_feature_dma_buf_revoke {
+ __s32 dmabuf_fd;
+};
+
/* -------- API for Type1 VFIO IOMMU -------- */
/**
|
| 3/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24201 output: 18 thoughts: 1086 Thoughts: **Analyzing the Implications of DMA-BUF Integration in VFIO-PCI** Okay, so this patch touches a significant number of files, indicating a pretty substantial shift. My initial thought is to dive into the core of these changes and understand the implications. The files listed are scattered across `dma-buf` and `vfio/pci` areas, so it's likely a cross-module integration effort. First, let's look at `dma-buf.c`. The fact that `dma_buf_set_name` is being exported and modified to work with kernel strings (`char *name`) instead of a user-space buffer via `strndup_user` is interesting. The addition of `DMA_BUF_SET_NAME_A` and `DMA_BUF_SET_NAME_B` in the ioctl handler, with `strndup_user`, suggests a potential hardening attempt, taking ownership of the name, and freeing the old one. We need to assess if this introduces new security considerations, especially around buffer overflows or incorrect string handling. Now, into the `vfio/pci` changes. The `VFIO_PCI_CORE` selecting `DMA_SHARED_BUFFER` is a key indicator. It suggests that DMA buffers are now becoming a core part of how VFIO-PCI operates. The constant inclusion of `vfio_pci_dmabuf.o` and the move to DMABUF-backed BAR mappings in `vfio_pci_core_mmap` are pretty monumental. It points to a complete shift in how devices are exposed and how memory is managed in a VFIO context. The inclusion of `vfio_pci_mmap_huge_fault` to handle page faults on these DMABUF-backed mappings means we're now dealing with DMA buffer memory management. And, introducing the `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl feature indicates a new mechanism for managing device memory revocation within VFIO. All this locking with `memory_lock`, `dmabuf_lock`, and `resv` needs close scrutiny. The potential for race conditions between fault handling, revoke, and cleanup/reset is a concern that demands a thorough review of the locking primitives and their implementation. The new functions introduced, like `vfio_pci_dma_buf_find_pfn`, `vfio_pci_core_mmap_prep_dmabuf`, and `vfio_pci_dma_buf_set_status`, all point to a deeper integration of DMA-BUF within the VFIO-PCI framework. Now, my mind jumps to fuzzing and exploitability. Can this all be fuzzed in standard virtualized environments like QEMU or with syzkaller? It's crucial to determine if `vfio-pci` or `dma-buf` is reachable in these environments. What devices bind to `vfio-pci` in a standard QEMU VM? Is it used, and can syzkaller bind PCI devices to VFIO-PCI? Does syzkaller offer any specific support for VFIO pseudo-syscalls or ioctls? We also have to evaluate the reachability of the `dma-buf` ioctls and how dma-bufs are created. Can userspace access `/dev/dma_heap`, or other `dma-buf` sources? The `dma_buf_ioctl` additions, especially `DMA_BUF_SET_NAME_A` and `DMA_BUF_SET_NAME_B`, are definitely entry points to investigate for potential vulnerabilities. Can syzkaller create a `dma-buf` instance? Yes, via `/dev/udmabuf`, `/dev/dma_heap/*`, or DRM drivers like vgem. We must verify VFIO accessibility using QEMU, and then check syzkaller for VFIO descriptions. I need to run `grepper` or scan the codebase quickly to confirm these suppositions. |
| 4/2 | 2026/09/26 21:06 | tool | grepper |
0mArgs:{
"Expression": "CONFIG_VFIO"
}Results: {
"Output": "Documentation/driver-api/vfio.rst=273=in a VFIO group.\nDocumentation/driver-api/vfio.rst-274-\nDocumentation/driver-api/vfio.rst:275:With CONFIG_VFIO_DEVICE_CDEV=y the user can now acquire a device fd\nDocumentation/driver-api/vfio.rst-276-by directly opening a character device /dev/vfio/devices/vfioX where\n--\nDocumentation/driver-api/vfio.rst=403=Note that the VFIO container case includes IOMMUFD provided VFIO compatibility\nDocumentation/driver-api/vfio.rst:404:interfaces when either CONFIG_VFIO_CONTAINER or CONFIG_IOMMUFD_VFIO_CONTAINER is\nDocumentation/driver-api/vfio.rst-405-enabled.\n--\nDocumentation/driver-api/vfio.rst=410=Kconfig Support Matrix\n--\nDocumentation/driver-api/vfio.rst-412-\nDocumentation/driver-api/vfio.rst:413:The visibility of CONFIG_VFIO_NOIOMMU depends on the combination of\nDocumentation/driver-api/vfio.rst:414:CONFIG_VFIO_GROUP, CONFIG_VFIO_DEVICE_CDEV, and whether a container backend\nDocumentation/driver-api/vfio.rst:415:(CONFIG_VFIO_CONTAINER or CONFIG_IOMMUFD_VFIO_CONTAINER) is configured. The\nDocumentation/driver-api/vfio.rst-416-Kconfig dependencies enforce the following constraints:\n--\nDocumentation/driver-api/vfio.rst=424=The resulting support matrix:\n--\nDocumentation/driver-api/vfio.rst-441-\nDocumentation/driver-api/vfio.rst:442:Container = CONFIG_VFIO_CONTAINER or CONFIG_IOMMUFD_VFIO_CONTAINER (either\nDocumentation/driver-api/vfio.rst-443-suffices). Case 4 is intentionally blocked: allowing NOIOMMU with GROUP\nDocumentation/driver-api/vfio.rst=444=enabled but no container would create unusable group nodes. Users who want\nDocumentation/driver-api/vfio.rst:445:cdev-only noiommu should set CONFIG_VFIO_GROUP=n (case 5).\nDocumentation/driver-api/vfio.rst-446-\nDocumentation/driver-api/vfio.rst-447-.. [#cdev] The cdev noiommu path requires CONFIG_GENERIC_ATOMIC64=n. When\nDocumentation/driver-api/vfio.rst:448: CONFIG_VFIO_GROUP=y, CONFIG_VFIO_DEVICE_CDEV=y, and\nDocumentation/driver-api/vfio.rst:449: CONFIG_GENERIC_ATOMIC64=y, CONFIG_VFIO_NOIOMMU remains selectable for the\nDocumentation/driver-api/vfio.rst-450- group path, but no noiommu device cdev is registered. Cdev-only noiommu is\n--\narch/arm64/configs/defconfig=1411=CONFIG_DMABUF_HEAPS_CMA=m\narch/arm64/configs/defconfig:1412:CONFIG_VFIO=y\narch/arm64/configs/defconfig:1413:CONFIG_VFIO_PCI=y\narch/arm64/configs/defconfig-1414-CONFIG_VIRTIO_PCI=y\n--\narch/loongarch/configs/loongson32_defconfig=901=CONFIG_UIO_PCI_GENERIC=m\narch/loongarch/configs/loongson32_defconfig:902:CONFIG_VFIO=m\narch/loongarch/configs/loongson32_defconfig:903:CONFIG_VFIO_PCI=m\narch/loongarch/configs/loongson32_defconfig-904-CONFIG_VIRTIO_PCI=y\n--\narch/loongarch/configs/loongson64_defconfig=930=CONFIG_UIO_PCI_GENERIC=m\narch/loongarch/configs/loongson64_defconfig:931:CONFIG_VFIO=m\narch/loongarch/configs/loongson64_defconfig:932:CONFIG_VFIO_PCI=m\narch/loongarch/configs/loongson64_defconfig-933-CONFIG_VIRTIO_PCI=y\n--\narch/s390/configs/debug_defconfig=58=CONFIG_CHSC_SCH=y\narch/s390/configs/debug_defconfig:59:CONFIG_VFIO_CCW=m\narch/s390/configs/debug_defconfig:60:CONFIG_VFIO_AP=m\narch/s390/configs/debug_defconfig:61:CONFIG_VFIO_DEVICE_CDEV=y\narch/s390/configs/debug_defconfig-62-CONFIG_IOMMUFD_DRIVER=y\n--\narch/s390/configs/debug_defconfig=635=CONFIG_MLX5_INFINIBAND=m\narch/s390/configs/debug_defconfig:636:CONFIG_VFIO=m\narch/s390/configs/debug_defconfig:637:CONFIG_VFIO_PCI=m\narch/s390/configs/debug_defconfig-638-CONFIG_MLX5_VFIO_PCI=m\n--\narch/s390/configs/defconfig=56=CONFIG_CHSC_SCH=y\narch/s390/configs/defconfig:57:CONFIG_VFIO_CCW=m\narch/s390/configs/defconfig:58:CONFIG_VFIO_AP=m\narch/s390/configs/defconfig:59:CONFIG_VFIO_DEVICE_CDEV=y\narch/s390/configs/defconfig-60-CONFIG_IOMMUFD_DRIVER=y\n--\narch/s390/configs/defconfig=625=CONFIG_MLX5_INFINIBAND=m\narch/s390/configs/defconfig:626:CONFIG_VFIO=m\narch/s390/configs/defconfig:627:CONFIG_VFIO_PCI=m\narch/s390/configs/defconfig-628-CONFIG_MLX5_VFIO_PCI=m\n--\narch/s390/kvm/s390/Makefile=11=kvm-y += $(gmap-y)\narch/s390/kvm/s390/Makefile-12-\narch/s390/kvm/s390/Makefile:13:kvm-$(CONFIG_VFIO_PCI_ZDEV_KVM) += pci.o\narch/s390/kvm/s390/Makefile-14-obj-$(CONFIG_KVM) += kvm.o\n--\narch/s390/kvm/s390/interrupt.c=3661=static void gib_alert_irq_handler(struct airq_struct *airq,\n--\narch/s390/kvm/s390/interrupt.c-3668-\tif ((info-\u003eforward || info-\u003eerror) \u0026\u0026\narch/s390/kvm/s390/interrupt.c:3669:\t IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM)) {\narch/s390/kvm/s390/interrupt.c-3670-\t\taen_process_gait(info-\u003eisc);\n--\narch/s390/kvm/s390/pci.h=48=static inline struct kvm *kvm_s390_pci_si_to_kvm(struct zpci_aift *aift,\n--\narch/s390/kvm/s390/pci.h-50-{\narch/s390/kvm/s390/pci.h:51:\tif (!IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM) || !aift-\u003ekzdev ||\narch/s390/kvm/s390/pci.h-52-\t !aift-\u003ekzdev[si])\n--\narch/s390/kvm/s390/pci.h=68=static inline bool kvm_s390_pci_interp_allowed(void)\n--\narch/s390/kvm/s390/pci.h-82-\tdefault:\narch/s390/kvm/s390/pci.h:83:\t\treturn (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM) \u0026\u0026\narch/s390/kvm/s390/pci.h-84-\t\t\tsclp.has_zpci_lsi \u0026\u0026 sclp.has_aeni \u0026\u0026 sclp.has_aisi \u0026\u0026\n--\narch/s390/kvm/s390/priv.c=541=static int handle_io_inst(struct kvm_vcpu *vcpu)\n--\narch/s390/kvm/s390/priv.c-569-\narch/s390/kvm/s390/priv.c:570:#if IS_ENABLED(CONFIG_VFIO_AP)\narch/s390/kvm/s390/priv.c-571-bool kvm_s390_is_gpa_in_memslot(struct kvm *kvm, gpa_t gpa)\n--\narch/s390/kvm/s390/s390.c=496=static int __init __kvm_s390_init(void)\n--\narch/s390/kvm/s390/s390.c-520-\narch/s390/kvm/s390/s390.c:521:\tif (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM)) {\narch/s390/kvm/s390/s390.c-522-\t\trc = kvm_s390_pci_init();\n--\narch/s390/kvm/s390/s390.c-538-err_gib:\narch/s390/kvm/s390/s390.c:539:\tif (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM))\narch/s390/kvm/s390/s390.c-540-\t\tkvm_s390_pci_exit();\n--\narch/s390/kvm/s390/s390.c=550=static void __kvm_s390_exit(void)\n--\narch/s390/kvm/s390/s390.c-555-\tkvm_s390_gib_destroy();\narch/s390/kvm/s390/s390.c:556:\tif (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM))\narch/s390/kvm/s390/s390.c-557-\t\tkvm_s390_pci_exit();\n--\narch/s390/kvm/s390/s390.c=2900=int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)\n--\narch/s390/kvm/s390/s390.c-3063-\t\tr = -EINVAL;\narch/s390/kvm/s390/s390.c:3064:\t\tif (!IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM))\narch/s390/kvm/s390/s390.c-3065-\t\t\tbreak;\n--\narch/s390/kvm/s390/s390.c=3232=void kvm_arch_free_vm(struct kvm *kvm)\narch/s390/kvm/s390/s390.c-3233-{\narch/s390/kvm/s390/s390.c:3234:\tif (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM))\narch/s390/kvm/s390/s390.c-3235-\t\tkvm_s390_pci_clear_list(kvm);\n--\narch/s390/kvm/s390/s390.c=3240=int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)\n--\narch/s390/kvm/s390/s390.c-3316-\narch/s390/kvm/s390/s390.c:3317:\tif (IS_ENABLED(CONFIG_VFIO_PCI_ZDEV_KVM)) {\narch/s390/kvm/s390/s390.c-3318-\t\tmutex_lock(\u0026kvm-\u003elock);\n--\ndrivers/Makefile=96=obj-$(CONFIG_UIO)\t\t+= uio/\ndrivers/Makefile:97:obj-$(CONFIG_VFIO)\t\t+= vfio/\ndrivers/Makefile-98-obj-y\t\t\t\t+= cdrom/\n--\ndrivers/iommu/iommufd/pages.c=1463=sym_vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,\n--\ndrivers/iommu/iommufd/pages.c-1472-\ndrivers/iommu/iommufd/pages.c:1473:\tif (!IS_ENABLED(CONFIG_VFIO_PCI_DMABUF))\ndrivers/iommu/iommufd/pages.c-1474-\t\treturn -EOPNOTSUPP;\n--\ndrivers/iommu/iommufd/vfio_compat.c=279=static int iommufd_vfio_check_extension(struct iommufd_ctx *ictx,\n--\ndrivers/iommu/iommufd/vfio_compat.c-288-\tcase VFIO_NOIOMMU_IOMMU:\ndrivers/iommu/iommufd/vfio_compat.c:289:\t\treturn IS_ENABLED(CONFIG_VFIO_NOIOMMU);\ndrivers/iommu/iommufd/vfio_compat.c-290-\n--\ndrivers/iommu/iommufd/vfio_compat.c=310=static int iommufd_vfio_set_iommu(struct iommufd_ctx *ictx, unsigned long type)\n--\ndrivers/iommu/iommufd/vfio_compat.c-320-\t */\ndrivers/iommu/iommufd/vfio_compat.c:321:\tif (IS_ENABLED(CONFIG_VFIO_NOIOMMU) \u0026\u0026 type == VFIO_NOIOMMU_IOMMU \u0026\u0026\ndrivers/iommu/iommufd/vfio_compat.c-322-\t no_iommu_mode) {\n--\ndrivers/s390/cio/Makefile=25=vfio_ccw-objs += vfio_ccw_drv.o vfio_ccw_cp.o vfio_ccw_ops.o vfio_ccw_fsm.o \\\ndrivers/s390/cio/Makefile-26-\tvfio_ccw_async.o vfio_ccw_trace.o vfio_ccw_chp.o\ndrivers/s390/cio/Makefile:27:obj-$(CONFIG_VFIO_CCW) += vfio_ccw.o\ndrivers/s390/cio/Makefile-28-\n--\ndrivers/s390/crypto/Makefile=37=vfio_ap-objs := vfio_ap_drv.o vfio_ap_ops.o\ndrivers/s390/crypto/Makefile:38:obj-$(CONFIG_VFIO_AP) += vfio_ap.o\n--\ndrivers/vfio/Makefile-1-# SPDX-License-Identifier: GPL-2.0\ndrivers/vfio/Makefile:2:obj-$(CONFIG_VFIO) += vfio.o\ndrivers/vfio/Makefile-3-\ndrivers/vfio/Makefile=4=vfio-y += vfio_main.o\ndrivers/vfio/Makefile:5:vfio-$(CONFIG_VFIO_DEVICE_CDEV) += device_cdev.o\ndrivers/vfio/Makefile:6:vfio-$(CONFIG_VFIO_GROUP) += group.o\ndrivers/vfio/Makefile-7-vfio-$(CONFIG_IOMMUFD) += iommufd.o\ndrivers/vfio/Makefile:8:vfio-$(CONFIG_VFIO_CONTAINER) += container.o\ndrivers/vfio/Makefile:9:vfio-$(CONFIG_VFIO_VIRQFD) += virqfd.o\ndrivers/vfio/Makefile:10:vfio-$(CONFIG_VFIO_DEBUGFS) += debugfs.o\ndrivers/vfio/Makefile-11-\ndrivers/vfio/Makefile:12:obj-$(CONFIG_VFIO_IOMMU_TYPE1) += vfio_iommu_type1.o\ndrivers/vfio/Makefile:13:obj-$(CONFIG_VFIO_IOMMU_SPAPR_TCE) += vfio_iommu_spapr_tce.o\ndrivers/vfio/Makefile:14:obj-$(CONFIG_VFIO_PCI_CORE) += pci/\ndrivers/vfio/Makefile:15:obj-$(CONFIG_VFIO_PLATFORM_BASE) += platform/\ndrivers/vfio/Makefile:16:obj-$(CONFIG_VFIO_MDEV) += mdev/\ndrivers/vfio/Makefile:17:obj-$(CONFIG_VFIO_FSL_MC) += fsl-mc/\ndrivers/vfio/Makefile:18:obj-$(CONFIG_VFIO_CDX) += cdx/\n--\ndrivers/vfio/cdx/Makefile-5-\ndrivers/vfio/cdx/Makefile:6:obj-$(CONFIG_VFIO_CDX) += vfio-cdx.o\ndrivers/vfio/cdx/Makefile-7-\n--\ndrivers/vfio/container.c=80=static bool vfio_iommu_driver_allowed(struct vfio_container *container,\n--\ndrivers/vfio/container.c-82-{\ndrivers/vfio/container.c:83:\tif (!IS_ENABLED(CONFIG_VFIO_NOIOMMU))\ndrivers/vfio/container.c-84-\t\treturn true;\n--\ndrivers/vfio/container.c=573=int __init vfio_container_init(void)\n--\ndrivers/vfio/container.c-585-\ndrivers/vfio/container.c:586:\tif (IS_ENABLED(CONFIG_VFIO_NOIOMMU)) {\ndrivers/vfio/container.c-587-\t\tret = vfio_register_iommu_driver(\u0026vfio_noiommu_ops);\n--\ndrivers/vfio/container.c=598=void vfio_container_cleanup(void)\ndrivers/vfio/container.c-599-{\ndrivers/vfio/container.c:600:\tif (IS_ENABLED(CONFIG_VFIO_NOIOMMU))\ndrivers/vfio/container.c-601-\t\tvfio_unregister_iommu_driver(\u0026vfio_noiommu_ops);\n--\ndrivers/vfio/fsl-mc/Makefile=3=vfio-fsl-mc-y := vfio_fsl_mc.o vfio_fsl_mc_intr.o\ndrivers/vfio/fsl-mc/Makefile:4:obj-$(CONFIG_VFIO_FSL_MC) += vfio-fsl-mc.o\n--\ndrivers/vfio/group.c=107=static int vfio_group_ioctl_set_container(struct vfio_group *group,\n--\ndrivers/vfio/group.c-139-\tif (!IS_ERR(iommufd)) {\ndrivers/vfio/group.c:140:\t\tif (IS_ENABLED(CONFIG_VFIO_NOIOMMU) \u0026\u0026\ndrivers/vfio/group.c-141-\t\t group-\u003etype == VFIO_NO_IOMMU)\n--\ndrivers/vfio/mdev/Makefile=3=mdev-y := mdev_core.o mdev_sysfs.o mdev_driver.o\ndrivers/vfio/mdev/Makefile-4-\ndrivers/vfio/mdev/Makefile:5:obj-$(CONFIG_VFIO_MDEV) += mdev.o\n--\ndrivers/vfio/pci/Makefile=3=vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o\ndrivers/vfio/pci/Makefile:4:vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o\ndrivers/vfio/pci/Makefile:5:obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o\ndrivers/vfio/pci/Makefile-6-\ndrivers/vfio/pci/Makefile=7=vfio-pci-y := vfio_pci.o\ndrivers/vfio/pci/Makefile:8:vfio-pci-$(CONFIG_VFIO_PCI_IGD) += vfio_pci_igd.o\ndrivers/vfio/pci/Makefile:9:obj-$(CONFIG_VFIO_PCI) += vfio-pci.o\ndrivers/vfio/pci/Makefile-10-\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1635=static void hisi_acc_vfio_debug_init(struct hisi_acc_vf_core_device *hisi_acc_vdev)\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1643-\tif (!debugfs_initialized() ||\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1644:\t !IS_ENABLED(CONFIG_VFIO_DEBUGFS))\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1645-\t\treturn;\n--\ndrivers/vfio/pci/vfio_pci.c=39=MODULE_PARM_DESC(nointxmask,\n--\ndrivers/vfio/pci/vfio_pci.c-41-\ndrivers/vfio/pci/vfio_pci.c:42:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci.c-43-static bool disable_vga;\n--\ndrivers/vfio/pci/vfio_pci.c=128=static int vfio_pci_init_dev(struct vfio_device *core_vdev)\n--\ndrivers/vfio/pci/vfio_pci.c-141-\tvdev-\u003edisable_idle_d3 = disable_idle_d3;\ndrivers/vfio/pci/vfio_pci.c:142:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci.c-143-\tvdev-\u003edisable_vga = disable_vga;\n--\ndrivers/vfio/pci/vfio_pci_config.c=1744=int vfio_config_init(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_config.c-1826-\ndrivers/vfio/pci/vfio_pci_config.c:1827:\tif (!IS_ENABLED(CONFIG_VFIO_PCI_INTX) || vdev-\u003enointx ||\ndrivers/vfio/pci/vfio_pci_config.c-1828-\t !vdev-\u003epdev-\u003eirq || vdev-\u003epdev-\u003eirq == IRQ_NOTCONNECTED)\n--\ndrivers/vfio/pci/vfio_pci_core.c=96=static inline bool vfio_vga_disabled(struct vfio_pci_core_device *vdev)\ndrivers/vfio/pci/vfio_pci_core.c-97-{\ndrivers/vfio/pci/vfio_pci_core.c:98:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_core.c-99-\treturn vdev-\u003edisable_vga;\n--\ndrivers/vfio/pci/vfio_pci_core.c-104-\ndrivers/vfio/pci/vfio_pci_core.c:105:#ifdef CONFIG_VFIO_DEBUGFS\ndrivers/vfio/pci/vfio_pci_core.c-106-static struct vfio_pci_core_device *\n--\ndrivers/vfio/pci/vfio_pci_core.c=154=static inline void vfio_pci_core_debugfs_init(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_core.c-156-}\ndrivers/vfio/pci/vfio_pci_core.c:157:#endif /* CONFIG_VFIO_DEBUGFS */\ndrivers/vfio/pci/vfio_pci_core.c-158-\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=11=MODULE_IMPORT_NS(\"DMA_BUF\");\ndrivers/vfio/pci/vfio_pci_dmabuf.c-12-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:13:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-14-static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=75=static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-84-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:85:#endif /* CONFIG_VFIO_PCI_DMABUF */\ndrivers/vfio/pci/vfio_pci_dmabuf.c-86-\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=149=static const struct dma_buf_ops vfio_pci_dmabuf_ops = {\ndrivers/vfio/pci/vfio_pci_dmabuf.c-150-\t.attach = vfio_pci_dma_buf_attach,\ndrivers/vfio/pci/vfio_pci_dmabuf.c:151:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-152-\t.mmap = vfio_pci_dma_buf_mmap,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=377=int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-433-\tpriv-\u003eprovider = pcim_p2pdma_provider(vdev-\u003epdev, res_index);\ndrivers/vfio/pci/vfio_pci_dmabuf.c:434:\tif (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) \u0026\u0026 !priv-\u003eprovider) {\ndrivers/vfio/pci/vfio_pci_dmabuf.c-435-\t\tret = -EINVAL;\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-476-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:477:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-478-/*\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=590=int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-675-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:676:#endif /* CONFIG_VFIO_PCI_DMABUF */\ndrivers/vfio/pci/vfio_pci_dmabuf.c-677-\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=760=void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-789-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:790:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-791-int vfio_pci_core_feature_dma_buf_revoke(\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-856-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:857:#endif /* CONFIG_VFIO_PCI_DMABUF */\n--\ndrivers/vfio/pci/vfio_pci_priv.h=65=ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_priv.h-67-\ndrivers/vfio/pci/vfio_pci_priv.h:68:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_priv.h-69-ssize_t vfio_pci_vga_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_priv.h=95=void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_priv.h-97-\ndrivers/vfio/pci/vfio_pci_priv.h:98:#ifdef CONFIG_VFIO_PCI_IGD\ndrivers/vfio/pci/vfio_pci_priv.h-99-bool vfio_pci_is_intel_display(struct pci_dev *pdev);\n--\ndrivers/vfio/pci/vfio_pci_priv.h=107=static inline int vfio_pci_igd_init(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_priv.h-112-\ndrivers/vfio/pci/vfio_pci_priv.h:113:#ifdef CONFIG_VFIO_PCI_ZDEV_KVM\ndrivers/vfio/pci/vfio_pci_priv.h-114-int vfio_pci_info_zdev_add_caps(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_priv.h=160=void vfio_pci_set_vma_ops(struct vm_area_struct *vma);\ndrivers/vfio/pci/vfio_pci_priv.h-161-\ndrivers/vfio/pci/vfio_pci_priv.h:162:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_priv.h-163-int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_rdwr.c=201=ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_rdwr.c-280-\ndrivers/vfio/pci/vfio_pci_rdwr.c:281:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_rdwr.c-282-ssize_t vfio_pci_vga_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/platform/Makefile=2=vfio-platform-base-y := vfio_platform_common.o vfio_platform_irq.o\ndrivers/vfio/platform/Makefile:3:obj-$(CONFIG_VFIO_PLATFORM_BASE) += vfio-platform-base.o\ndrivers/vfio/platform/Makefile:4:obj-$(CONFIG_VFIO_PLATFORM_BASE) += reset/\ndrivers/vfio/platform/Makefile-5-\ndrivers/vfio/platform/Makefile=6=vfio-platform-y := vfio_platform.o\ndrivers/vfio/platform/Makefile:7:obj-$(CONFIG_VFIO_PLATFORM) += vfio-platform.o\ndrivers/vfio/platform/Makefile-8-\ndrivers/vfio/platform/Makefile=9=vfio-amba-y := vfio_amba.o\ndrivers/vfio/platform/Makefile:10:obj-$(CONFIG_VFIO_AMBA) += vfio-amba.o\n--\ndrivers/vfio/platform/reset/Makefile=3=vfio-platform-amdxgbe-y := vfio_platform_amdxgbe.o\ndrivers/vfio/platform/reset/Makefile-4-\ndrivers/vfio/platform/reset/Makefile:5:obj-$(CONFIG_VFIO_PLATFORM_CALXEDAXGMAC_RESET) += vfio-platform-calxedaxgmac.o\ndrivers/vfio/platform/reset/Makefile:6:obj-$(CONFIG_VFIO_PLATFORM_AMDXGBE_RESET) += vfio-platform-amdxgbe.o\ndrivers/vfio/platform/reset/Makefile:7:obj-$(CONFIG_VFIO_PLATFORM_BCMFLEXRM_RESET) += vfio_platform_bcmflexrm.o\n--\ndrivers/vfio/vfio.h=37=extern const struct file_operations vfio_device_fops;\ndrivers/vfio/vfio.h-38-\ndrivers/vfio/vfio.h:39:#ifdef CONFIG_VFIO_NOIOMMU\ndrivers/vfio/vfio.h-40-extern bool vfio_noiommu __read_mostly;\n--\ndrivers/vfio/vfio.h=45=enum vfio_group_type {\n--\ndrivers/vfio/vfio.h-68-\ndrivers/vfio/vfio.h:69:#if IS_ENABLED(CONFIG_VFIO_GROUP)\ndrivers/vfio/vfio.h-70-struct vfio_group {\n--\ndrivers/vfio/vfio.h-85-\tstruct list_head\t\tvfio_next;\ndrivers/vfio/vfio.h:86:#if IS_ENABLED(CONFIG_VFIO_CONTAINER)\ndrivers/vfio/vfio.h-87-\tstruct list_head\t\tcontainer_next;\n--\ndrivers/vfio/vfio.h=182=static inline void vfio_group_cleanup(void)\n--\ndrivers/vfio/vfio.h-185-\ndrivers/vfio/vfio.h:186:#endif /* CONFIG_VFIO_GROUP */\ndrivers/vfio/vfio.h-187-\ndrivers/vfio/vfio.h=188=static inline bool vfio_device_is_noiommu(struct vfio_device *vdev)\ndrivers/vfio/vfio.h-189-{\ndrivers/vfio/vfio.h:190:#if IS_ENABLED(CONFIG_VFIO_GROUP)\ndrivers/vfio/vfio.h-191-\tif (vdev-\u003egroup \u0026\u0026 vdev-\u003egroup-\u003etype == VFIO_NO_IOMMU)\n--\ndrivers/vfio/vfio.h-197-\ndrivers/vfio/vfio.h:198:#if IS_ENABLED(CONFIG_VFIO_CONTAINER)\ndrivers/vfio/vfio.h-199-/**\n--\ndrivers/vfio/vfio.h=354=int vfio_df_ioctl_detach_pt(struct vfio_device_file *df,\n--\ndrivers/vfio/vfio.h-356-\ndrivers/vfio/vfio.h:357:#if IS_ENABLED(CONFIG_VFIO_DEVICE_CDEV)\ndrivers/vfio/vfio.h-358-void vfio_init_device_cdev(struct vfio_device *device);\n--\ndrivers/vfio/vfio.h=413=static inline void vfio_cdev_cleanup(void)\n--\ndrivers/vfio/vfio.h-415-}\ndrivers/vfio/vfio.h:416:#endif /* CONFIG_VFIO_DEVICE_CDEV */\ndrivers/vfio/vfio.h-417-\ndrivers/vfio/vfio.h:418:#if IS_ENABLED(CONFIG_VFIO_VIRQFD)\ndrivers/vfio/vfio.h-419-int __init vfio_virqfd_init(void);\n--\ndrivers/vfio/vfio.h=440=static inline void vfio_device_put_kvm(struct vfio_device *device)\n--\ndrivers/vfio/vfio.h-444-\ndrivers/vfio/vfio.h:445:#ifdef CONFIG_VFIO_DEBUGFS\ndrivers/vfio/vfio.h-446-void vfio_debugfs_create_root(void);\n--\ndrivers/vfio/vfio.h=456=static inline void vfio_device_debugfs_exit(struct vfio_device *vdev) { }\ndrivers/vfio/vfio.h:457:#endif /* CONFIG_VFIO_DEBUGFS */\ndrivers/vfio/vfio.h-458-\n--\ndrivers/vfio/vfio_main.c=51=static struct vfio {\n--\ndrivers/vfio/vfio_main.c-56-\ndrivers/vfio/vfio_main.c:57:#ifdef CONFIG_VFIO_NOIOMMU\ndrivers/vfio/vfio_main.c-58-bool vfio_noiommu __read_mostly;\n--\ndrivers/vfio/vfio_main.c=1369=static long vfio_device_fops_unl_ioctl(struct file *filep,\n--\ndrivers/vfio/vfio_main.c-1388-\t/* cdev only ioctls */\ndrivers/vfio/vfio_main.c:1389:\tif (IS_ENABLED(CONFIG_VFIO_DEVICE_CDEV) \u0026\u0026 !df-\u003egroup) {\ndrivers/vfio/vfio_main.c-1390-\t\tswitch (cmd) {\n--\ninclude/linux/sched/user.h=14=struct user_struct {\n--\ninclude/linux/sched/user.h-27-\tdefined(CONFIG_NET) || defined(CONFIG_IO_URING) || \\\ninclude/linux/sched/user.h:28:\tdefined(CONFIG_VFIO_PCI_ZDEV_KVM) || IS_ENABLED(CONFIG_IOMMUFD) || \\\ninclude/linux/sched/user.h-29-\tdefined(CONFIG_SECRETMEM)\n--\ninclude/linux/vfio.h=39=struct vfio_device {\n--\ninclude/linux/vfio.h-47-\tconst struct vfio_log_ops *log_ops;\ninclude/linux/vfio.h:48:#if IS_ENABLED(CONFIG_VFIO_GROUP)\ninclude/linux/vfio.h-49-\tstruct vfio_group *group;\n--\ninclude/linux/vfio.h-61-\tstruct device device;\t/* device.kref covers object life circle */\ninclude/linux/vfio.h:62:#if IS_ENABLED(CONFIG_VFIO_DEVICE_CDEV)\ninclude/linux/vfio.h-63-\tstruct cdev cdev;\n--\ninclude/linux/vfio.h=363=struct iommu_group *vfio_file_iommu_group(struct file *file);\ninclude/linux/vfio.h-364-\ninclude/linux/vfio.h:365:#if IS_ENABLED(CONFIG_VFIO_GROUP)\ninclude/linux/vfio.h-366-bool vfio_file_is_group(struct file *file);\n--\ninclude/linux/vfio_pci_core.h=60=struct vfio_pci_device_ops {\n--\ninclude/linux/vfio_pci_core.h-68-\ninclude/linux/vfio_pci_core.h:69:#if IS_ENABLED(CONFIG_VFIO_PCI_DMABUF)\ninclude/linux/vfio_pci_core.h-70-int vfio_pci_core_fill_phys_vec(struct phys_vec *phys_vec,\n"
} |
| 5/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36814 output: 21 thoughts: 1218 Thoughts: **Analyzing DMA-BUF and VFIO Context** Okay, let's break this down. First, I see references to `VFIO_PCI_CORE`, `VFIO_PCI`, and `DMA_BUF`. My initial thought is to consider whether VFIO is even accessible within syzkaller's QEMU environment. I need to quickly determine if there are any relevant pseudo-syscalls or configurations related to VFIO, or if VFIO devices are present at all. However, I'm reminded of a crucial limitation: I *cannot* use this tool to search syzkaller's syzlang descriptions or expressions containing "long syz_," because those pseudo-syscalls aren't in the kernel. Instead, I need to use `read-syz-spec` and `syz-grepper` for those. So, let's make sure I can even use those, and if I do need to search for those, since this tool is for Linux kernel files, POSIX headers, and sysfs/procfs paths. Alright, I see that I have `default_api:grepper`, `default_api:read-file`, `default_api:codesearch-dir-index`, and `default_api:set-results` available. It seems that `syz-grepper` and `read-syz-spec` are *not* available in this current environment, so I won't be able to use them anyway. Now, let's shift our focus to the patch itself and what it's changing. I'll examine the affected files. The patch modifies `drivers/dma-buf/dma-buf.c`, specifically the `dma_buf_set_name` function. Before the patch, it used `strndup_user` within `dma_buf_set_name`, which is now an exported symbol. The patch replaces `strndup_user` with a version that uses `strlen` to check the size, and now also returns an error if the name length exceeds `DMA_BUF_NAME_LEN`. Looking at `dma_buf_ioctl`, the patch seems to have an effect. The code used to use `strndup_user` with DMA_BUF_NAME_LEN as the size. Now, if the `name` is valid, `dma_buf_set_name` is called, which performs the length check. Now, I need to understand the attack surface. How does userspace actually *get* a `dma_buf` file descriptor? I need to figure out which kernel modules export dma_buf objects or create file descriptors for them. I remember `/dev/dma_heap/` and `/dev/udmabuf` (CONFIG_DMABUF_HEAPS, system-heap, CONFIG_UDMABUF). Then there are DRM ioctls like `DRM_IOCTL_PRIME_HANDLE_TO_FD` which export GEM objects as dma-buf fds, like vgem, virtio-gpu, dummy, and vkms. So, userspace can get a `dma_buf` file descriptor and then call an ioctl such as `ioctl(dmabuf_fd, DMA_BUF_SET_NAME_A, name)`. Now, the important bit: inside the kernel `strndup_user` is called and copies up to `DMA_BUF_NAME_LEN` (which is 32 bytes). If the userspace string is 32 bytes or longer *without* a null terminator within the first 31 bytes, then `strndup_user(buf, 32)` will return a 32-byte string (31 characters + null terminator). However, the new patch contains a check, so the code should return -E2BIG, but I need to double-check `DMA_BUF_NAME_LEN`. |
| 6/2 | 2026/09/26 21:06 | tool | grepper |
0mArgs:{
"Expression": "DMA_BUF_NAME_LEN"
}Results: {
"Output": "drivers/dma-buf/dma-buf.c=150=static char *dmabuffs_dname(struct dentry *dentry, char *buffer, int buflen)\n--\ndrivers/dma-buf/dma-buf.c-152-\tstruct dma_buf *dmabuf;\ndrivers/dma-buf/dma-buf.c:153:\tchar name[DMA_BUF_NAME_LEN];\ndrivers/dma-buf/dma-buf.c-154-\tssize_t ret = 0;\n--\ndrivers/dma-buf/dma-buf.c=512=static long dma_buf_ioctl(struct file *file,\n--\ndrivers/dma-buf/dma-buf.c-553-\t\tchar *name = strndup_user((const char __user *)arg,\ndrivers/dma-buf/dma-buf.c:554:\t\t\t\t\t DMA_BUF_NAME_LEN);\ndrivers/dma-buf/dma-buf.c-555-\n--\ndrivers/dma-buf/dma-buf.c=858=EXPORT_SYMBOL_NS_GPL(dma_buf_put, \"DMA_BUF\");\n--\ndrivers/dma-buf/dma-buf.c-872- * Returns 0 on success, -EINVAL if the name is NULL, or -E2BIG if the\ndrivers/dma-buf/dma-buf.c:873: * name (including terminator) exceeds DMA_BUF_NAME_LEN.\ndrivers/dma-buf/dma-buf.c-874- */\ndrivers/dma-buf/dma-buf.c=875=int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\n--\ndrivers/dma-buf/dma-buf.c-880-\t/* dmabuffs_dname() won't use the string if the length\ndrivers/dma-buf/dma-buf.c:881:\t * (including terminator) exceeds DMA_BUF_NAME_LEN:\ndrivers/dma-buf/dma-buf.c-882-\t */\ndrivers/dma-buf/dma-buf.c:883:\tif (strlen(name) \u003e= DMA_BUF_NAME_LEN)\ndrivers/dma-buf/dma-buf.c-884-\t\treturn -E2BIG;\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=377=int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-398-\t * Debug name: The absolute maximum size of the name\ndrivers/vfio/pci/vfio_pci_dmabuf.c:399:\t * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.\ndrivers/vfio/pci/vfio_pci_dmabuf.c-400-\t */\n--\ninclude/linux/dma-buf.h=294=struct dma_buf {\n--\ninclude/linux/dma-buf.h-345-\t * Userspace-provided name. Default value is NULL. If not NULL,\ninclude/linux/dma-buf.h:346:\t * length cannot be longer than DMA_BUF_NAME_LEN, including NIL\ninclude/linux/dma-buf.h-347-\t * char. Useful for accounting and debugging. Read/Write accesses\n--\ninclude/uapi/linux/dma-buf.h=55=struct dma_buf_sync {\n--\ninclude/uapi/linux/dma-buf.h-86-\ninclude/uapi/linux/dma-buf.h:87:#define DMA_BUF_NAME_LEN\t32\ninclude/uapi/linux/dma-buf.h-88-\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=22=static int udmabuf = -1;\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:23:static const char udmabuf_test_buffer_name[DMA_BUF_NAME_LEN] = \"udmabuf_test_buffer_for_iter\";\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-24-static size_t udmabuf_test_buffer_size;\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=25=static int sysheap_dmabuf = -1;\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:26:static const char sysheap_test_buffer_name[DMA_BUF_NAME_LEN] = \"sysheap_test_buffer_for_iter\";\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-27-static size_t sysheap_test_buffer_size;\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=29=static int create_udmabuf(void)\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-35-\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:36:\tif (!ASSERT_LE(sizeof(udmabuf_test_buffer_name), DMA_BUF_NAME_LEN, \"NAMETOOLONG\"))\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-37-\t\treturn -1;\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=76=static int create_sys_heap_dmabuf(size_t bytes)\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-85-\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:86:\tif (!ASSERT_LE(sizeof(sysheap_test_buffer_name), DMA_BUF_NAME_LEN, \"NAMETOOLONG\"))\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-87-\t\treturn -1;\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=131=struct DmabufInfo {\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-133-\tunsigned long size;\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:134:\tchar name[DMA_BUF_NAME_LEN];\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-135-\tchar exporter[32];\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c=242=static void subtest_dmabuf_iter_check_open_coded(struct dmabuf_iter *skel, int map_fd)\n--\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-244-\tLIBBPF_OPTS(bpf_test_run_opts, topts);\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c:245:\tchar key[DMA_BUF_NAME_LEN];\ntools/testing/selftests/bpf/prog_tests/dmabuf_iter.c-246-\tint err, fd;\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-7-/* From uapi/linux/dma-buf.h */\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:8:#define DMA_BUF_NAME_LEN 32\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-9-\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c=12=struct {\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-13-\t__uint(type, BPF_MAP_TYPE_HASH);\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:14:\t__uint(key_size, DMA_BUF_NAME_LEN);\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-15-\t__type(value, bool);\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c=31=int dmabuf_collector(struct bpf_iter__dmabuf *ctx)\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-37-\tconst char *pname, *exporter;\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:38:\tchar name[DMA_BUF_NAME_LEN] = {'\\0'};\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-39-\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c=63=int iter_dmabuf_for_each(const void *ctx)\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-67-\tbpf_for_each(dmabuf, d) {\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:68:\t\tchar name[DMA_BUF_NAME_LEN];\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-69-\t\tconst char *pname;\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-88-\t\t */\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:89:\t\tbpf_for(i, len, DMA_BUF_NAME_LEN)\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-90-\t\t\tname[i] = 0;\n"
} |
| 7/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41510 output: 21 thoughts: 414 Thoughts: **Analyzing DMA-BUF Name Setting** Okay, let's dive into this `dma_buf_set_name` function in `drivers/dma-buf/dma-buf.c`. My initial thought is that it's straightforward: It takes a `dma_buf` and a name, does some basic error checking for a NULL name, checks that the name length is within the defined `DMA_BUF_NAME_LEN` limit, and then, after acquiring a spinlock for thread safety, it frees any existing name and assigns the new name. Makes sense. But then, I need to check how it’s *called*. Okay, here's the `DMA_BUF_SET_NAME_A` ioctl handler. Ah, here's where it gets interesting. I see a `strndup_user` being used. That's good, because `strndup_user` is a user-space safe string duplication function, preventing potential buffer overflows. That's critical for security. It's used here with `DMA_BUF_NAME_LEN` as the length parameter. Hang on... I'm just quickly confirming the behavior of `strndup_user`. It duplicates at most `n` bytes *including* the null terminator. That's a crucial detail. I need to make sure I understand that interaction fully. I just want to be certain that behavior works correctly with the limit set by `DMA_BUF_NAME_LEN` and how that interacts with the overall design. So it's copying at most `DMA_BUF_NAME_LEN` characters, so if the user sends in a string that's the length of `DMA_BUF_NAME_LEN` it'll also be storing the null character. That seems like a good design overall. |
| 8/2 | 2026/09/26 21:06 | tool | grepper |
0mArgs:{
"Expression": "char \\*strndup_user"
}Results: {
"Output": "include/linux/string.h-16-\ninclude/linux/string.h:17:extern char *strndup_user(const char __user *, long);\ninclude/linux/string.h-18-extern void *memdup_user(const void __user *, size_t) __realloc_size(2);\n--\nmm/util.c=258=EXPORT_SYMBOL(vmemdup_user);\n--\nmm/util.c-266- */\nmm/util.c:267:char *strndup_user(const char __user *s, long n)\nmm/util.c-268-{\n"
} |
| 9/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 42494 output: 33 thoughts: 15 Thoughts: Okay, here's my interpretation of what's going on, formatted as you requested: **Deconstructing `mm/util.c:267`** Alright, let's dive into `mm/util.c:267`. It's time to examine this specific line of code. My expectation is that this line holds some key logic, probably a core component of memory management or related functionality, given the file path. Let's see what gems we can find here. I'm keen to understand its purpose within the broader context of the system's memory management strategy. |
| 10/2 | 2026/09/26 21:06 | tool | read-file |
0mArgs:{
"File": "mm/util.c",
"FirstLine": 265,
"LineCount": 35
}Results: {
"Contents": " 265:\t * Return: newly allocated copy of @s or an ERR_PTR() in case of error\n 266:\t */\n 267:\tchar *strndup_user(const char __user *s, long n)\n 268:\t{\n 269:\t\tchar *p;\n 270:\t\tlong length;\n 271:\t\n 272:\t\tlength = strnlen_user(s, n);\n 273:\t\n 274:\t\tif (!length)\n 275:\t\t\treturn ERR_PTR(-EFAULT);\n 276:\t\n 277:\t\tif (length \u003e n)\n 278:\t\t\treturn ERR_PTR(-EINVAL);\n 279:\t\n 280:\t\tp = memdup_user(s, length);\n 281:\t\n 282:\t\tif (IS_ERR(p))\n 283:\t\t\treturn p;\n 284:\t\n 285:\t\tp[length - 1] = '\\0';\n 286:\t\n 287:\t\treturn p;\n 288:\t}\n 289:\tEXPORT_SYMBOL(strndup_user);\n 290:\t\n 291:\t/**\n 292:\t * memdup_user_nul - duplicate memory region from user space and NUL-terminate\n 293:\t *\n 294:\t * @src: source address in user space\n 295:\t * @len: number of bytes to copy\n 296:\t *\n 297:\t * Return: an ERR_PTR() on failure.\n 298:\t */\n 299:\tvoid *memdup_user_nul(const void __user *src, size_t len)\n"
} |
| 11/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 43130 output: 46 thoughts: 301 Thoughts: **Analyzing String Length in DMA Buffer Naming** Okay, so I'm looking at this `strnlen_user(s, n)` function, and it's key to how DMA buffer names are handled. My thinking is, if the string's length is less than or equal to `n` (and includes the null terminator), it returns that length, which means the length can be up to `n`, in this case, `DMA_BUF_NAME_LEN`, which is 32. Then, the code sets `p[length - 1] = '\0'`. So, if I'm understanding this correctly, the effective length seen by `strlen(p)` can be up to `length - 1`, which means a maximum of 31 characters. Wait, if `length` was 32, then `strlen(p)` would be 31. So, `strlen(p) >= 32` would always be false in this scenario, right? But hold on... what if someone *does* try to call `dma_buf_set_name` with a string exactly 32 characters long? Oh, that's handled, because it returns `-E2BIG` if `strlen(name) >= DMA_BUF_NAME_LEN`. Good. But now I'm thinking about the bigger picture. I need to consider all the places this logic impacts, for example VFIO PCI. How does `vfio_pci_dmabuf.c` use this? What's going on in `vfio_pci_core_mmap`? Time to dig into the `vfio_pci_core.c` file around lines 1800-1980 and also check `vfio_pci_dmabuf.c` closely to ensure I'm getting the complete picture and not missing anything. |
| 12/2 | 2026/09/26 21:06 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_core.c",
"FirstLine": 1920,
"LineCount": 80
}Results: {
"Contents": "1920:\tvoid vfio_pci_set_vma_ops(struct vm_area_struct *vma)\n1921:\t{\n1922:\t\tvma-\u003evm_ops = \u0026vfio_pci_mmap_ops;\n1923:\t}\n1924:\t\n1925:\tint vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)\n1926:\t{\n1927:\t\tstruct vfio_pci_core_device *vdev =\n1928:\t\t\tcontainer_of(core_vdev, struct vfio_pci_core_device, vdev);\n1929:\t\tstruct pci_dev *pdev = vdev-\u003epdev;\n1930:\t\tunsigned int index;\n1931:\t\tu64 phys_len, req_len, pgoff, req_start;\n1932:\t\tvoid __iomem *bar_io;\n1933:\t\tint ret;\n1934:\t\n1935:\t\tindex = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n1936:\t\n1937:\t\tif (index \u003e= VFIO_PCI_NUM_REGIONS + vdev-\u003enum_regions)\n1938:\t\t\treturn -EINVAL;\n1939:\t\tif (vma-\u003evm_end \u003c vma-\u003evm_start)\n1940:\t\t\treturn -EINVAL;\n1941:\t\tif ((vma-\u003evm_flags \u0026 VM_SHARED) == 0)\n1942:\t\t\treturn -EINVAL;\n1943:\t\tif (index \u003e= VFIO_PCI_NUM_REGIONS) {\n1944:\t\t\tint regnum = index - VFIO_PCI_NUM_REGIONS;\n1945:\t\t\tstruct vfio_pci_region *region = vdev-\u003eregion + regnum;\n1946:\t\n1947:\t\t\tif (region-\u003eops \u0026\u0026 region-\u003eops-\u003emmap \u0026\u0026\n1948:\t\t\t (region-\u003eflags \u0026 VFIO_REGION_INFO_FLAG_MMAP))\n1949:\t\t\t\treturn region-\u003eops-\u003emmap(vdev, region, vma);\n1950:\t\t\treturn -EINVAL;\n1951:\t\t}\n1952:\t\tif (index \u003e= VFIO_PCI_ROM_REGION_INDEX)\n1953:\t\t\treturn -EINVAL;\n1954:\t\tif (!vdev-\u003ebar_mmap_supported[index])\n1955:\t\t\treturn -EINVAL;\n1956:\t\n1957:\t\tphys_len = PAGE_ALIGN(pci_resource_len(pdev, index));\n1958:\t\treq_len = vma-\u003evm_end - vma-\u003evm_start;\n1959:\t\tpgoff = vma-\u003evm_pgoff \u0026\n1960:\t\t\t((1U \u003c\u003c (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);\n1961:\t\treq_start = pgoff \u003c\u003c PAGE_SHIFT;\n1962:\t\n1963:\t\tif (req_start + req_len \u003e phys_len)\n1964:\t\t\treturn -EINVAL;\n1965:\t\n1966:\t\t/*\n1967:\t\t * Ensure the BAR resource region is reserved for use.\n1968:\t\t */\n1969:\t\tbar_io = vfio_pci_core_get_iomap(vdev, index);\n1970:\t\tif (IS_ERR(bar_io))\n1971:\t\t\treturn PTR_ERR(bar_io);\n1972:\t\n1973:\t\tret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,\n1974:\t\t\t\t\t\t pci_resource_start(pdev, index),\n1975:\t\t\t\t\t\t req_len, index);\n1976:\t\tif (ret)\n1977:\t\t\treturn ret;\n1978:\t\n1979:\t\tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n1980:\t\tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n1981:\t\n1982:\t\t/*\n1983:\t\t * Set vm_flags now, they should not be changed in the fault handler.\n1984:\t\t * We want the same flags and page protection (decrypted above) as\n1985:\t\t * io_remap_pfn_range() would set.\n1986:\t\t *\n1987:\t\t * VM_ALLOW_ANY_UNCACHED: The VMA flag is implemented for ARM64,\n1988:\t\t * allowing KVM stage 2 device mapping attributes to use Normal-NC\n1989:\t\t * rather than DEVICE_nGnRE, which allows guest mappings\n1990:\t\t * supporting write-combining attributes (WC). ARM does not\n1991:\t\t * architecturally guarantee this is safe, and indeed some MMIO\n1992:\t\t * regions like the GICv2 VCPU interface can trigger uncontained\n1993:\t\t * faults if Normal-NC is used.\n1994:\t\t *\n1995:\t\t * To safely use VFIO in KVM the platform must guarantee full\n1996:\t\t * safety in the guest where no action taken against a MMIO\n1997:\t\t * mapping can trigger an uncontained failure. The assumption is\n1998:\t\t * that most VFIO PCI platforms support this for both mapping types,\n1999:\t\t * at least in common flows, based on some expectations of how\n"
} |
| 13/2 | 2026/09/26 21:06 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45335 output: 47 thoughts: 163 Thoughts: Okay, let me break this down. **Diving into the VFIO PCI Memory Mapping Changes** Alright, I'm taking a look at this `vfio_pci_core_mmap` function, and it's clear there's been a shift in how things work. Before this patch, the memory mapping setup relied on directly assigning `vma->vm_private_data = vdev;` and using `vma_to_pfn` within `vfio_pci_mmap_huge_fault`, which is fairly standard in VFIO. Now, the situation's different. The critical change is this call to `vfio_pci_core_mmap_prep_dmabuf` within `vfio_pci_core_mmap`. That's where the interesting logic must be. Clearly, this `prep_dmabuf` function is introducing a different, possibly more sophisticated, method for handling DMA buffer preparation during the mmap operation. So, I need to understand what `vfio_pci_core_mmap_prep_dmabuf` does, especially in terms of its implications for DMA buffer management. I'll need to jump over to `drivers/vfio/pci/vfio_pci_dmabuf.c` to see how it works. That file should contain the key details. Let's see... |
| 14/2 | 2026/09/26 21:07 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_dmabuf.c",
"FirstLine": 375,
"LineCount": 95
}Results: {
"Contents": " 375:\t}\n 376:\t\n 377:\tint vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n 378:\t\t\t\t\t struct vm_area_struct *vma,\n 379:\t\t\t\t\t u64 phys_start, u64 req_len,\n 380:\t\t\t\t\t unsigned int res_index)\n 381:\t{\n 382:\t\tstruct vfio_pci_dma_buf *priv;\n 383:\t\tunsigned long vma_pgoff = vma-\u003evm_pgoff \u0026 (VFIO_PCI_OFFSET_MASK \u003e\u003e PAGE_SHIFT);\n 384:\t\tchar *bufname;\n 385:\t\tint ret;\n 386:\t\n 387:\t\tpriv = kzalloc_obj(*priv);\n 388:\t\tif (!priv)\n 389:\t\t\treturn -ENOMEM;\n 390:\t\n 391:\t\tpriv-\u003ephys_vec = kzalloc_obj(*priv-\u003ephys_vec);\n 392:\t\tif (!priv-\u003ephys_vec) {\n 393:\t\t\tret = -ENOMEM;\n 394:\t\t\tgoto err_free_priv;\n 395:\t\t}\n 396:\t\n 397:\t\t/*\n 398:\t\t * Debug name: The absolute maximum size of the name\n 399:\t\t * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.\n 400:\t\t */\n 401:\t\tbufname = kasprintf(GFP_KERNEL, \"vfio:%s/%x\",\n 402:\t\t\t\t pci_name(vdev-\u003epdev),\n 403:\t\t\t\t res_index);\n 404:\t\n 405:\t\tif (!bufname) {\n 406:\t\t\tret = -ENOMEM;\n 407:\t\t\tgoto err_free_phys;\n 408:\t\t}\n 409:\t\n 410:\t\t/*\n 411:\t\t * The DMABUF begins from the mmap()'s BAR offset, i.e. the\n 412:\t\t * start of the VMA corresponds to byte 0 of the DMABUF and\n 413:\t\t * byte (vma_pgoff \u003c\u003c PAGE_SHIFT) of the BAR.\n 414:\t\t *\n 415:\t\t * vfio_pci_dma_buf_find_pfn() reverses this offset using\n 416:\t\t * vma_pgoff_adjust, so that ultimately a fault's offset from\n 417:\t\t * the start of the _VMA_ has a consistent usage whether the\n 418:\t\t * VMA originates from an mmap() of the VFIO device here or a\n 419:\t\t * direct DMABUF mmap(). Note vma_pgoff_adjust also includes\n 420:\t\t * the encoded VFIO region index, which cancels out the index\n 421:\t\t * encoded in vm_pgoff.\n 422:\t\t */\n 423:\t\tpriv-\u003evdev = vdev;\n 424:\t\tpriv-\u003esize = req_len;\n 425:\t\tpriv-\u003enr_ranges = 1;\n 426:\t\tpriv-\u003evma_pgoff_adjust = vma-\u003evm_pgoff;\n 427:\t\n 428:\t\t/*\n 429:\t\t * The provider can be NULL _iff_ the DMABUF feature isn't\n 430:\t\t * supported, because it's only used by DMABUF import and\n 431:\t\t * attach is prohibited if the feature isn't present.\n 432:\t\t */\n 433:\t\tpriv-\u003eprovider = pcim_p2pdma_provider(vdev-\u003epdev, res_index);\n 434:\t\tif (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) \u0026\u0026 !priv-\u003eprovider) {\n 435:\t\t\tret = -EINVAL;\n 436:\t\t\tgoto err_free_name;\n 437:\t\t}\n 438:\t\n 439:\t\tpriv-\u003ephys_vec[0].paddr = phys_start + ((u64)vma_pgoff \u003c\u003c PAGE_SHIFT);\n 440:\t\tpriv-\u003ephys_vec[0].len = priv-\u003esize;\n 441:\t\n 442:\t\tret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);\n 443:\t\tif (ret)\n 444:\t\t\tgoto err_free_name;\n 445:\t\n 446:\t\tif (dma_buf_set_name(priv-\u003edmabuf, bufname)) {\n 447:\t\t\tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n 448:\t\t\t\t\t \"Failed to set map name '%s'\\n\",\n 449:\t\t\t\t\t bufname);\n 450:\t\t\tkfree(bufname);\n 451:\t\t}\n 452:\t\n 453:\t\t/*\n 454:\t\t * Ownership of the DMABUF file transfers to the VMA so that\n 455:\t\t * other users can locate the DMABUF via a VA. Ownership of\n 456:\t\t * the original VFIO device file being mmap()ed transfers to\n 457:\t\t * priv, and is put when the DMABUF is released. This\n 458:\t\t * intentionally does not use get_file()/vma_set_file()\n 459:\t\t * because the references are already held, and ownership\n 460:\t\t * moves.\n 461:\t\t */\n 462:\t\tpriv-\u003evfile = vma-\u003evm_file;\n 463:\t\tvma-\u003evm_file = priv-\u003edmabuf-\u003efile;\n 464:\t\tvma-\u003evm_private_data = priv;\n 465:\t\n 466:\t\treturn 0;\n 467:\t\n 468:\terr_free_name:\n 469:\t\tkfree(bufname);\n"
} |
| 15/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 47410 output: 20 thoughts: 1801 Thoughts: **VFIO-PCI, DMA-BUF, and Potential Fuzzing Targets** Okay, so here's the deal. I'm looking at how VFIO-PCI handles memory mapping. It's calling `vfio_pci_core_mmap`, which in turn creates a DMA buffer (`dma-buf`) for the BAR, swaps the `vma->vm_file`, sets `vma->vm_private_data`, calls `dma_buf_set_name`, and assigns the mmap operations. When a page fault occurs, it goes through a series of locking and PFN retrieval operations. The interesting bit is that `dma_buf_set_name` gets called. Now, I'm thinking about `dma_buf_ioctl`, and it hits me: the `DMA_BUF_SET_NAME` ioctl directly calls `dma_buf_set_name`! This means a user-space process with a DMA-BUF file descriptor can potentially set the name. Any process that accesses `/dev/dma_heap/*` or `/dev/udmabuf` can issue `ioctl(fd, DMA_BUF_SET_NAME, ...)` on the DMA-BUF file descriptor. Digging into `dma_buf_set_name`, I see it handles the name string. The kernel code checks for null names. Critically, the `dma_buf_ioctl` path calls `strndup_user` to copy the user-provided name and then calls `dma_buf_set_name` and frees the name if the call fails. Now I'm thinking about fuzzing and reachability. Can this path be hit from fuzzing environments? Is VFIO-PCI reachable? Can syzkaller reach it? Or even regular QEMU virtualized environments? I'm checking the instructions and it seems that fuzzing is possible if this code can be executed in regular QEMU virtualized environments. I'm focusing on "reachable executable code" in "core subsystems" or "accessible drivers". `dma-buf.c` is the core DMA-BUF subsystem. So this is a core piece of infrastructure. And I see that `dma_buf_set_name` has been recently modified and refactored with the addition of length checks, etc. I have to consider if VFIO-PCI can run in QEMU. I need to see if PCI devices can be bound to vfio-pci in QEMU. QEMU has e1000, virtio-pci, etc. Can a user bind a PCI device to vfio-pci in QEMU? Yes, via sysfs, like `/sys/bus/pci/drivers/vfio-pci/new_id` or `bind`. Then, I'm thinking that syzkaller runs as root! So, syzkaller could bind a device to vfio-pci. I have to see if syzkaller has vfio support. So I'll check the git logs or comments in the kernel for vfio-pci bugs or crashes. |
| 16/2 | 2026/09/26 21:07 | tool | grepper |
0mError:no matches Args: {
"Expression": "syzbot.*vfio"
}Results: {
"error": "no matches"
} |
| 17/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 49912 output: 20 thoughts: 21 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience and presented in the first person: **Investigating VFIO and Syzkaller Interactions** Alright, so the task at hand is to understand how `vfio` and `syzkaller` are interacting. Given my familiarity with both, my first instinct is to dive into the most obvious source of information: the `git log` related to `vfio`. I'm looking for direct mentions, sure, but I'm also casting a wider net, scanning commit messages for any sign of `syzkaller`, even if it's not explicitly tied to `vfio` at first glance. It's likely those are the initial entrypoints to find any problems. I'll need to use any heuristics that I can. That's the first step; let's see what the revision history can tell us about this relationship. This is the best place to start. |
| 18/2 | 2026/09/26 21:07 | tool | grepper |
0mError:no matches Args: {
"Expression": "syzkaller.*vfio"
}Results: {
"error": "no matches"
} |
| 19/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50167 output: 32 thoughts: 24 Thoughts: Here's my thought process, as if I were analyzing that prompt: **Investigating a Potential VFIO Bug Report** Okay, the prompt is directing me to investigate something within the `drivers/vfio` directory, and it looks like I'm hunting for information related to a "Reported-by: syzbot" tag. My immediate assumption is that `syzbot` is an automated kernel bug reporting tool. The goal is to identify bug reports specifically pertaining to Virtual Function I/O (VFIO) drivers. This suggests I'm likely dealing with a kernel-related issue. The fact that it's in the `drivers/vfio` directory tells me the problem centers around the VFIO subsystem, which handles things like device assignment and IOMMU management. "Reported-by: syzbot" is a key indicator, as it likely means an automated test or fuzzing process detected something in that code. So, I will immediately begin searching within the kernel's source code, specifically looking for any mentions of "Reported-by: syzbot" within the `drivers/vfio` directory. This search should return the specific bug reports, possibly with links to further information like the original report, the test case that triggered it, and any patches or discussions that followed. I will need to understand the details of these bug reports. |
| 20/2 | 2026/09/26 21:07 | tool | grepper |
0mError:no matches Args: {
"Expression": "Reported-by: syzbot",
"PathPrefix": "drivers/vfio"
}Results: {
"error": "no matches"
} |
| 21/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50523 output: 29 thoughts: 17 Thoughts: Here's my take on those thoughts, tailored for an expert audience: **Investigating VFIO Driver Reports** Okay, so I'm thinking, "Reported-by:" is a good starting point. I need to get a handle on what's been said about the `drivers/vfio` area lately. That's usually the best way to get a bird's eye view of the landscape, and see the issues that are hot in the community. Let's see what that search brings up. I'm expecting to find a collection of bug reports, maybe some patch descriptions, and potentially some discussion threads, all with relevant attribution. This search should give me a good jumping-off point to see what the current issues and contributors are. |
| 22/2 | 2026/09/26 21:07 | tool | grepper |
0mError:no matches Args: {
"Expression": "Reported-by:",
"PathPrefix": "drivers/vfio"
}Results: {
"error": "no matches"
} |
| 23/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 50741 output: 34 thoughts: 46 Thoughts: Okay, let's break this down. **Investigating Potential VFIO Interactions within DMA-BUF Reporting** Right, so I'm looking at potential issues, and my immediate thought is to check the `drivers/vfio` area, since reports have been filed against it. It's a common area for interaction with other subsystems. Hmm, but that’s not enough. To fully understand if a link exists, let's broaden the scope. Let me do a quick search within the `drivers/dma-buf` directory. A key indicator of a real problem, especially in this context, would be reports flagged by syzbot. Those are automated fuzzing results, and they are usually a great starting point for investigating subtle or complex interactions. Let me see what that turns up. |
| 24/2 | 2026/09/26 21:07 | tool | grepper |
0mError:no matches Args: {
"Expression": "Reported-by: syzbot",
"PathPrefix": "drivers/dma-buf"
}Results: {
"error": "no matches"
} |
| 25/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 51001 output: 17 thoughts: 64 Thoughts: Okay, here's my thought process on that: **Investigating a Syzbot-Related Absence in DMA-BUF Driver History** Hmm, that's odd. I ran `grep "Reported-by: syzbot" drivers/dma-buf` and got nothing back. That's unusual. I expected to find some traces of syzbot reports impacting that subsystem. Wait a second... I'm being an idiot! `grep` is searching the *working tree*, the current code, not the `git log` history. That explains why it's coming up empty. I need to look at commit messages to find that information. Okay, so how do I figure this out now? I should try searching the working tree for "syzkaller". Hopefully, that'll show me something useful. Let's start with that and then I can adjust from there, depending on what I find. |
| 26/2 | 2026/09/26 21:07 | tool | grepper |
0mArgs:{
"Expression": "syzkaller"
}Results: {
"Output": "Documentation/process/maintainer-kvm-x86.rst=303=In general, regression tests are preferred for any bug that is not trivial to\nDocumentation/process/maintainer-kvm-x86.rst:304:hit. E.g. even if the bug was originally found by a fuzzer such as syzkaller,\nDocumentation/process/maintainer-kvm-x86.rst-305-a targeted regression test may be warranted if the bug requires hitting a\n--\nDocumentation/translations/pt_BR/process/maintainer-kvm-x86.rst=341=trivial de ser atingido. Por exemplo, mesmo que o bug tenha sido originalmente\nDocumentation/translations/pt_BR/process/maintainer-kvm-x86.rst:342:encontrado por um fuzzer como o syzkaller, um teste de regressão direcionado\nDocumentation/translations/pt_BR/process/maintainer-kvm-x86.rst-343-pode ser justificável se o bug exigir que se atinja uma condição de corrida do\n--\nDocumentation/translations/sp_SP/process/maintainer-kvm-x86.rst=366=que no sea trivial de encontrar. Por ejemplo, incluso si el error fue\nDocumentation/translations/sp_SP/process/maintainer-kvm-x86.rst:367:encontrado originalmente por un fuzzer como syzkaller, una prueba de\nDocumentation/translations/sp_SP/process/maintainer-kvm-x86.rst-368-regresión dirigida puede estar justificada si el error requiere golpear una\n--\narch/x86/kernel/Makefile=45=KCOV_INSTRUMENT_unwind_guess.o\t\t\t\t:= n\n--\narch/x86/kernel/Makefile-49-#\narch/x86/kernel/Makefile:50:# As KCOV and KEXEC compatibility should be preserved (e.g. syzkaller is\narch/x86/kernel/Makefile-51-# using it to collect crash dumps during kernel fuzzing), disabling\n--\ndrivers/iommu/iommufd/selftest.c=52=static void mock_dev_disable_iopf(struct device *dev, struct iommu_domain *domain);\n--\ndrivers/iommu/iommufd/selftest.c-56- * to the map ioctl's output, and it has no ide about that. So, simplify things.\ndrivers/iommu/iommufd/selftest.c:57: * In syzkaller mode the 64 bit IOVA is converted into an nth area and offset\ndrivers/iommu/iommufd/selftest.c:58: * value. This has a much smaller randomization space and syzkaller can hit it.\ndrivers/iommu/iommufd/selftest.c-59- */\n--\ndrivers/iommu/iommufd/selftest.c=1540=static int iommufd_test_access_pages(struct iommufd_ucmd *ucmd,\n--\ndrivers/iommu/iommufd/selftest.c-1551-\ndrivers/iommu/iommufd/selftest.c:1552:\t/* Prevent syzkaller from triggering a WARN_ON in kvzalloc() */\ndrivers/iommu/iommufd/selftest.c-1553-\tif (length \u003e 16 * 1024 * 1024)\n--\ndrivers/iommu/iommufd/selftest.c-1595-\ndrivers/iommu/iommufd/selftest.c:1596:\t/* For syzkaller allow uptr to be NULL to skip this check */\ndrivers/iommu/iommufd/selftest.c-1597-\tif (uptr) {\n--\ndrivers/iommu/iommufd/selftest.c=1635=static int iommufd_test_access_rw(struct iommufd_ucmd *ucmd,\n--\ndrivers/iommu/iommufd/selftest.c-1644-\ndrivers/iommu/iommufd/selftest.c:1645:\t/* Prevent syzkaller from triggering a WARN_ON in kvzalloc() */\ndrivers/iommu/iommufd/selftest.c-1646-\tif (length \u003e 16 * 1024 * 1024)\n--\ndrivers/iommu/iommufd/viommu.c=308=iommufd_hw_queue_alloc_phys(struct iommu_hw_queue_alloc *cmd,\n--\ndrivers/iommu/iommufd/viommu.c-330-\t * Use kvcalloc() to avoid memory fragmentation for a large page array.\ndrivers/iommu/iommufd/viommu.c:331:\t * Set __GFP_NOWARN to avoid syzkaller blowups\ndrivers/iommu/iommufd/viommu.c-332-\t */\n--\nlib/Kconfig.debug=2193=config KCOV_INSTRUMENT_ALL\n--\nlib/Kconfig.debug-2197-\thelp\nlib/Kconfig.debug:2198:\t If you are doing generic system call fuzzing (like e.g. syzkaller),\nlib/Kconfig.debug-2199-\t then you will want to instrument the whole kernel and you should\n--\nnet/can/isotp.c=744=static void isotp_rcv(struct sk_buff *skb, void *data)\n--\nnet/can/isotp.c-767-\t * CAN frame reception time. This locking is not needed in real world\nnet/can/isotp.c:768:\t * use cases but the inconsistency can be triggered with syzkaller.\nnet/can/isotp.c-769-\t */\n--\nscripts/checkpatch.pl=2671=sub process {\n--\nscripts/checkpatch.pl-3269-\t\tif (!$in_header_lines \u0026\u0026 !$is_patch \u0026\u0026\nscripts/checkpatch.pl:3270:\t\t $line =~ /((?:(?:BUG: K.|UB)SAN: |Call Trace:|stable\\@|syzkaller))/) {\nscripts/checkpatch.pl-3271-\t\t\t$needs_fixes_tag = $1;\n--\ntools/testing/radix-tree/iteration_check.c=153=static void *tag_entries_fn(void *arg)\n--\ntools/testing/radix-tree/iteration_check.c-163-\ntools/testing/radix-tree/iteration_check.c:164:/* This is a unit test for a bug found by the syzkaller tester */\ntools/testing/radix-tree/iteration_check.c-165-void iteration_test(unsigned order, unsigned test_duration)\n--\ntools/testing/selftests/core/close_range_test.c=307=TEST(close_range_cloexec_unshare)\n--\ntools/testing/selftests/core/close_range_test.c-378-/*\ntools/testing/selftests/core/close_range_test.c:379: * Regression test for syzbot+96cfd2b22b3213646a93@syzkaller.appspotmail.com\ntools/testing/selftests/core/close_range_test.c-380- */\ntools/testing/selftests/core/close_range_test.c=381=TEST(close_range_cloexec_syzbot)\n--\ntools/testing/selftests/core/close_range_test.c-495-/*\ntools/testing/selftests/core/close_range_test.c:496: * Regression test for syzbot+96cfd2b22b3213646a93@syzkaller.appspotmail.com\ntools/testing/selftests/core/close_range_test.c-497- */\n"
} |
| 27/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 53276 output: 41 thoughts: 1435 Thoughts: **Analysis of the VFIO-PCI DMA-BUF Patch** Okay, let's break down this patch and assess its implications, especially concerning reachability for fuzzing. First, I need to methodically examine the modified files and their impact. * `drivers/dma-buf/dma-buf.c`: The introduction of `dma_buf_set_name` as a public, `EXPORT_SYMBOL_NS_GPL` function is interesting. It's called within `dma_buf_ioctl()`, which suggests a new interface to name DMA buffers. I need to keep in mind that this is a core subsystem. * `drivers/vfio/pci/Kconfig`: The `VFIO_PCI_DMABUF` configuration dependency is significantly changed. Previously conditional, now it is unconditionally part of the build when `PCI_P2PDMA` is enabled and `VFIO_PCI_CORE` is selected. This strongly indicates deeper integration and reliance on DMA buffers within VFIO-PCI. * `drivers/vfio/pci/Makefile`: `vfio_pci_dmabuf.o` is now unconditionally included in `vfio-pci-core-y`, which, again, points towards an increased dependency on DMA-BUF functionality. It is always built-in now. * `drivers/vfio/pci/vfio_pci_core.c`: This is where things get really interesting. The BAR mmap process has changed substantially. Instead of directly mapping the BAR, it now creates a `dma_buf`, exports it via `vfio_pci_dmabuf_export`, calls `dma_buf_set_name`, and replaces the `vma->vm_file` with the DMA-BUF's file. This means the entire mmap is now through the dma-buf layer, which is a major shift. The huge fault handler (`vfio_pci_mmap_huge_fault`) has been completely rewritten and incorporates a more robust locking scheme around DMA buffer access. It also includes new features such as `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` which means DMA-BUF's are revocable. The patch has added numerous locking assertions. The inclusion of `hisi_acc_vfio_pci.c` and disabling of `zap_bars_on_revoke` adds an interesting wrinkle. Now, the critical question: Is this reachable? Let's check against the "NEGATIVE CRITERIA": * Is `vfio-pci` structurally unreachable? No! `vfio-pci` is the generic driver and thus reachable. This is *not* a vendor-specific driver requiring special hardware. It binds to *any* PCI device, including emulated ones in QEMU like virtio-pci, e1000, and standard test devices. * `drivers/dma-buf/dma-buf.c`: `dma_buf_ioctl()` and `dma_buf_set_name()` are in `drivers/dma-buf/dma-buf.c`! The DMA-BUF subsystem is *absolutely* reachable, used by virtio-gpu, vgem, udmabuf, dma-heaps, etc.! * `dma_buf_ioctl()` calls: Let's confirm dma-buf call requirements. I'm confident this patch modifies reachable code. It is worthy of fuzzing! |
| 28/2 | 2026/09/26 21:07 | tool | read-file |
0mArgs:{
"File": "drivers/dma-buf/dma-buf.c",
"FirstLine": 545,
"LineCount": 35
}Results: {
"Contents": " 545:\t\t\t\tret = dma_buf_end_cpu_access(dmabuf, direction);\n 546:\t\t\telse\n 547:\t\t\t\tret = dma_buf_begin_cpu_access(dmabuf, direction);\n 548:\t\n 549:\t\t\treturn ret;\n 550:\t\n 551:\t\tcase DMA_BUF_SET_NAME_A:\n 552:\t\tcase DMA_BUF_SET_NAME_B: {\n 553:\t\t\tchar *name = strndup_user((const char __user *)arg,\n 554:\t\t\t\t\t\t DMA_BUF_NAME_LEN);\n 555:\t\n 556:\t\t\tif (IS_ERR(name))\n 557:\t\t\t\treturn PTR_ERR(name);\n 558:\t\n 559:\t\t\tret = dma_buf_set_name(dmabuf, name);\n 560:\t\t\tif (ret)\n 561:\t\t\t\tkfree(name);\n 562:\t\n 563:\t\t\treturn ret;\n 564:\t\t}\n 565:\t\n 566:\t#if IS_ENABLED(CONFIG_SYNC_FILE)\n 567:\t\tcase DMA_BUF_IOCTL_EXPORT_SYNC_FILE:\n 568:\t\t\treturn dma_buf_export_sync_file(dmabuf, (void __user *)arg);\n 569:\t\tcase DMA_BUF_IOCTL_IMPORT_SYNC_FILE:\n 570:\t\t\treturn dma_buf_import_sync_file(dmabuf, (const void __user *)arg);\n 571:\t#endif\n 572:\t\n 573:\t\tdefault:\n 574:\t\t\treturn -ENOTTY;\n 575:\t\t}\n 576:\t}\n 577:\t\n 578:\tstatic void dma_buf_show_fdinfo(struct seq_file *m, struct file *file)\n 579:\t{\n"
} |
| 29/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 56024 output: 29 thoughts: 804 Thoughts: **Analyzing a Promising Patch for Fuzzing** Okay, so I've been digging into this patch, and it's screaming "WorthFuzzing = true" at me. It's touching some really critical areas: the `dma_buf_ioctl` interface with those `DMA_BUF_SET_NAME_A` and `DMA_BUF_SET_NAME_B` ioctls – which, by the way, seem to trigger `dma_buf_set_name` – and, crucially, a bunch of changes within VFIO-PCI. We're talking about a significant rework of how BAR mmaps are handled, now translated into DMABUFs via `vfio_pci_core_mmap_prep_dmabuf`. And it doesn't stop there. This patch introduces `vfio_pci_mmap_huge_fault` to manage faults on these DMABUF-backed VMAs, complex PFN calculations in `vfio_pci_dma_buf_find_pfn`, and state management with `vfio_pci_dma_buf_set_status` using a state machine that handles revocation and completion. Plus, a new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl, which means new code paths and, of course, that always means a higher likelihood of new vulnerabilities being exposed. The best part is that Syzkaller already supports VFIO ioctls; this is the perfect setup for fuzzing this patch! The introduction of new lockdep assertions like `lockdep_assert_held_write(&vdev->memory_lock)`, `lockdep_assert_not_held`, and `lockdep_assert_held` immediately caught my eye. This aligns perfectly with the fuzzing goals: finding new crashes, regressions, exposed latent bugs, and importantly *new or modified assertions in reachable code*. Refactoring and code cleanups also apply here, especially considering the changes to reachable data structures and algorithm flows. So, this patch has all the hallmarks of something ripe for fuzzing: a large modification, new defensive checks, reachable drivers, and potentially complex concurrency scenarios. This is a clear case for prioritizing it! Now, the next step is to understand the dependencies. I need to figure out the required kernel configurations. I'll need to check the Kconfig dependencies in `dma-buf.c`, the VFIO-PCI `Kconfig`, and VFIO's `Kconfig` files. I'll do some grepping and reading to see what the critical dependencies are. I need to make sure I can build a configuration that will actually *exercise* this code. This will be critical to the success of any fuzzing campaign. |
| 30/2 | 2026/09/26 21:07 | tool | grepper |
0mArgs:{
"Expression": "config VFIO_PCI",
"PathPrefix": "drivers/vfio"
}Results: {
"Output": "drivers/vfio/pci/Kconfig=2=menu \"VFIO support for PCI devices\"\n--\ndrivers/vfio/pci/Kconfig-4-\ndrivers/vfio/pci/Kconfig:5:config VFIO_PCI_CORE\ndrivers/vfio/pci/Kconfig-6-\ttristate\n--\ndrivers/vfio/pci/Kconfig-10-\ndrivers/vfio/pci/Kconfig:11:config VFIO_PCI_INTX\ndrivers/vfio/pci/Kconfig-12-\tdef_bool y if !S390\n--\ndrivers/vfio/pci/Kconfig-14-\ndrivers/vfio/pci/Kconfig:15:config VFIO_PCI\ndrivers/vfio/pci/Kconfig-16-\ttristate \"Generic VFIO support for any PCI device\"\n--\ndrivers/vfio/pci/Kconfig=24=if VFIO_PCI\ndrivers/vfio/pci/Kconfig:25:config VFIO_PCI_VGA\ndrivers/vfio/pci/Kconfig-26-\tbool \"Generic VFIO PCI support for VGA devices\"\n--\ndrivers/vfio/pci/Kconfig-34-\ndrivers/vfio/pci/Kconfig:35:config VFIO_PCI_IGD\ndrivers/vfio/pci/Kconfig-36-\tbool \"Generic VFIO PCI extensions for Intel graphics (GVT-d)\"\n--\ndrivers/vfio/pci/Kconfig=46=endif\ndrivers/vfio/pci/Kconfig-47-\ndrivers/vfio/pci/Kconfig:48:config VFIO_PCI_ZDEV_KVM\ndrivers/vfio/pci/Kconfig-49-\tbool \"VFIO PCI extensions for s390x KVM passthrough\"\n--\ndrivers/vfio/pci/Kconfig-58-\ndrivers/vfio/pci/Kconfig:59:config VFIO_PCI_DMABUF\ndrivers/vfio/pci/Kconfig-60-\tdef_bool y if PCI_P2PDMA\n"
} |
| 31/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 57965 output: 36 thoughts: 24 Thoughts: Okay, here's my take on those first 65 lines of `drivers/vfio/pci/Kconfig`. Let's get to it. **Delving into VFIO/PCI Configuration** Alright, I'm diving into this Kconfig file, specifically the VFIO/PCI section. Seems like the initial lines are setting the stage, defining the overall structure and dependencies. I'm seeing the standard `config` directives right off the bat, which is expected. The hierarchy is starting to become clear – nested within `VFIO` generally. The first few options, `VFIO_PCI`, are likely the core enabler for the VFIO PCI driver itself. I'm noting the basic help text describing what this functionality provides, the ability to grant PCI device access to user-space in a secure manner and with performance. The dependencies, like `PCI`, are non-negotiable, of course. We can't really do anything PCI-related without the PCI core functionality enabled in the kernel, so seeing these dependencies makes perfect sense. The `HAS_IOMMU` dependency is important and something I need to keep in mind, and that the `VFIO_IOMMU_TYPE1` dependency is key here as well. I'm guessing that this relates to the IOMMU configuration, and how VFIO will utilize it. It's looking like these early options are fundamental, setting up the foundation for all the VFIO PCI features. The dependencies here, `HAS_IOMMU` and `VFIO_IOMMU_TYPE1`, are going to be key pieces, the latter of which defines the Type 1 IOMMU support, critical for I/O virtualization, and now that is the base for more configurable options. Next, it looks like it's starting to define options for the specific vendors and devices. It seems like it defines a series of settings depending on the devices detected by the driver. I am going to have to dive into those options, and determine their meaning as I parse the code. |
| 32/2 | 2026/09/26 21:07 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/Kconfig",
"FirstLine": 1,
"LineCount": 65
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0-only\n 2:\tmenu \"VFIO support for PCI devices\"\n 3:\t\tdepends on PCI\n 4:\t\n 5:\tconfig VFIO_PCI_CORE\n 6:\t\ttristate\n 7:\t\tselect VFIO_VIRQFD\n 8:\t\tselect IRQ_BYPASS_MANAGER\n 9:\t\tselect DMA_SHARED_BUFFER\n 10:\t\n 11:\tconfig VFIO_PCI_INTX\n 12:\t\tdef_bool y if !S390\n 13:\t\tdepends on VFIO_PCI_CORE\n 14:\t\n 15:\tconfig VFIO_PCI\n 16:\t\ttristate \"Generic VFIO support for any PCI device\"\n 17:\t\tselect VFIO_PCI_CORE\n 18:\t\thelp\n 19:\t\t Support for the generic PCI VFIO bus driver which can connect any\n 20:\t\t PCI device to the VFIO framework.\n 21:\t\n 22:\t\t If you don't know what to do here, say N.\n 23:\t\n 24:\tif VFIO_PCI\n 25:\tconfig VFIO_PCI_VGA\n 26:\t\tbool \"Generic VFIO PCI support for VGA devices\"\n 27:\t\tdepends on X86 \u0026\u0026 VGA_ARB\n 28:\t\thelp\n 29:\t\t Support for VGA extension to VFIO PCI. This exposes an additional\n 30:\t\t region on VGA devices for accessing legacy VGA addresses used by\n 31:\t\t BIOS and generic video drivers.\n 32:\t\n 33:\t\t If you don't know what to do here, say N.\n 34:\t\n 35:\tconfig VFIO_PCI_IGD\n 36:\t\tbool \"Generic VFIO PCI extensions for Intel graphics (GVT-d)\"\n 37:\t\tdepends on X86\n 38:\t\tdefault y\n 39:\t\thelp\n 40:\t\t Support for Intel IGD specific extensions to enable direct\n 41:\t\t assignment to virtual machines. This includes exposing an IGD\n 42:\t\t specific firmware table and read-only copies of the host bridge\n 43:\t\t and LPC bridge config space.\n 44:\t\n 45:\t\t To enable Intel IGD assignment through vfio-pci, say Y.\n 46:\tendif\n 47:\t\n 48:\tconfig VFIO_PCI_ZDEV_KVM\n 49:\t\tbool \"VFIO PCI extensions for s390x KVM passthrough\"\n 50:\t\tdepends on S390 \u0026\u0026 KVM\n 51:\t\tdefault y\n 52:\t\thelp\n 53:\t\t Support s390x-specific extensions to enable support for enhancements\n 54:\t\t to KVM passthrough capabilities, such as interpretive execution of\n 55:\t\t zPCI instructions.\n 56:\t\n 57:\t\t To enable s390x KVM vfio-pci extensions, say Y.\n 58:\t\n 59:\tconfig VFIO_PCI_DMABUF\n 60:\t\tdef_bool y if PCI_P2PDMA\n 61:\t\tdepends on VFIO_PCI_CORE\n 62:\t\n 63:\tsource \"drivers/vfio/pci/mlx5/Kconfig\"\n 64:\t\n 65:\tsource \"drivers/vfio/pci/ism/Kconfig\"\n"
} |
| 33/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 59321 output: 21 thoughts: 100 Thoughts: **Analyzing VFIO_PCI_DMABUF Configuration** Okay, let's break this down. My initial observation is that `VFIO_PCI_DMABUF` is automatically enabled if `PCI_P2PDMA` is also enabled and, of course, if `VFIO_PCI_CORE` is enabled. Seems straightforward enough – a DMA buffer setup dependent on Peer-to-Peer DMA support and the core VFIO PCI functionality. But, I need to be completely certain about the `PCI_P2PDMA` dependency chain. I should probably check the Kconfig entry for `PCI_P2PDMA` to see what _it_ depends on. It's crucial to understand the full enabling path to fully grasp its implications. Gotta follow the breadcrumbs and ensure nothing is missed that could affect VFIO's DMA buffer functionality. |
| 34/2 | 2026/09/26 21:07 | tool | grepper |
0mArgs:{
"Expression": "config PCI_P2PDMA"
}Results: {
"Output": "drivers/pci/Kconfig=197=config PCIE_TPH\n--\ndrivers/pci/Kconfig-205-\ndrivers/pci/Kconfig:206:config PCI_P2PDMA\ndrivers/pci/Kconfig-207-\tbool \"PCI peer-to-peer transfer support\"\n"
} |
| 35/2 | 2026/09/26 21:07 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 59712 output: 35 thoughts: 20 Thoughts: Okay, here's the summary, framed as your thought process: **Diving into PCI Configuration Options** Alright, let's get down to brass tacks. I need to get a better handle on these PCI configuration options. Where would they be defined? Hmm... based on the context, and what I know about kernel configuration, `drivers/pci/Kconfig` is probably the right place to start. And, given that I'm looking for *specific* options, I'll focus around line 206 for starters. Let's see what kind of configuration knobs are being exposed in that area and see if I can glean the specifics I need from there. That should provide a good jumping-off point. Let's see what we've got. |
| 36/2 | 2026/09/26 21:08 | tool | read-file |
0mArgs:{
"File": "drivers/pci/Kconfig",
"FirstLine": 200,
"LineCount": 25
}Results: {
"Contents": " 200:\t\t This option adds support for PCIe TLP Processing Hints (TPH).\n 201:\t\t TPH allows endpoint devices to provide optimization hints, such as\n 202:\t\t desired caching behavior, for requests that target memory space.\n 203:\t\t These hints, called Steering Tags, can empower the system hardware\n 204:\t\t to optimize the utilization of platform resources.\n 205:\t\n 206:\tconfig PCI_P2PDMA\n 207:\t\tbool \"PCI peer-to-peer transfer support\"\n 208:\t\tdepends on ZONE_DEVICE\n 209:\t\t#\n 210:\t\t# The need for the scatterlist DMA bus address flag means PCI P2PDMA\n 211:\t\t# requires 64bit\n 212:\t\t#\n 213:\t\tdepends on 64BIT\n 214:\t\tselect GENERIC_ALLOCATOR\n 215:\t\tselect NEED_SG_DMA_FLAGS\n 216:\t\thelp\n 217:\t\t Enables drivers to do PCI peer-to-peer transactions to and from\n 218:\t\t BARs that are exposed in other devices that are the part of\n 219:\t\t the hierarchy where peer-to-peer DMA is guaranteed by the PCI\n 220:\t\t specification to work (ie. anything below a single PCI bridge).\n 221:\t\n 222:\t\t Many PCIe root complexes do not support P2P transactions and\n 223:\t\t it's hard to tell which support it at all, so at this time,\n 224:\t\t P2P DMA transactions must be between devices behind the same root\n"
} |
| 37/2 | 2026/09/26 21:08 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 60354 output: 18 thoughts: 242 Thoughts: **Analyzing the Dependency Chain** Okay, here's what I'm seeing. `PCI_P2PDMA` has a dependency on both `ZONE_DEVICE` and `64BIT`. That's important to keep in mind, right? Now, the `VFIO_PCI_DMABUF` configuration option is also really interesting, especially since it enables several functions: `vfio_pci_dma_buf_mmap`, `vfio_pci_dma_buf_attach`, and `vfio_pci_core_feature_dma_buf`, which are key pieces to enabling the DMA buffer functionality. But here's the kicker: I see that `vfio_pci_core_feature_dma_buf_revoke` is also enabled by `VFIO_PCI_DMABUF`, and it was introduced in the linked patch. I looked at line 790 in `drivers/vfio/pci/vfio_pci_dmabuf.c` and saw the `#ifdef CONFIG_VFIO_PCI_DMABUF` wrapper around `vfio_pci_core_feature_dma_buf_revoke`. That's a strong indicator that `VFIO_PCI_DMABUF` and therefore `PCI_P2PDMA` are critically important, especially with this new revoke feature tied to the former. Now I should probably double check the `ZONE_DEVICE` dependency too, and its implications in this context. |
| 38/2 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "config ZONE_DEVICE"
}Results: {
"Output": "mm/Kconfig=1220=config ZONE_DMA32\n--\nmm/Kconfig-1224-\nmm/Kconfig:1225:config ZONE_DEVICE\nmm/Kconfig-1226-\tbool \"Device memory (pmem, HMM, etc...) hotplug support\"\n"
} |
| 39/2 | 2026/09/26 21:08 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 61020 output: 33 thoughts: 17 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Navigating the Kconfig Labyrinth** Alright, time to dive into the core configuration specifics. I'm thinking we should zero in on `mm/Kconfig` around line 1225. That likely puts us smack in the middle of a crucial section, probably dealing with memory management options, features, or dependencies. Given the context, I'm expecting to find either a detailed breakdown of a specific memory allocation policy, perhaps related to a new feature under development, or maybe a set of conditional definitions related to a platform's unique memory layout. It could even be some complex interaction between various memory management subsystems. Going straight to that specific line is the fastest way to get to the heart of what's currently relevant. I'll need to remember the surrounding context, of course, because a single line rarely tells the whole story. Let's get reading. |
| 40/2 | 2026/09/26 21:08 | tool | read-file |
0mArgs:{
"File": "mm/Kconfig",
"FirstLine": 1225,
"LineCount": 20
}Results: {
"Contents": "1225:\tconfig ZONE_DEVICE\n1226:\t\tbool \"Device memory (pmem, HMM, etc...) hotplug support\"\n1227:\t\tdepends on MEMORY_HOTREMOVE\n1228:\t\tselect XARRAY_MULTI\n1229:\t\n1230:\t\thelp\n1231:\t\t Device memory hotplug support allows for establishing pmem,\n1232:\t\t or other device driver discovered memory regions, in the\n1233:\t\t memmap. This allows pfn_to_page() lookups of otherwise\n1234:\t\t \"device-physical\" addresses which is needed for DAX, PCI_P2PDMA, and\n1235:\t\t DEVICE_PRIVATE features among others.\n1236:\t\n1237:\t\t Enabling this option will reduce the entropy of x86 KASLR memory\n1238:\t\t regions. For example - on a 46 bit system, the entropy goes down\n1239:\t\t from 16 bits to 15 bits. The actual reduction in entropy depends\n1240:\t\t on the physical address bits, on processor features, kernel config\n1241:\t\t (5 level page table) and physical memory present on the system.\n1242:\t\n1243:\t#\n1244:\t# Helpers to mirror range of the CPU page tables of a process into device page\n"
} |
| 41/2 | 2026/09/26 21:08 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 61632 output: 27 thoughts: 39 Thoughts: Okay, let's break this down. **Investigating VFIO_PCI Dependencies** Alright, I need to understand the relationship between `VFIO_PCI` and `VFIO`. Is `VFIO_PCI` a standalone module, or does it rely on the core `VFIO` framework? It's essential to grasp this dependency for any configuration or troubleshooting work. The most direct path to finding that information will be within the kernel configuration files. Let me go check `drivers/vfio/Kconfig`. That should have the answers I'm looking for. |
| 42/2 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "config VFIO",
"PathPrefix": "drivers/vfio"
}Results: {
"Output": "drivers/vfio/Kconfig-1-# SPDX-License-Identifier: GPL-2.0-only\ndrivers/vfio/Kconfig:2:menuconfig VFIO\ndrivers/vfio/Kconfig-3-\ttristate \"VFIO Non-Privileged userspace driver framework\"\n--\ndrivers/vfio/Kconfig=16=if VFIO\ndrivers/vfio/Kconfig:17:config VFIO_DEVICE_CDEV\ndrivers/vfio/Kconfig-18-\tbool \"Support for the VFIO cdev /dev/vfio/devices/vfioX\"\n--\ndrivers/vfio/Kconfig-28-\ndrivers/vfio/Kconfig:29:config VFIO_GROUP\ndrivers/vfio/Kconfig-30-\tbool \"Support for the VFIO group /dev/vfio/$group_id\"\n--\ndrivers/vfio/Kconfig-38-\ndrivers/vfio/Kconfig:39:config VFIO_CONTAINER\ndrivers/vfio/Kconfig-40-\tbool \"Support for the VFIO container /dev/vfio/vfio\"\n--\ndrivers/vfio/Kconfig=51=if VFIO_CONTAINER\ndrivers/vfio/Kconfig:52:config VFIO_IOMMU_TYPE1\ndrivers/vfio/Kconfig-53-\ttristate\n--\ndrivers/vfio/Kconfig-55-\ndrivers/vfio/Kconfig:56:config VFIO_IOMMU_SPAPR_TCE\ndrivers/vfio/Kconfig-57-\ttristate\n--\ndrivers/vfio/Kconfig=60=endif\ndrivers/vfio/Kconfig-61-\ndrivers/vfio/Kconfig:62:config VFIO_NOIOMMU\ndrivers/vfio/Kconfig-63-\tbool \"VFIO No-IOMMU support\"\n--\ndrivers/vfio/Kconfig-78-\ndrivers/vfio/Kconfig:79:config VFIO_VIRQFD\ndrivers/vfio/Kconfig-80-\tbool\n--\ndrivers/vfio/Kconfig-83-\ndrivers/vfio/Kconfig:84:config VFIO_DEBUGFS\ndrivers/vfio/Kconfig-85-\tbool \"Export VFIO internals in DebugFS\"\n--\ndrivers/vfio/cdx/Kconfig-7-\ndrivers/vfio/cdx/Kconfig:8:config VFIO_CDX\ndrivers/vfio/cdx/Kconfig-9-\ttristate \"VFIO support for CDX bus devices\"\n--\ndrivers/vfio/fsl-mc/Kconfig=1=menu \"VFIO support for FSL_MC bus devices\"\n--\ndrivers/vfio/fsl-mc/Kconfig-3-\ndrivers/vfio/fsl-mc/Kconfig:4:config VFIO_FSL_MC\ndrivers/vfio/fsl-mc/Kconfig-5-\ttristate \"VFIO support for QorIQ DPAA2 fsl-mc bus devices\"\n--\ndrivers/vfio/mdev/Kconfig-2-\ndrivers/vfio/mdev/Kconfig:3:config VFIO_MDEV\ndrivers/vfio/mdev/Kconfig-4-\ttristate\n--\ndrivers/vfio/pci/Kconfig=2=menu \"VFIO support for PCI devices\"\n--\ndrivers/vfio/pci/Kconfig-4-\ndrivers/vfio/pci/Kconfig:5:config VFIO_PCI_CORE\ndrivers/vfio/pci/Kconfig-6-\ttristate\n--\ndrivers/vfio/pci/Kconfig-10-\ndrivers/vfio/pci/Kconfig:11:config VFIO_PCI_INTX\ndrivers/vfio/pci/Kconfig-12-\tdef_bool y if !S390\n--\ndrivers/vfio/pci/Kconfig-14-\ndrivers/vfio/pci/Kconfig:15:config VFIO_PCI\ndrivers/vfio/pci/Kconfig-16-\ttristate \"Generic VFIO support for any PCI device\"\n--\ndrivers/vfio/pci/Kconfig=24=if VFIO_PCI\ndrivers/vfio/pci/Kconfig:25:config VFIO_PCI_VGA\ndrivers/vfio/pci/Kconfig-26-\tbool \"Generic VFIO PCI support for VGA devices\"\n--\ndrivers/vfio/pci/Kconfig-34-\ndrivers/vfio/pci/Kconfig:35:config VFIO_PCI_IGD\ndrivers/vfio/pci/Kconfig-36-\tbool \"Generic VFIO PCI extensions for Intel graphics (GVT-d)\"\n--\ndrivers/vfio/pci/Kconfig=46=endif\ndrivers/vfio/pci/Kconfig-47-\ndrivers/vfio/pci/Kconfig:48:config VFIO_PCI_ZDEV_KVM\ndrivers/vfio/pci/Kconfig-49-\tbool \"VFIO PCI extensions for s390x KVM passthrough\"\n--\ndrivers/vfio/pci/Kconfig-58-\ndrivers/vfio/pci/Kconfig:59:config VFIO_PCI_DMABUF\ndrivers/vfio/pci/Kconfig-60-\tdef_bool y if PCI_P2PDMA\n--\ndrivers/vfio/platform/Kconfig=2=menu \"VFIO support for platform devices\"\n--\ndrivers/vfio/platform/Kconfig-4-\ndrivers/vfio/platform/Kconfig:5:config VFIO_PLATFORM_BASE\ndrivers/vfio/platform/Kconfig-6-\ttristate\n--\ndrivers/vfio/platform/Kconfig-8-\ndrivers/vfio/platform/Kconfig:9:config VFIO_PLATFORM\ndrivers/vfio/platform/Kconfig-10-\ttristate \"Generic VFIO support for any platform device\"\n--\ndrivers/vfio/platform/Kconfig-18-\ndrivers/vfio/platform/Kconfig:19:config VFIO_AMBA\ndrivers/vfio/platform/Kconfig-20-\ttristate \"VFIO support for AMBA devices (DEPRECATED)\"\n--\ndrivers/vfio/platform/reset/Kconfig=2=if VFIO_PLATFORM\ndrivers/vfio/platform/reset/Kconfig:3:config VFIO_PLATFORM_CALXEDAXGMAC_RESET\ndrivers/vfio/platform/reset/Kconfig-4-\ttristate \"VFIO support for calxeda xgmac reset (DEPRECATED)\"\n--\ndrivers/vfio/platform/reset/Kconfig-9-\ndrivers/vfio/platform/reset/Kconfig:10:config VFIO_PLATFORM_AMDXGBE_RESET\ndrivers/vfio/platform/reset/Kconfig-11-\ttristate \"VFIO support for AMD XGBE reset (DEPRECATED)\"\n--\ndrivers/vfio/platform/reset/Kconfig-16-\ndrivers/vfio/platform/reset/Kconfig:17:config VFIO_PLATFORM_BCMFLEXRM_RESET\ndrivers/vfio/platform/reset/Kconfig-18-\ttristate \"VFIO support for Broadcom FlexRM reset (DEPRECATED)\"\n"
} |
| 43/2 | 2026/09/26 21:08 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 63579 output: 20 thoughts: 22 Thoughts: Here's my thought process, as requested: **Debugging the DMA Shared Buffer** Okay, time to get down to brass tacks. I need to understand what's happening with the `DMA_SHARED_BUFFER`. My first step, as always, is to establish its provenance. Where in the code is `DMA_SHARED_BUFFER` actually defined? I need to track down its declaration – the data type, any associated attributes, initial values, and, critically, which modules are aware of its existence. Knowing the definition is fundamental to understanding its behavior. I need to get a handle on this key variable immediately. |
| 44/2 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "config DMA_SHARED_BUFFER"
}Results: {
"Output": "drivers/base/Kconfig=213=source \"drivers/base/regmap/Kconfig\"\ndrivers/base/Kconfig-214-\ndrivers/base/Kconfig:215:config DMA_SHARED_BUFFER\ndrivers/base/Kconfig-216-\tbool\n"
} |
| 45/2 | 2026/09/26 21:08 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 63830 output: 372 thoughts: 2170 Thoughts: **VFIO PCI DMABUF Integration: Fuzzing Justification** Okay, let's break this down. My analysis indicates a significant shift in how VFIO PCI handles BAR memory mappings. The core change involves integrating DMABUF functionality, which warrants careful scrutiny. `DMA_SHARED_BUFFER` is key here, and I need to ensure the correct dependencies are in place, particularly `VFIO`, `VFIO_PCI`, `VFIO_PCI_CORE`, `VFIO_PCI_DMABUF`, `PCI_P2PDMA`, and that these are all correctly enabled. The fact that `PCI_P2PDMA` enables `VFIO_PCI_DMABUF` confirms this. I've also verified that the dependencies for `PCI_P2PDMA` such as `ZONE_DEVICE`, and `MEMORY_HOTREMOVE`, are appropriately handled. Now, the focus should be on the new and modified functions that implement this integration, not ubiquitous functions. I've pinpointed several critical areas to focus on in the fuzzing: 1. `dma_buf_set_name`: This function, newly exported, now plays a role in both `vfio-pci` and dmabuf ioctls, so it needs careful attention. 2. `vfio_pci_core_mmap_prep_dmabuf`: This is entirely new. It's the preparation stage for DMABUF-backed BAR mmaps. 3. `vfio_pci_dma_buf_find_pfn`: Another new addition, crucial for PFN search and alignment calculations during faults. 4. `vfio_pci_mmap_huge_fault`: A critical function, completely rewritten to handle BAR memory mappings and DMABUF-based mappings. This is a high-priority target. 5. `vfio_pci_core_feature_dma_buf_revoke`: The entry point for the new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl. 6. `vfio_pci_lock_revoke_bars` and `vfio_pci_unrevoke_bars`: These are critical for locking and unlocking BARs to ensure they are consistent across reset/PM/config writes. 7. `vfio_pci_dma_buf_move`: A modified function to handle DMABUF revocation status changes. 8. `vfio_pci_core_mmap`: A modified function and the main entry point to setting up DMABUF-backed BAR mapping. 9. `vfio_pci_dma_buf_mmap`: A new mmap callback for `vfio_pci` dmabufs. These are the functions that directly implement the core logic of the DMABUF integration. They are feature-specific and non-ubiquitous, making them ideal targets for fuzzing. I've verified the configurations needed for this new functionality are `VFIO`, `VFIO_PCI`, `VFIO_PCI_CORE`, `VFIO_PCI_DMABUF`, `PCI_P2PDMA`, and `DMA_SHARED_BUFFER`. I'm confident in the analysis and the identified symbols. Now, let's call `set-results`. |
| 46/2 | 2026/09/26 21:08 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_core_mmap",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch modifies reachable code in DMA-BUF and VFIO-PCI. It converts VFIO PCI BAR memory mappings to be backed by DMABUFs, introduces a new fault handler (vfio_pci_mmap_huge_fault) and PFN search logic (vfio_pci_dma_buf_find_pfn), adds a new user-facing device feature for buffer revocation (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE), restructures BAR revocation synchronization with lockdep assertions, and refactors dma_buf_set_name and its ioctl handling. Fuzzing is warranted to detect races, regressions, and assertion failures.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_core_mmap",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch modifies reachable code in DMA-BUF and VFIO-PCI. It converts VFIO PCI BAR memory mappings to be backed by DMABUFs, introduces a new fault handler (vfio_pci_mmap_huge_fault) and PFN search logic (vfio_pci_dma_buf_find_pfn), adds a new user-facing device feature for buffer revocation (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE), restructures BAR revocation synchronization with lockdep assertions, and refactors dma_buf_set_name and its ioctl handling. Fuzzing is warranted to detect races, regressions, and assertion failures.",
"WorthFuzzing": true
} |
| 47/1 | 2026/09/26 21:08 | action | If |
1mArgs:{
"WorthFuzzing": true
}Results: null |
| 48/2 | 2026/09/26 21:08 | agent | kmsan-evaluator |
1mModel:core Args: null Results: {
"KMSANReasoning": "The patch series refactors DMA-BUF naming (allowing kernel exporters to set debug names) and integrates DMA-BUF backing into the VFIO PCI BAR mmap infrastructure, introducing a new VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE feature and granular locking with a new rwsem (dmabuf_lock).\n\nKMSAN vs KASAN Applicability:\n1. Uninitialized Memory / Info-leak Analysis:\n - All heap allocations for newly introduced structures (struct vfio_pci_dma_buf, phys_vec) use kzalloc_obj, ensuring all fields are zero-initialized.\n - The string handling in dma_buf_set_name operates on NULL-terminated, initialized strings produced by kasprintf() and strndup_user().\n - The new ioctl feature (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE) is an input-only command (VFIO_DEVICE_FEATURE_SET) that uses copy_from_user; no kernel data or padding is returned or copied to user space.\n - Local variables in the fault handler (vfio_pci_mmap_huge_fault, vfio_pci_dma_buf_find_pfn) such as pfn are explicitly initialized prior to being used or checked.\n\n2. Bug Detector Suitability:\n - The potential bugs introduced by these changes are concurrency/locking deadlocks (dmabuf_lock vs memory_lock vs DMA-BUF resv lock), object lifetime/racing teardown issues (DMABUF vs VFIO device vs VMA references), and potential use-after-free or invalid PFN insertion.\n - These issues are comprehensively covered by LOCKDEP and KASAN.\n - There are no uninitialized memory reads or info-leaks, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 0b43d95e7d80a0760fdb300edd1dd05f46dba412
Author: syz-cluster <triage@syzkaller.com>
Date: Sat Sep 26 21:06:15 2026 +0000
syz-cluster: applied patch under review
diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
index 4c9add51f9ef5..e4ec504e92884 100644
--- a/drivers/dma-buf/dma-buf.c
+++ b/drivers/dma-buf/dma-buf.c
@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)
return events;
}
-/**
- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.
- * It could support changing the name of the dma-buf if the same
- * piece of memory is used for multiple purpose between different devices.
- *
- * @dmabuf: [in] dmabuf buffer that will be renamed.
- * @buf: [in] A piece of userspace memory that contains the name of
- * the dma-buf.
- *
- * Returns 0 on success. If the dma-buf buffer is already attached to
- * devices, return -EBUSY.
- *
- */
-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)
-{
- char *name = strndup_user(buf, DMA_BUF_NAME_LEN);
-
- if (IS_ERR(name))
- return PTR_ERR(name);
-
- spin_lock(&dmabuf->name_lock);
- kfree(dmabuf->name);
- dmabuf->name = name;
- spin_unlock(&dmabuf->name_lock);
-
- return 0;
-}
-
#if IS_ENABLED(CONFIG_SYNC_FILE)
static long dma_buf_export_sync_file(struct dma_buf *dmabuf,
void __user *user_data)
@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,
return ret;
case DMA_BUF_SET_NAME_A:
- case DMA_BUF_SET_NAME_B:
- return dma_buf_set_name(dmabuf, (const char __user *)arg);
+ case DMA_BUF_SET_NAME_B: {
+ char *name = strndup_user((const char __user *)arg,
+ DMA_BUF_NAME_LEN);
+
+ if (IS_ERR(name))
+ return PTR_ERR(name);
+
+ ret = dma_buf_set_name(dmabuf, name);
+ if (ret)
+ kfree(name);
+
+ return ret;
+ }
#if IS_ENABLED(CONFIG_SYNC_FILE)
case DMA_BUF_IOCTL_EXPORT_SYNC_FILE:
@@ -874,6 +857,41 @@ void dma_buf_put(struct dma_buf *dmabuf)
}
EXPORT_SYMBOL_NS_GPL(dma_buf_put, "DMA_BUF");
+/**
+ * dma_buf_set_name - Set a dmabuf's name
+ * Intended to be used by the exporter to set a name for debug
+ * purposes. This can also change an existing name if the same piece
+ * of memory is used for multiple purposes over time.
+ *
+ * @dmabuf: [in] dmabuf buffer that will be renamed.
+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or
+ * similar. This takes ownership of the allocation
+ * on success, which will be kfree()d when the
+ * dmabuf is released or a new name assigned.
+ *
+ * Returns 0 on success, -EINVAL if the name is NULL, or -E2BIG if the
+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.
+ */
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)
+{
+ if (!name)
+ return -EINVAL;
+
+ /* dmabuffs_dname() won't use the string if the length
+ * (including terminator) exceeds DMA_BUF_NAME_LEN:
+ */
+ if (strlen(name) >= DMA_BUF_NAME_LEN)
+ return -E2BIG;
+
+ spin_lock(&dmabuf->name_lock);
+ kfree(dmabuf->name);
+ dmabuf->name = name;
+ spin_unlock(&dmabuf->name_lock);
+
+ return 0;
+}
+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, "DMA_BUF");
+
static int dma_buf_wrap_sg_table(struct sg_table **sg_table)
{
struct scatterlist *to_sg, *from_sg;
diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig
index 296bf01e185ec..c6d6fb09af86c 100644
--- a/drivers/vfio/pci/Kconfig
+++ b/drivers/vfio/pci/Kconfig
@@ -6,6 +6,7 @@ config VFIO_PCI_CORE
tristate
select VFIO_VIRQFD
select IRQ_BYPASS_MANAGER
+ select DMA_SHARED_BUFFER
config VFIO_PCI_INTX
def_bool y if !S390
@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM
To enable s390x KVM vfio-pci extensions, say Y.
config VFIO_PCI_DMABUF
- def_bool y if VFIO_PCI_CORE && PCI_P2PDMA && DMA_SHARED_BUFFER
+ def_bool y if PCI_P2PDMA
+ depends on VFIO_PCI_CORE
source "drivers/vfio/pci/mlx5/Kconfig"
diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile
index 6138f1bf241df..881452ea89be0 100644
--- a/drivers/vfio/pci/Makefile
+++ b/drivers/vfio/pci/Makefile
@@ -1,8 +1,7 @@
# SPDX-License-Identifier: GPL-2.0-only
-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o
+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o
vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o
-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o
obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o
vfio-pci-y := vfio_pci.o
diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
index 86362ec424a50..14622556355eb 100644
--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);
struct pci_dev *pdev = to_pci_dev(core_vdev->dev);
struct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);
+ int ret;
hisi_acc_vdev->vf_id = pci_iov_vf_id(pdev) + 1;
hisi_acc_vdev->pf_qm = pf_qm;
@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
core_vdev->migration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;
core_vdev->mig_ops = &hisi_acc_vfio_pci_migrn_state_ops;
- return vfio_pci_core_init_dev(core_vdev);
+ ret = vfio_pci_core_init_dev(core_vdev);
+ if (ret)
+ return ret;
+ /*
+ * hisi_acc_vfio_pci_mmap() calls down to
+ * vfio_pci_core_mmap(), so BAR mappings are still
+ * DMABUF-backed. They don't require a zap on revoke, so opt
+ * out:
+ */
+ hisi_acc_vdev->core_device.zap_bars_on_revoke = false;
+
+ return 0;
}
static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index 9914f3ac69aef..cef337f4e8f2e 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
virt_mem = !!(le16_to_cpu(*virt_cmd) & PCI_COMMAND_MEMORY);
new_mem = !!(new_cmd & PCI_COMMAND_MEMORY);
- if (!new_mem) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (!new_mem)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
/*
* If the user is writing mem/io enable (new_mem/io) and we
@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
*virt_cmd |= cpu_to_le16(new_cmd & mask);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state)
{
- if (state >= PCI_D3hot) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (state >= PCI_D3hot)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
vfio_pci_set_power_state(vdev, state);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_EXP_DEVCAP_FLR)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_AF_CAP_FLR) && (cap & PCI_AF_CAP_TP)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d875..8860185cff49f 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -13,6 +13,8 @@
#include <linux/aperture.h>
#include <linux/debugfs.h>
#include <linux/device.h>
+#include <linux/dma-buf.h>
+#include <linux/dma-resv.h>
#include <linux/eventfd.h>
#include <linux/file.h>
#include <linux/interrupt.h>
@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
* The vdev power related flags are protected with 'memory_lock'
* semaphore.
*/
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
if (vdev->pm_runtime_engaged) {
up_write(&vdev->memory_lock);
@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)
down_write(&vdev->memory_lock);
__vfio_pci_runtime_pm_exit(vdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)
*/
down_write(&vdev->memory_lock);
if (vdev->pm_wake_eventfd_ctx) {
- eventfd_signal(vdev->pm_wake_eventfd_ctx);
+ struct eventfd_ctx *ctx = vdev->pm_wake_eventfd_ctx;
+
+ vdev->pm_wake_eventfd_ctx = NULL;
__vfio_pci_runtime_pm_exit(vdev);
+ if (__vfio_pci_memory_enabled(vdev))
+ vfio_pci_unrevoke_bars(vdev);
+ eventfd_signal(ctx);
+ eventfd_ctx_put(ctx);
}
up_write(&vdev->memory_lock);
@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
vdev->has_vga = true;
vfio_pci_core_map_bars(vdev);
+ vdev->bars_revoked = false;
return 0;
@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
return ret;
}
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);
+
static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
void __user *arg)
{
@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
if (!vdev->reset_works)
return -EINVAL;
- vfio_pci_zap_and_down_write_memory_lock(vdev);
+ down_write(&vdev->memory_lock);
/*
* This function can be invoked while the power state is non-D0. If
@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
* have NoSoftRst-, the reset function can cause the PCI config space
* reset without restoring the original state (saved locally in
* 'vdev->pm_save').
+ *
+ * The zap is done after making the device accessible in D0,
+ * because a DMABUF importer could access the device as part
+ * of its revocation cleanup.
*/
vfio_pci_set_power_state(vdev, PCI_D0);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_revoke_bars(vdev);
+
ret = pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
return ret;
@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
return vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);
case VFIO_DEVICE_FEATURE_ZPCI_ERROR:
return vfio_pci_zdev_feature_err(device, flags, arg, argsz);
+ case VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:
+ return vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);
default:
return -ENOTTY;
}
@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu
}
EXPORT_SYMBOL_GPL(vfio_pci_core_write);
-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)
{
- struct vfio_device *core_vdev = &vdev->vdev;
- loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
- loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
- loff_t len = end - start;
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, true);
- unmap_mapping_range(core_vdev->inode->i_mapping, start, len, true);
+ /*
+ * If a driver could possibly create BAR mappings in the
+ * vdev's address_space, do an additional zap on revoke. See
+ * vfio_pci_core_init_dev().
+ */
+ if (vdev->zap_bars_on_revoke) {
+ struct vfio_device *core_vdev = &vdev->vdev;
+ loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
+ loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
+ loff_t len = end - start;
+
+ unmap_mapping_range(core_vdev->inode->i_mapping,
+ start, len, true);
+ }
}
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)
{
down_write(&vdev->memory_lock);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
+}
+
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)
+{
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, false);
}
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)
@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c
up_write(&vdev->memory_lock);
}
-static unsigned long vma_to_pfn(struct vm_area_struct *vma)
-{
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- int index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
- u64 pgoff;
-
- pgoff = vma->vm_pgoff &
- ((1U << (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);
-
- return (pci_resource_start(vdev->pdev, index) >> PAGE_SHIFT) + pgoff;
-}
-
vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,
struct vm_fault *vmf,
unsigned long pfn,
@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,
unsigned int order)
{
struct vm_area_struct *vma = vmf->vma;
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- unsigned long addr = vmf->address & ~((PAGE_SIZE << order) - 1);
- unsigned long pgoff = linear_page_delta(vma, addr);
- unsigned long pfn = vma_to_pfn(vma) + pgoff;
- vm_fault_t ret = VM_FAULT_FALLBACK;
-
- if (is_aligned_for_order(vma, addr, pfn, order)) {
- scoped_guard(rwsem_read, &vdev->memory_lock)
- ret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);
+ struct vfio_pci_dma_buf *priv = vma->vm_private_data;
+ struct vfio_pci_core_device *vdev;
+ unsigned long pfn = 0;
+ vm_fault_t ret = VM_FAULT_SIGBUS;
+
+ /*
+ * The only thing this can rely on is that the DMABUF relating
+ * to the VMA's vm_file exists (priv).
+ *
+ * A DMABUF for a VFIO device fd mmap() holds a reference to
+ * the original VFIO device fd, but an explicitly-exported
+ * DMABUF does not. The original fd might have closed,
+ * meaning this fault can race with
+ * vfio_pci_dma_buf_cleanup(), meaning the buffer could have
+ * been revoked (in which case priv->vdev might be NULL), and
+ * the VFIO device registration might have been dropped.
+ *
+ * With the goal of taking vdev locks in a world where vdev
+ * might not still exist:
+ *
+ * 1. Take the resv lock on the DMABUF:
+ * - If racing cleanup got in first, the buffer is revoked;
+ * stop/exit if so.
+ * - If we got in first, the buffer is not revoked so vdev is
+ * non-NULL, accessible, and cleanup _has not yet put the
+ * VFIO device registration_. So, the device refcount must
+ * be >0.
+ *
+ * 2. Take vfio_device registration (refcount guaranteed >0
+ * hereafter).
+ *
+ * 3. Unlock the DMABUF's resv lock:
+ * - A racing cleanup can now complete.
+ * - But, the device refcount >0, meaning the vfio_device
+ * (and vfio_pcie_core device vdev) have not yet been
+ * freed. vdev is accessible, even if the DMABUF has been
+ * revoked or cleanup has happened, because
+ * vfio_unregister_group_dev() can't complete.
+ *
+ * 4. Take the vdev->memory_lock then vdev->dmabuf_lock:
+ * - Either the DMABUF is usable, or has been cleaned up.
+ * - It's not necessary to also take the resv lock, because
+ * the status/vdev can't change while dmabuf_lock is held.
+ * - Test the DMABUF revocation status again: if it was
+ * revoked between 1 and 4, return a SIGBUS. Otherwise,
+ * return a PFN.
+ *
+ * 5. Unlock, done.
+ */
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+
+ if (priv->status != VFIO_PCI_DMABUF_OK) {
+ pr_debug_ratelimited("%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\n",
+ __func__, vmf->address, vma->vm_pgoff);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+
+ /* If the buffer isn't revoked, vdev is valid */
+ vdev = priv->vdev;
+
+ if (!vfio_device_try_get_registration(&vdev->vdev)) {
+ /*
+ * If vdev != NULL (above), the registration should
+ * already be >0 and so this try_get should never
+ * fail.
+ */
+ dev_warn_ratelimited(&vdev->pdev->dev,
+ "%s: Unexpected registration failure\n",
+ __func__);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+ dma_resv_unlock(priv->dmabuf->resv);
+
+ /* memory_lock for vfio_pci_vmf_insert_pfn() */
+ down_read(&vdev->memory_lock);
+ /* Re-test revocation status under dmabuf_lock */
+ down_read(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_OK) {
+ int pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,
+ vmf->address,
+ order, &pfn);
+
+ if (pres == 0)
+ ret = vfio_pci_vmf_insert_pfn(vdev, vmf,
+ pfn, order);
+ else if (pres == -ERANGE)
+ ret = VM_FAULT_FALLBACK;
}
+ up_read(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
dev_dbg_ratelimited(&vdev->pdev->dev,
- "%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\n",
- __func__, order,
- vma->vm_pgoff >>
- (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),
- pgoff, (unsigned int)ret);
+ "%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\n",
+ __func__, order, pfn, vmf->address,
+ vma->vm_pgoff, (unsigned int)ret);
+ vfio_device_put_registration(&vdev->vdev);
return ret;
}
@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {
#endif
};
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)
+{
+ vma->vm_ops = &vfio_pci_mmap_ops;
+}
+
int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)
{
struct vfio_pci_core_device *vdev =
@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
unsigned int index;
u64 phys_len, req_len, pgoff, req_start;
void __iomem *bar_io;
+ int ret;
index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
if (IS_ERR(bar_io))
return PTR_ERR(bar_io);
- vma->vm_private_data = vdev;
+ ret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,
+ pci_resource_start(pdev, index),
+ req_len, index);
+ if (ret)
+ return ret;
+
vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
return ret;
INIT_LIST_HEAD(&vdev->dmabufs);
init_rwsem(&vdev->memory_lock);
+ init_rwsem(&vdev->dmabuf_lock);
xa_init(&vdev->ctx);
+ /*
+ * If a driver overrides .mmap, it has to be assumed that it
+ * might not use the DMABUF-backed core mmap; this flag
+ * enables a zap at revoke time. A driver can opt out by
+ * clearing this flag at init, if their .mmap override calls
+ * down to vfio_pci_core_mmap().
+ */
+ if (vdev->vdev.ops->mmap != vfio_pci_core_mmap)
+ vdev->zap_bars_on_revoke = true;
+
return 0;
}
EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);
@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
}
/*
- * Take the memory write lock for each device and zap BAR
- * mappings to prevent the user accessing the device while in
- * reset. Locking multiple devices is prone to deadlock,
+ * Take the memory write lock for each device and
+ * zap/revoke BAR mappings to prevent the user (or
+ * peers) accessing the device while in reset.
+ * Locking multiple devices is prone to deadlock,
* runaway and unwind if we hit contention.
*/
if (!down_write_trylock(&vdev->memory_lock)) {
@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
break;
}
- vfio_pci_dma_buf_move(vdev, true);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
}
if (!list_entry_is_head(vdev,
@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
list_for_each_entry_from_reverse(vdev, &dev_set->device_list,
vdev.dev_set_list) {
if (vdev->vdev.open_count && __vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d68..b57bfaefd9fae 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -3,25 +3,14 @@
*/
#include <linux/dma-buf-mapping.h>
#include <linux/pci-p2pdma.h>
+#include <linux/dma-buf.h>
#include <linux/dma-resv.h>
#include "vfio_pci_priv.h"
MODULE_IMPORT_NS("DMA_BUF");
-struct vfio_pci_dma_buf {
- struct dma_buf *dmabuf;
- struct vfio_pci_core_device *vdev;
- struct list_head dmabufs_elm;
- size_t size;
- struct phys_vec *phys_vec;
- struct p2pdma_provider *provider;
- u32 nr_ranges;
- struct kref kref;
- struct completion comp;
- u8 revoked : 1;
-};
-
+#ifdef CONFIG_VFIO_PCI_DMABUF
static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
struct dma_buf_attachment *attachment)
{
@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
if (!attachment->peer2peer)
return -EOPNOTSUPP;
- if (priv->revoked)
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
return -ENODEV;
if (!dma_buf_attach_revocable(attachment))
@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
return 0;
}
+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)
+{
+ struct vfio_pci_dma_buf *priv = dmabuf->priv;
+
+ /*
+ * dma_buf_mmap_internal() has asserted that the VMA is
+ * contained within the DMABUF size before calling this.
+ *
+ * Also, if we observe that the buffer is revoked now then
+ * refuse the mmap(). This is a belt-and-braces early failure
+ * to ease debugging a revoked buffer being used. Userspace
+ * might also race an mmap() against an explicit revocation,
+ * or an action causing a revoke; race scenarios are still
+ * safe because the fault handler ultimately prevents access
+ * to a revoked buffer if it isn't caught here.
+ */
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
+ return -ENODEV;
+ /*
+ * Make clear that anything with an offset adjustment is
+ * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()
+ * maths would underflow; this doesn't happen through the
+ * regular DMABUF export path used with this mmap(). A DMABUF
+ * implicitly created for BAR mmap could have adjust > 0, but
+ * these can't currently be re-opened and mmap()ed again.
+ * Catch here in case that assumption ever changes.
+ */
+ if (priv->vma_pgoff_adjust)
+ return -EINVAL;
+ if ((vma->vm_flags & VM_SHARED) == 0)
+ return -EINVAL;
+
+ vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
+ vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
+
+ /* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */
+ vm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |
+ VM_DONTEXPAND | VM_DONTDUMP);
+ vma->vm_private_data = priv;
+ vfio_pci_set_vma_ops(vma);
+
+ return 0;
+}
+#else
+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
+ struct dma_buf_attachment *attachment)
+{
+ /*
+ * Explicit export can't occur without the DMABUF feature, but
+ * DMABUFs are implicitly created for BAR mappings. An
+ * .attach that fails prevents dma_buf_attach().
+ */
+ return -EOPNOTSUPP;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
static void vfio_pci_dma_buf_done(struct kref *kref)
{
struct vfio_pci_dma_buf *priv =
@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,
dma_resv_assert_held(priv->dmabuf->resv);
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return ERR_PTR(-ENODEV);
ret = dma_buf_phys_vec_to_sgt(attachment, priv->provider,
@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)
* The refcount prevents both.
*/
if (priv->vdev) {
- down_write(&priv->vdev->memory_lock);
+ down_write(&priv->vdev->dmabuf_lock);
list_del_init(&priv->dmabufs_elm);
- up_write(&priv->vdev->memory_lock);
+ up_write(&priv->vdev->dmabuf_lock);
vfio_device_put_registration(&priv->vdev->vdev);
}
+ if (priv->vfile)
+ fput(priv->vfile);
kfree(priv->phys_vec);
kfree(priv);
}
static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
.attach = vfio_pci_dma_buf_attach,
+#ifdef CONFIG_VFIO_PCI_DMABUF
+ .mmap = vfio_pci_dma_buf_mmap,
+#endif
.map_dma_buf = vfio_pci_dma_buf_map,
.unmap_dma_buf = vfio_pci_dma_buf_unmap,
.release = vfio_pci_dma_buf_release,
};
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long fault_addr,
+ unsigned int order,
+ unsigned long *out_pfn)
+{
+ /*
+ * Given a VMA (start, end, pgoffs) and a fault address,
+ * search the corresponding DMABUF's phys_vec[] to find the
+ * range representing the address's offset into the VMA, and
+ * its PFN. vdev must be the device that the DMABUF priv was
+ * exported from; vdev->dmabuf_lock must be held, and priv
+ * must not be revoked.
+ *
+ * The phys_vec[] ranges represent contiguous spans of VAs
+ * upwards from the buffer offset 0; the actual PFNs might be
+ * in any order, overlap/alias, etc. Calculate an offset of
+ * the desired page given VMA start/pgoff and address, then
+ * search upwards from 0 to find which span contains it.
+ *
+ * On success, a valid PFN for a page sized by 'order' is
+ * returned into out_pfn.
+ *
+ * Failure occurs if:
+ * - A hugepage would cross the edge of the VMA,
+ * - A hugepage isn't entirely contained within a range
+ * (including where it straddles the boundary between
+ * ranges),
+ * - We find a range, but the final PFN isn't aligned to the
+ * requested order.
+ *
+ * Upon failure, -ERANGE is returned and the caller is
+ * expected to try again with a smaller order, which will
+ * eventually succeed.
+ *
+ * It's suboptimal if DMABUFs are created with neighbouring
+ * ranges that are physically contiguous, since hugepages
+ * can't straddle range boundaries. (The construction of the
+ * ranges should merge them in this case.)
+ *
+ * Finally, vma_pgoff_adjust is used with a DMABUF created for
+ * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
+ * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
+ * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
+ * To avoid double-offsetting in this scenario, subtracting
+ * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
+ * the effective offset. This also removes the VFIO region
+ * index encoded in vm_pgoff for VFIO BAR mmaps.
+ */
+
+ const unsigned long pagesize = PAGE_SIZE << order;
+ unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) <<
+ PAGE_SHIFT;
+ unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
+ unsigned long rounded_page_end = rounded_page_addr + pagesize;
+ unsigned long fault_offset;
+ unsigned long fault_offset_end;
+ unsigned long range_start_offset = 0;
+ unsigned int i;
+ int ret;
+
+ if (unlikely(!vdev))
+ return -ENODEV;
+
+ /* This prevents the dmabuf revocation state from changing under us */
+ lockdep_assert_held(&vdev->dmabuf_lock);
+
+ if (unlikely(priv->vdev != vdev || priv->status != VFIO_PCI_DMABUF_OK))
+ return -ENODEV;
+
+ if (rounded_page_addr < vma->vm_start || rounded_page_end > vma->vm_end) {
+ if (order > 0)
+ return -ERANGE;
+
+ /* A fault address outside of the VMA is absurd. */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "Fault addr 0x%lx outside VMA 0x%lx-0x%lx\n",
+ fault_addr, vma->vm_start, vma->vm_end);
+ return -EFAULT;
+ }
+
+ /*
+ * fault_offset[_end] is the span within the DMABUF
+ * corresponding to the faulting page:
+ */
+ if (unlikely(check_add_overflow(rounded_page_addr - vma->vm_start,
+ vma_off, &fault_offset) ||
+ check_add_overflow(fault_offset, pagesize,
+ &fault_offset_end)))
+ return -EFAULT;
+
+ /*
+ * Iterate over ranges in the buffer, summing their lengths:
+ * range_start_offset represents the current range's starting
+ * offset in the buffer (from 0 upwards).
+ *
+ * A failure for order == 0 is unexpected, and triggers a
+ * fault/warn.
+ */
+ ret = (order == 0) ? -EFAULT : -ERANGE;
+
+ for (i = 0; i < priv->nr_ranges; i++) {
+ size_t range_len = priv->phys_vec[i].len;
+
+ /* Early exit if range starts after the page end */
+ if (fault_offset_end <= range_start_offset)
+ break;
+
+ if (fault_offset >= range_start_offset &&
+ fault_offset_end <= range_start_offset + range_len) {
+ /*
+ * The faulting page is wholly contained
+ * within the span represented by this range,
+ * so validate PFN alignment for the order.
+ * The if() condition ensures the pfn
+ * arithmetic won't overflow.
+ */
+ unsigned long pfn =
+ ((fault_offset - range_start_offset) +
+ priv->phys_vec[i].paddr) >> PAGE_SHIFT;
+
+ if (IS_ALIGNED(pfn, 1 << order)) {
+ *out_pfn = pfn;
+ ret = 0;
+ }
+ /*
+ * Else order > 0; ERANGE retries with smaller
+ * order
+ */
+ break;
+ }
+ range_start_offset += range_len;
+ }
+
+ if (order == 0 && ret != 0)
+ /*
+ * The address fell outside of the span represented by
+ * the (concatenated) ranges. As setup of a mapping
+ * ensures that the VMA is <= the total size of the
+ * ranges this should never happen. If it does, warn
+ * and SIGBUS.
+ */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\n",
+ fault_addr, order, vma->vm_start, vma->vm_end,
+ vma->vm_pgoff, priv->nr_ranges, priv->size);
+
+ return ret;
+}
+
+/*
+ * Create a DMABUF corresponding to priv, add it to vdev->dmabufs list
+ * for tracking (meaning cleanup or revocation will zap it), and take
+ * a vfio_device registration.
+ */
+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv, u32 flags)
+{
+ DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
+
+ if (!vfio_device_try_get_registration(&vdev->vdev))
+ return -ENODEV;
+
+ exp_info.ops = &vfio_pci_dmabuf_ops;
+ exp_info.size = priv->size;
+ exp_info.flags = flags;
+ exp_info.priv = priv;
+
+ priv->dmabuf = dma_buf_export(&exp_info);
+ if (IS_ERR(priv->dmabuf)) {
+ vfio_device_put_registration(&vdev->vdev);
+ return PTR_ERR(priv->dmabuf);
+ }
+
+ kref_init(&priv->kref);
+ init_completion(&priv->comp);
+
+ /* dma_buf_put() now frees priv */
+ INIT_LIST_HEAD(&priv->dmabufs_elm);
+
+ /*
+ * dmabuf_lock synchronises access (R) or updates (W) to the
+ * vdev->dmabufs list and to bars_revoked (see below). The
+ * revocation state of DMABUF elements in the list is written
+ * holding both dmabuf_lock(W) and resv, and tested with
+ * either.
+ *
+ * (memory_lock, if held ->) dmabuf_lock -> resv
+ *
+ * NOTE: memory_lock is strictly avoided here, to avoid a
+ * dependency on memory_lock when mmap_lock is held, when
+ * mmap() leads to export. vfio-pci variant drivers are
+ * permitted to hold memory_lock across actions that might
+ * fault (such as user access); a deadlock could result when
+ * that fault path attempts to take mmap_lock (if held by an
+ * export waiting for memory_lock).
+ *
+ * vdev->bars_revoked tracks the BAR revocation status updated
+ * via vfio_pci_dma_buf_move(), so the initial DMABUF state
+ * follows the same criteria that later update the DMABUF
+ * state (BAR zap, etc.).
+ */
+ lockdep_assert_not_held(&vdev->memory_lock);
+
+ down_write(&vdev->dmabuf_lock);
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ priv->status = vdev->bars_revoked ? VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK;
+ list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
+ dma_resv_unlock(priv->dmabuf->resv);
+ up_write(&vdev->dmabuf_lock);
+
+ return 0;
+}
+
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index)
+{
+ struct vfio_pci_dma_buf *priv;
+ unsigned long vma_pgoff = vma->vm_pgoff & (VFIO_PCI_OFFSET_MASK >> PAGE_SHIFT);
+ char *bufname;
+ int ret;
+
+ priv = kzalloc_obj(*priv);
+ if (!priv)
+ return -ENOMEM;
+
+ priv->phys_vec = kzalloc_obj(*priv->phys_vec);
+ if (!priv->phys_vec) {
+ ret = -ENOMEM;
+ goto err_free_priv;
+ }
+
+ /*
+ * Debug name: The absolute maximum size of the name
+ * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.
+ */
+ bufname = kasprintf(GFP_KERNEL, "vfio:%s/%x",
+ pci_name(vdev->pdev),
+ res_index);
+
+ if (!bufname) {
+ ret = -ENOMEM;
+ goto err_free_phys;
+ }
+
+ /*
+ * The DMABUF begins from the mmap()'s BAR offset, i.e. the
+ * start of the VMA corresponds to byte 0 of the DMABUF and
+ * byte (vma_pgoff << PAGE_SHIFT) of the BAR.
+ *
+ * vfio_pci_dma_buf_find_pfn() reverses this offset using
+ * vma_pgoff_adjust, so that ultimately a fault's offset from
+ * the start of the _VMA_ has a consistent usage whether the
+ * VMA originates from an mmap() of the VFIO device here or a
+ * direct DMABUF mmap(). Note vma_pgoff_adjust also includes
+ * the encoded VFIO region index, which cancels out the index
+ * encoded in vm_pgoff.
+ */
+ priv->vdev = vdev;
+ priv->size = req_len;
+ priv->nr_ranges = 1;
+ priv->vma_pgoff_adjust = vma->vm_pgoff;
+
+ /*
+ * The provider can be NULL _iff_ the DMABUF feature isn't
+ * supported, because it's only used by DMABUF import and
+ * attach is prohibited if the feature isn't present.
+ */
+ priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
+ if (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) && !priv->provider) {
+ ret = -EINVAL;
+ goto err_free_name;
+ }
+
+ priv->phys_vec[0].paddr = phys_start + ((u64)vma_pgoff << PAGE_SHIFT);
+ priv->phys_vec[0].len = priv->size;
+
+ ret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);
+ if (ret)
+ goto err_free_name;
+
+ if (dma_buf_set_name(priv->dmabuf, bufname)) {
+ dev_dbg_ratelimited(&vdev->pdev->dev,
+ "Failed to set map name '%s'\n",
+ bufname);
+ kfree(bufname);
+ }
+
+ /*
+ * Ownership of the DMABUF file transfers to the VMA so that
+ * other users can locate the DMABUF via a VA. Ownership of
+ * the original VFIO device file being mmap()ed transfers to
+ * priv, and is put when the DMABUF is released. This
+ * intentionally does not use get_file()/vma_set_file()
+ * because the references are already held, and ownership
+ * moves.
+ */
+ priv->vfile = vma->vm_file;
+ vma->vm_file = priv->dmabuf->file;
+ vma->vm_private_data = priv;
+
+ return 0;
+
+err_free_name:
+ kfree(bufname);
+err_free_phys:
+ kfree(priv->phys_vec);
+err_free_priv:
+ kfree(priv);
+ return ret;
+}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
/*
* This is a temporary "private interconnect" between VFIO DMABUF and iommufd.
* It allows the two co-operating drivers to exchange the physical address of
@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
return -EOPNOTSUPP;
priv = attachment->dmabuf->priv;
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return -ENODEV;
/* More than one range to iommufd will require proper DMABUF support */
@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
struct vfio_device_feature_dma_buf get_dma_buf = {};
struct vfio_region_dma_range *dma_ranges;
- DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
struct vfio_pci_dma_buf *priv;
size_t length;
int ret;
@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
dma_ranges = NULL;
- if (!vfio_device_try_get_registration(&vdev->vdev)) {
- ret = -ENODEV;
+ ret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);
+ if (ret)
goto err_free_phys;
- }
-
- exp_info.ops = &vfio_pci_dmabuf_ops;
- exp_info.size = priv->size;
- exp_info.flags = get_dma_buf.open_flags;
- exp_info.priv = priv;
-
- priv->dmabuf = dma_buf_export(&exp_info);
- if (IS_ERR(priv->dmabuf)) {
- ret = PTR_ERR(priv->dmabuf);
- goto err_dev_put;
- }
-
- kref_init(&priv->kref);
- init_completion(&priv->comp);
-
- /* dma_buf_put() now frees priv */
- INIT_LIST_HEAD(&priv->dmabufs_elm);
- down_write(&vdev->memory_lock);
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = !__vfio_pci_memory_enabled(vdev);
- list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
- dma_resv_unlock(priv->dmabuf->resv);
- up_write(&vdev->memory_lock);
-
/*
* dma_buf_fd() consumes the reference, when the file closes the dmabuf
* will be released.
@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
return ret;
-err_dev_put:
- vfio_device_put_registration(&vdev->vdev);
err_free_phys:
kfree(priv->phys_vec);
err_free_priv:
@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
return ret;
}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
+/*
+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives
+ * the guarantee that all future map/attach attempts will fail no
+ * matter what, whereas REVOKED can transition back to OK.
+ */
+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,
+ enum vfio_pci_dma_buf_status new_status)
+{
+ bool was_revoked;
+
+ /*
+ * Changes to the DMABUF's revocation status are synchronised
+ * using dmabuf_lock:
+ */
+ lockdep_assert_held_write(&priv->vdev->dmabuf_lock);
+
+ /* If DEAD, state can no longer change */
+ if (priv->status == VFIO_PCI_DMABUF_DEAD ||
+ priv->status == new_status)
+ return;
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ was_revoked = (priv->status == VFIO_PCI_DMABUF_REVOKED);
+
+ if (new_status != VFIO_PCI_DMABUF_OK) {
+ priv->status = new_status;
+
+ if (was_revoked) {
+ /*
+ * A REVOKED buffer is being marked DEAD.
+ * invalidate_mappings/unmap wait happened
+ * when it became REVOKED, don't wait again.
+ */
+ dma_resv_unlock(priv->dmabuf->resv);
+ return;
+ }
+ dma_buf_invalidate_mappings(priv->dmabuf);
+ dma_resv_wait_timeout(priv->dmabuf->resv,
+ DMA_RESV_USAGE_BOOKKEEP, false,
+ MAX_SCHEDULE_TIMEOUT);
+ dma_resv_unlock(priv->dmabuf->resv);
+ kref_put(&priv->kref, vfio_pci_dma_buf_done);
+ wait_for_completion(&priv->comp);
+ unmap_mapping_range(priv->dmabuf->file->f_mapping,
+ 0, 0, true);
+ /*
+ * Re-arm the registered kref reference and the
+ * completion so the post-revoke state matches the
+ * post-creation state. An un-revoke followed by a
+ * new mapping needs the kref to be non-zero before
+ * kref_get(), and vfio_pci_dma_buf_cleanup()
+ * delegates its drain back through this revoke
+ * path on a possibly-already-revoked dma-buf.
+ */
+ kref_init(&priv->kref);
+ reinit_completion(&priv->comp);
+ } else {
+ priv->status = VFIO_PCI_DMABUF_OK;
+ dma_resv_unlock(priv->dmabuf->resv);
+ }
+}
void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
{
@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
lockdep_assert_held_write(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ vdev->bars_revoked = revoked;
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
-
- if (priv->revoked != revoked) {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- if (revoked)
- priv->revoked = true;
- dma_buf_invalidate_mappings(priv->dmabuf);
- dma_resv_wait_timeout(priv->dmabuf->resv,
- DMA_RESV_USAGE_BOOKKEEP, false,
- MAX_SCHEDULE_TIMEOUT);
- dma_resv_unlock(priv->dmabuf->resv);
- if (revoked) {
- kref_put(&priv->kref, vfio_pci_dma_buf_done);
- wait_for_completion(&priv->comp);
- /*
- * Re-arm the registered kref reference and the
- * completion so the post-revoke state matches the
- * post-creation state. An un-revoke followed by a
- * new mapping needs the kref to be non-zero before
- * kref_get(), and vfio_pci_dma_buf_cleanup()
- * delegates its drain back through this revoke
- * path on a possibly-already-revoked dma-buf.
- */
- kref_init(&priv->kref);
- reinit_completion(&priv->comp);
- } else {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = false;
- dma_resv_unlock(priv->dmabuf->resv);
- }
- }
+ vfio_pci_dma_buf_set_status(priv, revoked ?
+ VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
}
void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
*/
vfio_pci_dma_buf_move(vdev, true);
+ down_write(&vdev->dmabuf_lock);
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
list_del_init(&priv->dmabufs_elm);
- priv->vdev = NULL;
+ WRITE_ONCE(priv->vdev, NULL);
vfio_device_put_registration(&vdev->vdev);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
up_write(&vdev->memory_lock);
}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
+{
+ struct vfio_device_feature_dma_buf_revoke db_revoke;
+ struct vfio_pci_dma_buf *priv;
+ struct dma_buf *dmabuf;
+ int ret;
+
+ if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
+ return -EOPNOTSUPP;
+
+ ret = vfio_check_feature(flags, argsz,
+ VFIO_DEVICE_FEATURE_SET,
+ sizeof(db_revoke));
+ if (ret != 1)
+ return ret;
+
+ if (copy_from_user(&db_revoke, arg, sizeof(db_revoke)))
+ return -EFAULT;
+
+ dmabuf = dma_buf_get(db_revoke.dmabuf_fd);
+ if (IS_ERR(dmabuf))
+ return PTR_ERR(dmabuf);
+
+ priv = dmabuf->priv;
+ /*
+ * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_
+ * relates to the VFIO device it was provided with.
+ *
+ * If the DMABUF relates to this vdev then priv->vdev is
+ * stable because this open fd prevents cleanup.
+ *
+ * If it relates to a different vdev, reading priv->vdev might
+ * race with a concurrent cleanup on that device. But if so,
+ * it points to a non-matching vdev or NULL and is unusable
+ * either way.
+ */
+ if (dmabuf->ops != &vfio_pci_dmabuf_ops ||
+ READ_ONCE(priv->vdev) != vdev) {
+ ret = -ENODEV;
+ goto out_put_buf;
+ }
+
+ /*
+ * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()
+ * from getting it and then blocking all devices in the dev_set behind
+ * this revoke's drain.
+ */
+ down_read(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_DEAD) {
+ ret = -EBADFD;
+ } else {
+ vfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);
+ ret = 0;
+ }
+ up_write(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
+
+out_put_buf:
+ dma_buf_put(dmabuf);
+
+ return ret;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 4e7162234a2eb..ca12221af5553 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {
bool test_mem;
};
+enum vfio_pci_dma_buf_status {
+ VFIO_PCI_DMABUF_OK = 0,
+ VFIO_PCI_DMABUF_REVOKED = 1,
+ VFIO_PCI_DMABUF_DEAD = 2,
+};
+
+struct vfio_pci_dma_buf {
+ struct dma_buf *dmabuf;
+ struct vfio_pci_core_device *vdev;
+ struct list_head dmabufs_elm;
+ size_t size;
+ struct phys_vec *phys_vec;
+ struct p2pdma_provider *provider;
+ struct file *vfile;
+ u32 nr_ranges;
+ struct kref kref;
+ struct completion comp;
+ unsigned long vma_pgoff_adjust;
+ enum vfio_pci_dma_buf_status status;
+};
+
bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);
void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);
@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);
int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state);
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
u16 cmd);
@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)
return (pdev->class >> 8) == PCI_CLASS_DISPLAY_VGA;
}
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long address,
+ unsigned int order,
+ unsigned long *out_pfn);
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index);
+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);
+
#ifdef CONFIG_VFIO_PCI_DMABUF
int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
struct vfio_device_feature_dma_buf __user *arg,
size_t argsz);
-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz);
#else
static inline int
vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
return -ENOTTY;
}
-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
-{
-}
-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,
- bool revoked)
+static inline int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
{
+ return -ENOTTY;
}
#endif
diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
index d15b2b31d3c91..952a2c196ad42 100644
--- a/include/linux/dma-buf.h
+++ b/include/linux/dma-buf.h
@@ -571,6 +571,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);
struct dma_buf *dma_buf_get(int fd);
void dma_buf_put(struct dma_buf *dmabuf);
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);
+
struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,
enum dma_data_direction);
void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa2..44891fdb7c76e 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -129,11 +129,13 @@ struct vfio_pci_core_device {
bool disable_idle_d3:1;
bool nointxmask:1;
bool disable_vga:1;
+ bool zap_bars_on_revoke:1;
/* Flags modified at runtime - dedicated storage unit */
bool needs_reset;
bool pm_intx_masked;
bool pm_runtime_engaged;
bool sriov_active;
+ bool bars_revoked;
struct pci_saved_state *pci_saved_state;
struct pci_saved_state *pm_save;
int ioeventfds_nr;
@@ -148,6 +150,7 @@ struct vfio_pci_core_device {
struct vfio_pci_core_device *sriov_pf_core_dev;
struct notifier_block nb;
struct rw_semaphore memory_lock;
+ struct rw_semaphore dmabuf_lock;
struct list_head dmabufs;
};
diff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h
index e41437fa17ad0..d3c6057983e09 100644
--- a/include/uapi/linux/vfio.h
+++ b/include/uapi/linux/vfio.h
@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {
#define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
+/**
+ * Given a DMABUF fd previously exported from the same device by
+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that
+ * access to the corresponding DMABUF is immediately revoked. On
+ * successful return, the buffer is no longer accessible through any
+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future
+ * mmap()s and map/attach requests from any new/existing importer.
+ *
+ * Return: 0 on success, -1 and errno is set on failure:
+ *
+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.
+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.
+ * ENODEV: The DMABUF was not exported from this device.
+ * EBADFD: The DMABUF is already revoked by this feature.
+ *
+ * Additionally, common errors can occur: EFAULT accessing the struct,
+ * or EINVAL requesting an unsupported feature op.
+ */
+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14
+
+struct vfio_device_feature_dma_buf_revoke {
+ __s32 dmabuf_fd;
+};
+
/* -------- API for Type1 VFIO IOMMU -------- */
/**
|
| 49/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 23413 output: 20 thoughts: 2204 Thoughts: **KMSAN Fuzzing Analysis: VFIO DMA-BUF Patch Series**
Okay, let's break this down. The task is to determine if this VFIO DMA-BUF patch series warrants a dedicated KMSAN fuzzing session. The key here is to determine if uninitialized memory reads or information leaks are being introduced. I have to apply the rules: if KASAN, LOCKDEP, or standard detectors can catch it, then no KMSAN is needed. KMSAN is for the hard stuff: *only* uninitialized memory bugs. I also need to look holistically at the surrounding code.
From what I can tell, the patch series introduces a new DMA-BUF mechanism within the VFIO/PCI framework. It appears that `dma_buf` is used for sharing device memory between the host and guest, and there are many calls into `vfio_pci_dma_buf` and related functions.
My initial analysis reveals the following:
* **`drivers/dma-buf/dma-buf.c`**: The change in `dma_buf_set_name` seems safe. The `strndup_user` function, which is now called in `dma_buf_ioctl`, copies a string from userspace and null-terminates it. The subsequent operations use standard functions, and I see no risk of uninitialized memory reads or information leaks there. I confirmed this using strndup_user's man page as well.
* **`drivers/vfio/pci/*`**: This is where it gets interesting. I need to make sure I am fully aware of the context.
* The `vfio_pci_core_mmap_prep_dmabuf` function looks important. It allocates a `vfio_pci_dma_buf` structure and `phys_vec`. Crucially, both are `kzalloc`'d. This means that they *are* initialized to zero, which makes me think there is no chance that there would be uninitialized data. The code then populates `priv` and `priv->phys_vec` with derived values based on the physical address and size of the buffer. The function calls `vfio_pci_dmabuf_export` and, importantly, `dma_buf_set_name`. However, I've already determined this function to be safe.
* `vfio_pci_core_feature_dma_buf_revoke` takes a `db_revoke` struct from userspace, which I checked, and sees if `VFIO_DEVICE_FEATURE_SET` applies to it. I looked for `copy_to_user` calls to ensure that no data goes back to userspace, and it seems this one is safe in that respect. I used grep for this and found no instances. This `VFIO_DEVICE_FEATURE_SET` seems important; I might need to look into it to make sure I am not missing anything. I will need to clarify the implications of `VFIO_DEVICE_FEATURE_SET` for completeness, but it is unlikely to change the results.
Given the code so far:
* Most of the new structures are `kzalloc`'d, which means they are zero-initialized.
* The code that handles names uses `strndup_user`, which handles strings.
* No `copy_to_user` calls are obvious that would lead to information leaks.
Therefore, my initial inclination is that this patch series does *not* need a dedicated KMSAN fuzzing session. The primary risks appear to be bounds checks, NULL pointer dereferences, or race conditions, which are caught by existing detectors. However, I want to confirm this by going through the code for the `VFIO_DEVICE_FEATURE_SET`. I need to review all paths that go through the new code, but right now, I have little doubt.
**Conclusion: NeedsKMSAN=false**
|
| 50/3 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "vfio_check_feature"
}Results: {
"Output": "drivers/vfio/cdx/main.c=84=static int vfio_cdx_bm_ctrl(struct vfio_device *core_vdev, u32 flags,\n--\ndrivers/vfio/cdx/main.c-97-\ndrivers/vfio/cdx/main.c:98:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/cdx/main.c-99-\t\t\t\t sizeof(ops));\n--\ndrivers/vfio/pci/vfio_pci_core.c=396=static int vfio_pci_core_pm_entry(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_core.c-400-\ndrivers/vfio/pci/vfio_pci_core.c:401:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/pci/vfio_pci_core.c-402-\tif (ret != 1)\n--\ndrivers/vfio/pci/vfio_pci_core.c=414=static int vfio_pci_core_pm_entry_with_wakeup(\n--\ndrivers/vfio/pci/vfio_pci_core.c-422-\ndrivers/vfio/pci/vfio_pci_core.c:423:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/pci/vfio_pci_core.c-424-\t\t\t\t sizeof(entry));\n--\ndrivers/vfio/pci/vfio_pci_core.c=471=static int vfio_pci_core_pm_exit(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_core.c-475-\ndrivers/vfio/pci/vfio_pci_core.c:476:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/pci/vfio_pci_core.c-477-\tif (ret != 1)\n--\ndrivers/vfio/pci/vfio_pci_core.c=1598=static int vfio_pci_core_feature_token(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1610-\t */\ndrivers/vfio/pci/vfio_pci_core.c:1611:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/pci/vfio_pci_core.c-1612-\t\t\t\t sizeof(uuid));\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=590=int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-602-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:603:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/pci/vfio_pci_dmabuf.c-604-\t\t\t\t sizeof(get_dma_buf));\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=791=int vfio_pci_core_feature_dma_buf_revoke(\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-803-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:804:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/pci/vfio_pci_dmabuf.c-805-\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n--\ndrivers/vfio/pci/vfio_pci_zdev.c=145=int vfio_pci_zdev_feature_err(struct vfio_device *device, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_zdev.c-158-\ndrivers/vfio/pci/vfio_pci_zdev.c:159:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/pci/vfio_pci_zdev.c-160-\t\t\t\t sizeof(err));\n--\ndrivers/vfio/vfio_main.c=931=vfio_ioctl_device_feature_mig_device_state(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-943-\ndrivers/vfio/vfio_main.c:944:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-945-\t\t\t\t VFIO_DEVICE_FEATURE_SET |\n--\ndrivers/vfio/vfio_main.c=981=vfio_ioctl_device_feature_migration_data_size(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-991-\ndrivers/vfio/vfio_main.c:992:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/vfio_main.c-993-\t\t\t\t sizeof(data_size));\n--\ndrivers/vfio/vfio_main.c=1009=vfio_ioctl_device_feature_migration_precopy_info_v2(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1016-\ndrivers/vfio/vfio_main.c:1017:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/vfio_main.c-1018-\tif (ret != 1)\n--\ndrivers/vfio/vfio_main.c=1025=static int vfio_ioctl_device_feature_migration(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1036-\ndrivers/vfio/vfio_main.c:1037:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/vfio_main.c-1038-\t\t\t\t sizeof(mig));\n--\ndrivers/vfio/vfio_main.c=1108=vfio_ioctl_device_feature_logging_start(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1126-\ndrivers/vfio/vfio_main.c:1127:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1128-\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n--\ndrivers/vfio/vfio_main.c=1192=vfio_ioctl_device_feature_logging_stop(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1200-\ndrivers/vfio/vfio_main.c:1201:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1202-\t\t\t\t VFIO_DEVICE_FEATURE_SET, 0);\n--\ndrivers/vfio/vfio_main.c=1219=vfio_ioctl_device_feature_logging_report(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1233-\ndrivers/vfio/vfio_main.c:1234:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1235-\t\t\t\t VFIO_DEVICE_FEATURE_GET,\n--\ninclude/linux/vfio.h=250=struct vfio_log_ops {\n--\ninclude/linux/vfio.h-259-/**\ninclude/linux/vfio.h:260: * vfio_check_feature - Validate user input for the VFIO_DEVICE_FEATURE ioctl\ninclude/linux/vfio.h-261- * @flags: Arg from the device_feature op\n--\ninclude/linux/vfio.h-271- */\ninclude/linux/vfio.h:272:static inline int vfio_check_feature(u32 flags, size_t argsz, u32 supported_ops,\ninclude/linux/vfio.h-273-\t\t\t\t size_t minsz)\n"
} |
| 51/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 28783 output: 31 thoughts: 38 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Investigating DMA-BUF Name Handling** Alright, so I'm diving into DMA-BUF naming conventions again. I want to quickly ascertain exactly *how* these `dmabuf->name` fields are being utilized within the kernel. My immediate thought is to leverage `grepper` – it’s the most efficient tool for this kind of code archaeology, especially when dealing with such a fundamental data structure. Let's see exactly where and how this name is being accessed and used throughout the kernel source tree. I’m expecting a high number of hits, but I'm looking for the *patterns* of usage: reads, comparisons, assignments. What are the common operations done on the name and where? |
| 52/3 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "-\u003ename\\b",
"PathPrefix": "drivers/dma-buf/"
}Results: {
"Output": "drivers/dma-buf/dma-buf.c-39-/*\ndrivers/dma-buf/dma-buf.c:40: * dmabuf-\u003ename must be accessed with holding dmabuf-\u003ename_lock.\ndrivers/dma-buf/dma-buf.c-41- * we need to take the lock around the tracepoint call itself where\n--\ndrivers/dma-buf/dma-buf.c=150=static char *dmabuffs_dname(struct dentry *dentry, char *buffer, int buflen)\n--\ndrivers/dma-buf/dma-buf.c-157-\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:158:\tif (dmabuf-\u003ename)\ndrivers/dma-buf/dma-buf.c:159:\t\tret = strscpy(name, dmabuf-\u003ename, sizeof(name));\ndrivers/dma-buf/dma-buf.c-160-\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n--\ndrivers/dma-buf/dma-buf.c=166=static void dma_buf_release(struct dentry *dentry)\n--\ndrivers/dma-buf/dma-buf.c-189-\tmodule_put(dmabuf-\u003eowner);\ndrivers/dma-buf/dma-buf.c:190:\tkfree(dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c-191-\tkfree(dmabuf);\n--\ndrivers/dma-buf/dma-buf.c=578=static void dma_buf_show_fdinfo(struct seq_file *m, struct file *file)\n--\ndrivers/dma-buf/dma-buf.c-586-\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:587:\tif (dmabuf-\u003ename)\ndrivers/dma-buf/dma-buf.c:588:\t\tseq_printf(m, \"name:\\t%s\\n\", dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c-589-\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n--\ndrivers/dma-buf/dma-buf.c=875=int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\n--\ndrivers/dma-buf/dma-buf.c-886-\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:887:\tkfree(dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c:888:\tdmabuf-\u003ename = name;\ndrivers/dma-buf/dma-buf.c-889-\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n--\ndrivers/dma-buf/dma-buf.c=1771=static int dma_buf_debug_show(struct seq_file *s, void *unused)\n--\ndrivers/dma-buf/dma-buf.c-1801-\t\t\t\tfile_inode(buf_obj-\u003efile)-\u003ei_ino,\ndrivers/dma-buf/dma-buf.c:1802:\t\t\t\tbuf_obj-\u003ename ?: \"\u003cnone\u003e\");\ndrivers/dma-buf/dma-buf.c-1803-\t\tspin_unlock(\u0026buf_obj-\u003ename_lock);\n--\ndrivers/dma-buf/dma-heap.c=222=const char *dma_heap_get_name(struct dma_heap *heap)\ndrivers/dma-buf/dma-heap.c-223-{\ndrivers/dma-buf/dma-heap.c:224:\treturn heap-\u003ename;\ndrivers/dma-buf/dma-heap.c-225-}\n--\ndrivers/dma-buf/dma-heap.c=232=struct dma_heap *dma_heap_add(const struct dma_heap_export_info *exp_info)\n--\ndrivers/dma-buf/dma-heap.c-238-\ndrivers/dma-buf/dma-heap.c:239:\tif (!exp_info-\u003ename || !strcmp(exp_info-\u003ename, \"\")) {\ndrivers/dma-buf/dma-heap.c-240-\t\tpr_err(\"dma_heap: Cannot add heap without a name\\n\");\n--\ndrivers/dma-buf/dma-heap.c-252-\ndrivers/dma-buf/dma-heap.c:253:\theap-\u003ename = exp_info-\u003ename;\ndrivers/dma-buf/dma-heap.c-254-\theap-\u003eops = exp_info-\u003eops;\n--\ndrivers/dma-buf/dma-heap.c-280-\t\t\t\tNULL,\ndrivers/dma-buf/dma-heap.c:281:\t\t\t\theap-\u003ename);\ndrivers/dma-buf/dma-heap.c-282-\tif (IS_ERR(dev_ret)) {\n--\ndrivers/dma-buf/dma-heap.c-290-\tlist_for_each_entry(h, \u0026heap_list, list) {\ndrivers/dma-buf/dma-heap.c:291:\t\tif (!strcmp(h-\u003ename, exp_info-\u003ename)) {\ndrivers/dma-buf/dma-heap.c-292-\t\t\tmutex_unlock(\u0026heap_list_lock);\ndrivers/dma-buf/dma-heap.c-293-\t\t\tpr_err(\"dma_heap: Already registered heap named %s\\n\",\ndrivers/dma-buf/dma-heap.c:294:\t\t\t exp_info-\u003ename);\ndrivers/dma-buf/dma-heap.c-295-\t\t\terr_ret = ERR_PTR(-EINVAL);\n--\ndrivers/dma-buf/sw_sync.c=100=static struct sync_timeline *sync_timeline_create(const char *name)\n--\ndrivers/dma-buf/sw_sync.c-109-\tobj-\u003econtext = dma_fence_context_alloc(1);\ndrivers/dma-buf/sw_sync.c:110:\tstrscpy(obj-\u003ename, name, sizeof(obj-\u003ename));\ndrivers/dma-buf/sw_sync.c-111-\n--\ndrivers/dma-buf/sw_sync.c=146=static const char *timeline_fence_get_timeline_name(struct dma_fence *fence)\n--\ndrivers/dma-buf/sw_sync.c-149-\ndrivers/dma-buf/sw_sync.c:150:\treturn parent-\u003ename;\ndrivers/dma-buf/sw_sync.c-151-}\n--\ndrivers/dma-buf/sync_debug.c=45=static void sync_print_fence(struct seq_file *s,\n--\ndrivers/dma-buf/sync_debug.c-53-\tseq_printf(s, \" %s%sfence %s\",\ndrivers/dma-buf/sync_debug.c:54:\t\t show ? parent-\u003ename : \"\",\ndrivers/dma-buf/sync_debug.c-55-\t\t show ? \"_\" : \"\",\n--\ndrivers/dma-buf/sync_debug.c=70=static void sync_print_obj(struct seq_file *s, struct sync_timeline *obj)\n--\ndrivers/dma-buf/sync_debug.c-73-\ndrivers/dma-buf/sync_debug.c:74:\tseq_printf(s, \"%s: %d\\n\", obj-\u003ename, obj-\u003evalue);\ndrivers/dma-buf/sync_debug.c-75-\n--\ndrivers/dma-buf/sync_trace.h=12=TRACE_EVENT(sync_timeline,\n--\ndrivers/dma-buf/sync_trace.h-17-\tTP_STRUCT__entry(\ndrivers/dma-buf/sync_trace.h:18:\t\t\t__string(name, timeline-\u003ename)\ndrivers/dma-buf/sync_trace.h-19-\t\t\t__field(u32, value)\n"
} |
| 53/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31108 output: 28 thoughts: 2040 Thoughts: **Deep Dive into DMA Buffer Handling in VFIO-PCI** Alright, let's break down this DMA buffer stuff in VFIO-PCI, starting with `dma_buf_set_name`. The core of this function is pretty straightforward: it sets a name for a DMA buffer. There's a crucial length check here, `DMA_BUF_NAME_LEN`. If the provided name is too long, we get `-E2BIG`; otherwise, it updates the `dmabuf->name`. We see this function being used in `dma_buf_ioctl`, where `strndup_user` ensures the name is null-terminated and no longer than `DMA_BUF_NAME_LEN`. Interestingly, we also see it used in `vfio_pci_core_mmap_prep_dmabuf` where `kasprintf` generates the buffer name. Note the potential for a name collision. If `kasprintf` fails, `bufname` is `NULL`. If `dma_buf_set_name` fails, it frees `bufname`. Now, let's jump to `vfio_pci_core_mmap_prep_dmabuf`, which is where things get interesting. We're setting up a DMA buffer for a PCI device. The `kzalloc_obj` calls zero out memory. Crucially, the code then populates a `vfio_pci_dma_buf` struct, along with `phys_vec`. I need to ensure that those zeroed fields are not causing problems. The initialization of various fields like `size`, `nr_ranges`, `vma_pgoff_adjust`, and the association with the `p2pdma_provider` and the VMA is all critical. The `priv->phys_vec` contains the physical address and length. Note also that the VMA's file is replaced with the DMA buffer's file, which looks like a pointer swap. The `vfio_pci_dmabuf_export` function is where the DMA buffer is actually exported. `dma_buf_export` creates the actual `dmabuf` and associates it with our `priv`. We see the initialization of `kref`, `comp`, and `dmabufs_elm`, and the setting of the `status` based on `vdev->bars_revoked`. Lastly, let's scrutinize `vfio_pci_core_feature_dma_buf_revoke`. This seems to be the function that handles revoking DMA buffers. The call to `vfio_check_feature` is good, but `copy_from_user` indicates that we are *reading* data from userspace, which is as expected for a revoke operation. I also see that `vfio_pci_core_ioctl_feature` has several feature requests that it handles, and then the call to `vfio_ioctl_device_feature` is key. I need to understand that call's context to confirm there are no vulnerabilities here. I'll need to double check how these pieces fit together, especially the VMA and DMA buffer associations, and ensure that revoking doesn't have race conditions or access violations. |
| 54/3 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "ioctl_feature",
"PathPrefix": "drivers/vfio/"
}Results: {
"Output": "drivers/vfio/cdx/main.c=84=static int vfio_cdx_bm_ctrl(struct vfio_device *core_vdev, u32 flags,\n--\ndrivers/vfio/cdx/main.c-115-\ndrivers/vfio/cdx/main.c:116:static int vfio_cdx_ioctl_feature(struct vfio_device *device, u32 flags,\ndrivers/vfio/cdx/main.c-117-\t\t\t\t void __user *arg, size_t argsz)\n--\ndrivers/vfio/cdx/main.c=291=static const struct vfio_device_ops vfio_cdx_ops = {\n--\ndrivers/vfio/cdx/main.c-298-\t.get_region_info_caps = vfio_cdx_ioctl_get_region_info,\ndrivers/vfio/cdx/main.c:299:\t.device_feature = vfio_cdx_ioctl_feature,\ndrivers/vfio/cdx/main.c-300-\t.mmap\t\t= vfio_cdx_mmap,\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1593=static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1600-\t.get_region_info_caps = hisi_acc_vfio_ioctl_get_region,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1601:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1602-\t.read = hisi_acc_vfio_pci_read,\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1614=static const struct vfio_device_ops hisi_acc_vfio_pci_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1621-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1622:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1623-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/ism/main.c=335=static const struct vfio_device_ops ism_pci_ops = {\n--\ndrivers/vfio/pci/ism/main.c-342-\t.get_region_info_caps = ism_vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/ism/main.c:343:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/ism/main.c-344-\t.read = ism_vfio_pci_read,\n--\ndrivers/vfio/pci/mlx5/main.c=1386=static const struct vfio_device_ops mlx5vf_pci_ops = {\n--\ndrivers/vfio/pci/mlx5/main.c-1393-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/mlx5/main.c:1394:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/mlx5/main.c-1395-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1065=static const struct vfio_device_ops nvgrace_gpu_pci_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1072-\t.get_region_info_caps = nvgrace_gpu_ioctl_get_region_info,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1073:\t.device_feature\t= vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1074-\t.read\t\t= nvgrace_gpu_read,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1090=static const struct vfio_device_ops nvgrace_gpu_pci_core_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1097-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1098:\t.device_feature\t= vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1099-\t.read\t\t= vfio_pci_core_read,\n--\ndrivers/vfio/pci/pds/vfio_dev.c=191=static const struct vfio_device_ops pds_vfio_ops = {\n--\ndrivers/vfio/pci/pds/vfio_dev.c-198-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/pds/vfio_dev.c:199:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/pds/vfio_dev.c-200-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/vfio_pci.c=149=static const struct vfio_device_ops vfio_pci_ops = {\n--\ndrivers/vfio/pci/vfio_pci.c-156-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/vfio_pci.c:157:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/vfio_pci.c-158-\t.read\t\t= vfio_pci_core_read,\n--\ndrivers/vfio/pci/vfio_pci_core.c=1598=static int vfio_pci_core_feature_token(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1624-\ndrivers/vfio/pci/vfio_pci_core.c:1625:int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,\ndrivers/vfio/pci/vfio_pci_core.c-1626-\t\t\t\tvoid __user *arg, size_t argsz)\n--\ndrivers/vfio/pci/vfio_pci_core.c-1650-}\ndrivers/vfio/pci/vfio_pci_core.c:1651:EXPORT_SYMBOL_GPL(vfio_pci_core_ioctl_feature);\ndrivers/vfio/pci/vfio_pci_core.c-1652-\n--\ndrivers/vfio/pci/virtio/main.c=84=static const struct vfio_device_ops virtiovf_vfio_pci_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-91-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:92:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-93-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/virtio/main.c=106=static const struct vfio_device_ops virtiovf_vfio_pci_tran_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-113-\t.get_region_info_caps = virtiovf_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:114:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-115-\t.read = virtiovf_pci_core_read,\n--\ndrivers/vfio/pci/virtio/main.c=128=static const struct vfio_device_ops virtiovf_vfio_pci_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-135-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:136:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-137-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/xe/main.c=522=static const struct vfio_device_ops xe_vfio_pci_ops = {\n--\ndrivers/vfio/pci/xe/main.c-529-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/xe/main.c:530:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/xe/main.c-531-\t.read = vfio_pci_core_read,\n"
} |
| 55/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36482 output: 34 thoughts: 28 Thoughts: Here's a summary of my thought process, as requested: **Diving into VFIO Device Feature Control** Alright, let's get down to brass tacks. I need to understand how the VFIO subsystem handles device features, specifically through the ioctl interface. My instinct is to head straight for `drivers/vfio/vfio_main.c`. That's where I'd expect to find the central handling of VFIO's core functionality. My focus is on the `vfio_ioctl_device_feature` ioctl handler. I need to dissect this code to understand how features are queried, enabled, disabled, or generally managed at the device level from userspace. I'll be looking for things like: how the driver identifies and indexes the feature being targeted, what data structures are used to represent these features internally, what the security considerations are around managing a potentially sensitive operation, and how the kernel interacts with the underlying hardware to enact these feature changes. Let's see what we've got. |
| 56/3 | 2026/09/26 21:08 | tool | grepper |
0mArgs:{
"Expression": "vfio_ioctl_device_feature\\(",
"PathPrefix": "drivers/vfio/"
}Results: {
"Output": "drivers/vfio/vfio_main.c=1219=vfio_ioctl_device_feature_logging_report(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1262-\ndrivers/vfio/vfio_main.c:1263:static int vfio_ioctl_device_feature(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1264-\t\t\t\t struct vfio_device_feature __user *arg)\n--\ndrivers/vfio/vfio_main.c=1369=static long vfio_device_fops_unl_ioctl(struct file *filep,\n--\ndrivers/vfio/vfio_main.c-1402-\tcase VFIO_DEVICE_FEATURE:\ndrivers/vfio/vfio_main.c:1403:\t\tret = vfio_ioctl_device_feature(device, uptr);\ndrivers/vfio/vfio_main.c-1404-\t\tbreak;\n"
} |
| 57/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37024 output: 40 thoughts: 25 Thoughts: Okay, let's dive into this. Here's my take on those instructions, as if they're my train of thought: **Diving into vfio_main.c** Alright, time to get my hands dirty. It looks like we're homing in on the `drivers/vfio/vfio_main.c` file. Specifically, I need to focus my attention around line 1263. That's usually where the core of some important logic lies, given the location within a VFIO main source file. I anticipate there's something critical there that's relevant to whatever I'm currently investigating – likely related to device assignment, resource management, or perhaps the interrupt handling flow. Let's see what we can find. |
| 58/3 | 2026/09/26 21:08 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/vfio_main.c",
"FirstLine": 1260,
"LineCount": 40
}Results: {
"Contents": "1260:\t\treturn ret;\n1261:\t}\n1262:\t\n1263:\tstatic int vfio_ioctl_device_feature(struct vfio_device *device,\n1264:\t\t\t\t\t struct vfio_device_feature __user *arg)\n1265:\t{\n1266:\t\tsize_t minsz = offsetofend(struct vfio_device_feature, flags);\n1267:\t\tstruct vfio_device_feature feature;\n1268:\t\n1269:\t\tif (copy_from_user(\u0026feature, arg, minsz))\n1270:\t\t\treturn -EFAULT;\n1271:\t\n1272:\t\tif (feature.argsz \u003c minsz)\n1273:\t\t\treturn -EINVAL;\n1274:\t\n1275:\t\t/* Check unknown flags */\n1276:\t\tif (feature.flags \u0026\n1277:\t\t ~(VFIO_DEVICE_FEATURE_MASK | VFIO_DEVICE_FEATURE_SET |\n1278:\t\t VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_PROBE))\n1279:\t\t\treturn -EINVAL;\n1280:\t\n1281:\t\t/* GET \u0026 SET are mutually exclusive except with PROBE */\n1282:\t\tif (!(feature.flags \u0026 VFIO_DEVICE_FEATURE_PROBE) \u0026\u0026\n1283:\t\t (feature.flags \u0026 VFIO_DEVICE_FEATURE_SET) \u0026\u0026\n1284:\t\t (feature.flags \u0026 VFIO_DEVICE_FEATURE_GET))\n1285:\t\t\treturn -EINVAL;\n1286:\t\n1287:\t\tswitch (feature.flags \u0026 VFIO_DEVICE_FEATURE_MASK) {\n1288:\t\tcase VFIO_DEVICE_FEATURE_MIGRATION:\n1289:\t\t\treturn vfio_ioctl_device_feature_migration(\n1290:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1291:\t\t\t\tfeature.argsz - minsz);\n1292:\t\tcase VFIO_DEVICE_FEATURE_MIG_DEVICE_STATE:\n1293:\t\t\treturn vfio_ioctl_device_feature_mig_device_state(\n1294:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1295:\t\t\t\tfeature.argsz - minsz);\n1296:\t\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_START:\n1297:\t\t\treturn vfio_ioctl_device_feature_logging_start(\n1298:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1299:\t\t\t\tfeature.argsz - minsz);\n"
} |
| 59/3 | 2026/09/26 21:08 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37920 output: 40 thoughts: 17 Thoughts: Okay, here's my attempt at summarizing the thought process, as requested, geared toward an expert audience, and written in the first person: **Diving into `vfio_ioctl_device_feature`** Alright, let's get into the details of `vfio_ioctl_device_feature`. I need to understand this ioctl. Given my experience, I'm assuming this function is core to how features are handled at the device level within VFIO. I bet it handles the low-level negotiation and enabling/disabling of specific device capabilities presented to the guest. I need to figure out: * **What capabilities is it operating on?** Is it dealing with MSI-X vectors, DMA remapping, interrupt delivery modes, or something else entirely? The naming suggests a broad scope, so it's probably not limited to a single specific capability. * **What are the exact data structures and parameters involved?** I need to see the `ioctl` arguments. Are we talking about a simple bitmask, a complex structure containing feature IDs and values, or something more intricate? The definition of that struct is *crucial* to understanding its purpose and potential pitfalls. * **What's the control flow?** Does it use a switch statement to handle different feature IDs? What are the checks and security validations? How does it interact with the device's internal state? I'll be looking for potential race conditions or other vulnerabilities. The context switching in VFIO is often complex, and I'll need to carefully track how these features impact the state. * **Error handling is a must** I'll be analyzing the error handling. How does it report failures? Does it return detailed error codes, or just generic error flags? Are the error codes meaningful enough to debug? * **How does this fit in with the rest of the VFIO ecosystem?** I'll need to see how it integrates with other VFIO ioctls, especially those related to resource management and device assignment. This is likely how a guest interacts with this function, but the host driver could also leverage this functionality. * **Performance:** I'll be looking for areas that might cause performance bottlenecks, especially if locks are used within the critical path. * **Security:** How does this function interact with user-provided data and where does it get that from? This is most important! I'll be keeping an eye out for potential security vulnerabilities like buffer overflows, format string bugs, integer overflows, or off-by-one errors that malicious guests could exploit. Time to dig in. Let's see what we've got. |
| 60/3 | 2026/09/26 21:09 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/vfio_main.c",
"FirstLine": 1300,
"LineCount": 35
}Results: {
"Contents": "1300:\t\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_STOP:\n1301:\t\t\treturn vfio_ioctl_device_feature_logging_stop(\n1302:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1303:\t\t\t\tfeature.argsz - minsz);\n1304:\t\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_REPORT:\n1305:\t\t\treturn vfio_ioctl_device_feature_logging_report(\n1306:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1307:\t\t\t\tfeature.argsz - minsz);\n1308:\t\tcase VFIO_DEVICE_FEATURE_MIG_DATA_SIZE:\n1309:\t\t\treturn vfio_ioctl_device_feature_migration_data_size(\n1310:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1311:\t\t\t\tfeature.argsz - minsz);\n1312:\t\tcase VFIO_DEVICE_FEATURE_MIG_PRECOPY_INFOv2:\n1313:\t\t\treturn vfio_ioctl_device_feature_migration_precopy_info_v2(\n1314:\t\t\t\tdevice, feature.flags, feature.argsz - minsz);\n1315:\t\tdefault:\n1316:\t\t\tif (unlikely(!device-\u003eops-\u003edevice_feature))\n1317:\t\t\t\treturn -ENOTTY;\n1318:\t\t\treturn device-\u003eops-\u003edevice_feature(device, feature.flags,\n1319:\t\t\t\t\t\t\t arg-\u003edata,\n1320:\t\t\t\t\t\t\t feature.argsz - minsz);\n1321:\t\t}\n1322:\t}\n1323:\t\n1324:\tstatic long vfio_get_region_info(struct vfio_device *device,\n1325:\t\t\t\t\t struct vfio_region_info __user *arg)\n1326:\t{\n1327:\t\tunsigned long minsz = offsetofend(struct vfio_region_info, offset);\n1328:\t\tstruct vfio_region_info info = {};\n1329:\t\tstruct vfio_info_cap caps = {};\n1330:\t\tint ret;\n1331:\t\n1332:\t\tif (unlikely(!device-\u003eops-\u003eget_region_info_caps))\n1333:\t\t\treturn -EINVAL;\n1334:\t\n"
} |
| 61/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 39133 output: 23 thoughts: 1043 Thoughts: **Analyzing Potential Vulnerabilities in DMA-BUF and VFIO-PCI Interactions** Alright, let's break this down systematically. I see that `arg->data` is passed into `vfio_pci_core_ioctl_feature`, which calls `vfio_pci_core_feature_dma_buf_revoke`. My immediate concern here is with the data flow in `vfio_pci_core_feature_dma_buf_revoke`. Specifically, the code checks `VFIO_DEVICE_FEATURE_SET` using `vfio_check_feature` and then copies data from user space via `copy_from_user` into `db_revoke`. When `VFIO_DEVICE_FEATURE_PROBE` is used, nothing is copied. I need to ascertain if `db_revoke` is used after the copy, and if so, how. I need to confirm that there isn't some scenario where uninitialized memory from `db_revoke` could be read or leaked to userspace, or that other avenues of data exposure exist. Now, let's dive into the details. I see changes in `drivers/dma-buf/dma-buf.c`. I'm especially interested in `dma_buf_ioctl` and `dma_buf_set_name`. Before, we had a simple `dma_buf_set_name` with `strndup_user`. Now, with the changes, `dma_buf_set_name` is called within the ioctl handler. The crucial point here is that the `name` passed to `dma_buf_set_name` comes from `strndup_user`. This means `name` *should* be initialized, since `strndup_user` copies from userspace, and it's null-terminated. Still, I should confirm where `dma_buf_set_name` is being called from. Okay, let's track down those calls. I see it's called from both `dma_buf_ioctl` (where the memory is initialized by `strndup_user`), and in `vfio_pci_core_mmap_prep_dmabuf`. In the latter, the name comes from `kasprintf`, which *also* guarantees initialization and null-termination. Good. I'd want to be certain there are *no other* callers of `dma_buf_set_name` in the kernel, just to be absolutely sure no unexpected interactions are present. I have to be completely certain that there aren't any potential leaks or reads of uninitialized memory. |
| 62/3 | 2026/09/26 21:09 | tool | grepper |
0mArgs:{
"Expression": "dma_buf_set_name\\("
}Results: {
"Output": "drivers/dma-buf/dma-buf.c=512=static long dma_buf_ioctl(struct file *file,\n--\ndrivers/dma-buf/dma-buf.c-558-\ndrivers/dma-buf/dma-buf.c:559:\t\tret = dma_buf_set_name(dmabuf, name);\ndrivers/dma-buf/dma-buf.c-560-\t\tif (ret)\n--\ndrivers/dma-buf/dma-buf.c=858=EXPORT_SYMBOL_NS_GPL(dma_buf_put, \"DMA_BUF\");\n--\ndrivers/dma-buf/dma-buf.c-874- */\ndrivers/dma-buf/dma-buf.c:875:int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\ndrivers/dma-buf/dma-buf.c-876-{\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=377=int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-445-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:446:\tif (dma_buf_set_name(priv-\u003edmabuf, bufname)) {\ndrivers/vfio/pci/vfio_pci_dmabuf.c-447-\t\tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n--\ninclude/linux/dma-buf.h=572=void dma_buf_put(struct dma_buf *dmabuf);\ninclude/linux/dma-buf.h-573-\ninclude/linux/dma-buf.h:574:int dma_buf_set_name(struct dma_buf *dmabuf, char *name);\ninclude/linux/dma-buf.h-575-\n"
} |
| 63/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41298 output: 26 thoughts: 153 Thoughts: **Analyzing DMA Buffer Naming**
Okay, so I'm focusing in on where `dma_buf_set_name` is being called. It looks like it's only in two places, which narrows down the scope considerably. I need to understand why this function is being used, so let's start with `vfio_pci_dmabuf_export`. It looks like this is where the `dma_buf` is being *exported* within the VFIO-PCI context.
```c
static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,
struct vfio_pci_dma_buf *priv, u32 flags)
{
DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
...
```
Right, here it is, the `vfio_pci_dmabuf_export` function, and it seems like the key is `DEFINE_DMA_BUF_EXPORT_INFO`. Let's take a look at how this macro initializes the `exp_info` structure. It's likely setting up the necessary data for the export process, including the name if I recall correctly. I need to verify what the initialization of `exp_info` entails; specifically, I need to understand what gets zeroed out, since that gives a baseline for what fields *aren't* being populated directly. That should tell me a lot. So, let's trace this to `DEFINE_DMA_BUF_EXPORT_INFO`.
|
| 64/3 | 2026/09/26 21:09 | tool | grepper |
0mArgs:{
"Expression": "#define DEFINE_DMA_BUF_EXPORT_INFO"
}Results: {
"Output": "include/linux/dma-buf.h=508=struct dma_buf_export_info {\n--\ninclude/linux/dma-buf.h-524- */\ninclude/linux/dma-buf.h:525:#define DEFINE_DMA_BUF_EXPORT_INFO(name)\t\\\ninclude/linux/dma-buf.h-526-\tstruct dma_buf_export_info name = { .exp_name = KBUILD_MODNAME, \\\n"
} |
| 65/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41942 output: 38 thoughts: 897 Thoughts: **Analyzing the VFIO PCI DMA Buffer PFN Handling** Okay, let's break this down. `exp_info` is initialized to zero using designated initializers, nothing unexpected there. Now, the interesting part starts with the `vfio_pci_dma_buf_find_pfn` function. I see a structure `vfio_pci_core_device` (vdev) which seems to represent the core device, a `vfio_pci_dma_buf` (priv) representing the DMA buffer, a virtual memory area structure (vma), fault address, order, and an output pointer `out_pfn`. We are looking at how a Page Frame Number (PFN) is located. In the calling function `vfio_pci_mmap_huge_fault`, the `pfn` variable is initialized to 0. The function then acquires read locks on `vdev->memory_lock` and `vdev->dmabuf_lock` for concurrency control. The code checks `priv->status` and if it is `VFIO_PCI_DMABUF_OK`, it calls `vfio_pci_dma_buf_find_pfn` to find a PFN, and the result is stored in `pres`. If `pres` is zero, meaning a PFN was found, it calls `vfio_pci_vmf_insert_pfn` to insert the PFN into the virtual memory structure. It's the `vfio_pci_vmf_insert_pfn` that takes the `pfn` and the `order` as arguments. Note that it specifically passes the result of `vfio_pci_dma_buf_find_pfn`, which is `pfn`. If the order is not valid, it might return a `-ERANGE` and fall back to try another order, in that case `ret` is set to `VM_FAULT_FALLBACK`. Looking at `vfio_pci_dma_buf_find_pfn` itself, the key here is the loop iterating through DMA buffer ranges. It calculates `range_start_offset`, `fault_offset`, and `fault_offset_end`. The core logic checks if the faulting address falls within a valid DMA range. If so, it calculates a potential `pfn` using a bitshift operation, checking for alignment with the order. If the alignment check using `IS_ALIGNED` passes, then it assigns `pfn` to the `*out_pfn` pointer, and sets `ret = 0`. The loop breaks at this point, returning a positive result. If no range is found or the alignment is invalid, it proceeds to the next range. So, in the case of a successful PFN lookup, the `pfn` that's passed to `vfio_pci_vmf_insert_pfn` is the `pfn` calculated and validated within `vfio_pci_dma_buf_find_pfn`. The original initialization to zero at the beginning of `vfio_pci_mmap_huge_fault` is overwritten with the correct PFN if a valid range is found and alignment is met. Now I'm interested in the `vfio_pci_vmf_insert_pfn` to understand how the PFN is ultimately used in the mapping process. |
| 66/3 | 2026/09/26 21:09 | tool | grepper |
0mArgs:{
"Expression": "vfio_pci_vmf_insert_pfn",
"PathPrefix": "drivers/vfio/"
}Results: {
"Output": "drivers/vfio/pci/nvgrace-gpu/main.c=392=static vm_fault_t nvgrace_gpu_vfio_pci_huge_fault(struct vm_fault *vmf,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-440-\ndrivers/vfio/pci/nvgrace-gpu/main.c:441:\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);\ndrivers/vfio/pci/nvgrace-gpu/main.c-442-\t\t}\n--\ndrivers/vfio/pci/vfio_pci_core.c=1772=void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 cmd)\n--\ndrivers/vfio/pci/vfio_pci_core.c-1777-\ndrivers/vfio/pci/vfio_pci_core.c:1778:vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\ndrivers/vfio/pci/vfio_pci_core.c-1779-\t\t\t\t struct vm_fault *vmf,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1798-}\ndrivers/vfio/pci/vfio_pci_core.c:1799:EXPORT_SYMBOL_GPL(vfio_pci_vmf_insert_pfn);\ndrivers/vfio/pci/vfio_pci_core.c-1800-\ndrivers/vfio/pci/vfio_pci_core.c=1801=static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1880-\ndrivers/vfio/pci/vfio_pci_core.c:1881:\t/* memory_lock for vfio_pci_vmf_insert_pfn() */\ndrivers/vfio/pci/vfio_pci_core.c-1882-\tdown_read(\u0026vdev-\u003ememory_lock);\n--\ndrivers/vfio/pci/vfio_pci_core.c-1890-\t\tif (pres == 0)\ndrivers/vfio/pci/vfio_pci_core.c:1891:\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf,\ndrivers/vfio/pci/vfio_pci_core.c-1892-\t\t\t\t\t\t pfn, order);\n"
} |
| 67/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44310 output: 46 thoughts: 39 Thoughts: Okay, let's dive into this. Here's my take on `vfio_pci_vmf_insert_pfn` in `drivers/vfio/pci/vfio_pci_core.c`: **Analyzing `vfio_pci_vmf_insert_pfn`** Right, so I need to understand how this function, `vfio_pci_vmf_insert_pfn`, fits into the larger picture of VFIO PCI device assignment. Specifically, how it handles inserting a physical frame number (PFN) into the Virtual Machine's (VM) address space. This suggests we're dealing with the core mechanics of memory mapping for direct device access. It's the building block, the crucial step in enabling the VM to talk directly to the PCI device. I'm assuming there's some sort of address translation and protection going on. I'll need to see how the system ensures that the VM only accesses memory it's authorized to. I'll be looking for any potential race conditions or security vulnerabilities around the PFN insertion process, and how memory is validated at this point. I'm keen to see if this function handles page faults correctly. I'll be interested to see how it manages the page tables and any required flushing, and how it considers caching. My main goal is to understand the performance implications and potential pitfalls of this memory mapping function, so I'll need to keep an eye out for any locks, synchronization mechanisms, and potential bottlenecks. |
| 68/3 | 2026/09/26 21:09 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_core.c",
"FirstLine": 1778,
"LineCount": 25
}Results: {
"Contents": "1778:\tvm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\n1779:\t\t\t\t\t struct vm_fault *vmf,\n1780:\t\t\t\t\t unsigned long pfn,\n1781:\t\t\t\t\t unsigned int order)\n1782:\t{\n1783:\t\tlockdep_assert_held_read(\u0026vdev-\u003ememory_lock);\n1784:\t\n1785:\t\tif (vdev-\u003epm_runtime_engaged || !__vfio_pci_memory_enabled(vdev))\n1786:\t\t\treturn VM_FAULT_SIGBUS;\n1787:\t\n1788:\t\tif (!order)\n1789:\t\t\treturn vmf_insert_pfn(vmf-\u003evma, vmf-\u003eaddress, pfn);\n1790:\t\n1791:\t\tif (IS_ENABLED(CONFIG_ARCH_SUPPORTS_PMD_PFNMAP) \u0026\u0026 order == PMD_ORDER)\n1792:\t\t\treturn vmf_insert_pfn_pmd(vmf, pfn, false);\n1793:\t\n1794:\t\tif (IS_ENABLED(CONFIG_ARCH_SUPPORTS_PUD_PFNMAP) \u0026\u0026 order == PUD_ORDER)\n1795:\t\t\treturn vmf_insert_pfn_pud(vmf, pfn, false);\n1796:\t\n1797:\t\treturn VM_FAULT_FALLBACK;\n1798:\t}\n1799:\tEXPORT_SYMBOL_GPL(vfio_pci_vmf_insert_pfn);\n1800:\t\n1801:\tstatic vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n1802:\t\t\t\t\t\t unsigned int order)\n"
} |
| 69/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45193 output: 31 thoughts: 264 Thoughts: **Analyzing VFIO PCI Runtime Resume and Potential NVGrace GPU Interactions** Okay, let's dive in. I'm looking at `vfio_pci_core_runtime_resume`. Right, so we have the write lock on `vdev->memory_lock`, which is expected. Then, the crucial check: `vdev->pm_wake_eventfd_ctx`. This is where the runtime PM wake eventfd context comes in. If it exists, we grab the context pointer, `ctx`, and null out the `pm_wake_eventfd_ctx`. So far, so good – we're clearing the wake event, which triggers the actual wake event. It then exits runtime PM, checks memory enabled status and unrevokes BARs if appropriate. Then, it signals the `eventfd`, and then puts the `eventfd_ctx` at the end. Classic cleanup sequence, the removal of wake event and unrevoking of BARs is the key here. Now, my primary concern is to see if any of this runtime resume is relevant to the NVGrace GPU. I need to see if the new logic affects the device. Let's check `nvgrace_gpu_vfio_pci_huge_fault` to check if NVGrace GPU uses `vfio_pci_core_runtime_resume`. Specifically, I'm trying to determine if anything *changed* that might affect the NVGrace GPU. Hold on, I need to know: Does the NVGrace GPU module override mmap behavior? This is a crucial area for VFIO interactions. I need to be sure it's handling mapping and unmapping correctly. So, I need to head over to `drivers/vfio/pci/nvgrace-gpu/main.c` to see if there's any custom mmap logic in there. I'm going to scan the source for mmap calls or overridden functions immediately. That's the next logical step. |
| 70/3 | 2026/09/26 21:09 | tool | grepper |
0mArgs:{
"Expression": "\\.mmap",
"PathPrefix": "drivers/vfio/pci/"
}Results: {
"Output": "drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1593=static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1603-\t.write = hisi_acc_vfio_pci_write,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1604:\t.mmap = hisi_acc_vfio_pci_mmap,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1605-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1614=static const struct vfio_device_ops hisi_acc_vfio_pci_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1624-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1625:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1626-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/mlx5/main.c=1386=static const struct vfio_device_ops mlx5vf_pci_ops = {\n--\ndrivers/vfio/pci/mlx5/main.c-1396-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/mlx5/main.c:1397:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/mlx5/main.c-1398-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1065=static const struct vfio_device_ops nvgrace_gpu_pci_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1075-\t.write\t\t= nvgrace_gpu_write,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1076:\t.mmap\t\t= nvgrace_gpu_mmap,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1077-\t.request\t= vfio_pci_core_request,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1090=static const struct vfio_device_ops nvgrace_gpu_pci_core_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1100-\t.write\t\t= vfio_pci_core_write,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1101:\t.mmap\t\t= vfio_pci_core_mmap,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1102-\t.request\t= vfio_pci_core_request,\n--\ndrivers/vfio/pci/pds/vfio_dev.c=191=static const struct vfio_device_ops pds_vfio_ops = {\n--\ndrivers/vfio/pci/pds/vfio_dev.c-201-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/pds/vfio_dev.c:202:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/pds/vfio_dev.c-203-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/qat/main.c=604=static const struct vfio_device_ops qat_vf_pci_ops = {\n--\ndrivers/vfio/pci/qat/main.c-613-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/qat/main.c:614:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/qat/main.c-615-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/vfio_pci.c=149=static const struct vfio_device_ops vfio_pci_ops = {\n--\ndrivers/vfio/pci/vfio_pci.c-159-\t.write\t\t= vfio_pci_core_write,\ndrivers/vfio/pci/vfio_pci.c:160:\t.mmap\t\t= vfio_pci_core_mmap,\ndrivers/vfio/pci/vfio_pci.c-161-\t.request\t= vfio_pci_core_request,\n--\ndrivers/vfio/pci/vfio_pci_core.c=2296=int vfio_pci_core_init_dev(struct vfio_device *core_vdev)\n--\ndrivers/vfio/pci/vfio_pci_core.c-2318-\t/*\ndrivers/vfio/pci/vfio_pci_core.c:2319:\t * If a driver overrides .mmap, it has to be assumed that it\ndrivers/vfio/pci/vfio_pci_core.c-2320-\t * might not use the DMABUF-backed core mmap; this flag\ndrivers/vfio/pci/vfio_pci_core.c-2321-\t * enables a zap at revoke time. A driver can opt out by\ndrivers/vfio/pci/vfio_pci_core.c:2322:\t * clearing this flag at init, if their .mmap override calls\ndrivers/vfio/pci/vfio_pci_core.c-2323-\t * down to vfio_pci_core_mmap().\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=149=static const struct dma_buf_ops vfio_pci_dmabuf_ops = {\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-151-#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c:152:\t.mmap = vfio_pci_dma_buf_mmap,\ndrivers/vfio/pci/vfio_pci_dmabuf.c-153-#endif\n--\ndrivers/vfio/pci/virtio/main.c=84=static const struct vfio_device_ops virtiovf_vfio_pci_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-94-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/virtio/main.c:95:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/virtio/main.c-96-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/virtio/main.c=106=static const struct vfio_device_ops virtiovf_vfio_pci_tran_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-116-\t.write = virtiovf_pci_core_write,\ndrivers/vfio/pci/virtio/main.c:117:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/virtio/main.c-118-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/virtio/main.c=128=static const struct vfio_device_ops virtiovf_vfio_pci_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-138-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/virtio/main.c:139:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/virtio/main.c-140-\t.request = vfio_pci_core_request,\n--\ndrivers/vfio/pci/xe/main.c=522=static const struct vfio_device_ops xe_vfio_pci_ops = {\n--\ndrivers/vfio/pci/xe/main.c-532-\t.write = vfio_pci_core_write,\ndrivers/vfio/pci/xe/main.c:533:\t.mmap = vfio_pci_core_mmap,\ndrivers/vfio/pci/xe/main.c-534-\t.request = vfio_pci_core_request,\n"
} |
| 71/3 | 2026/09/26 21:09 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 48273 output: 423 thoughts: 1656 Thoughts: **Analysis of VFIO DMA-BUF Integration and KMSAN Applicability** Okay, let's break this down. First, I'm seeing a conditional related to memory mapping and BAR revocation, which depends on whether the `.mmap` operation is overridden. The key is that for certain drivers (like `hisi_acc_vfio_pci_migrn_ops`), this is *intentionally* cleared because the underlying mapping function still calls `vfio_pci_core_mmap`, but with others like `nvgrace-gpu`, it isn't. This impacts when `vfio_pci_revoke_bars` is called. Got it. Now, the main question: Does this patch introduce or expose uninitialized memory risks, and therefore, do we *really* need KMSAN? I need to methodically check against those KMSAN criteria. First, kernel structures sent to userspace. The new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl *only* reads data from user space. Similarly, `DMA_BUF_SET_NAME` takes user-provided strings. The sysfs/procfs output is either null or properly initialized strings from `strndup_user` or `kasprintf`. No leaks. Second, conditional logic depending on uninitialized variables? Nope. `kzalloc_obj` is used for allocations, zeroing the structures. `pfn` is initialized, `bars_revoked` is initialized, and `vdev->dmabuf_lock` is initialized during device initialization. Everything appears safe there. Third, complex data structures with potential uninitialized fields? Again, `vfio_pci_dma_buf` and `phys_vec` are allocated with `kzalloc_obj`. All used fields are initialized explicitly. Finally, bounds checks or logic that could lead to uninitialized memory access? `vfio_pci_dma_buf_find_pfn` does checks on offsets, sizes, alignment, bounds against the VMA, etc., to handle the MMIO/P2P DMA BAR mapping. Seems rigorous, and importantly, it operates on *mapped* regions, not uninitialized memory. So, what bugs *could* be introduced here? Lock issues, deadlocks (dmabuf_lock vs memory_lock vs resv lock), Use-After-Free or races (dmabuf release, vdev lifetime), out-of-bounds or arithmetic errors in `vfio_pci_dma_buf_find_pfn`, and NULL pointer dereferences. But *none* of these are uninitialized memory bugs or info leaks! KASAN and LOCKDEP are fully capable of catching the bugs here. The KMSAN rule is clear: *If* KASAN, LOCKDEP, or standard detectors can catch the bug, KMSAN is a waste. And in this case, all of the risks are concurrency related (use-after-free, double free, race conditions, locking deadlocks) or memory access violations that will be caught by existing tools. KMSAN is redundant. Essentially, this patch refactors DMA-BUF integration within VFIO, which involves setting names, mapping BARs through DMA-BUFs, page fault handlers, and locking. All the new allocations are zeroed. The only data that goes from kernel to user is already controlled and initialized. The risks are primarily race conditions and memory access errors that are already handled by KASAN and LOCKDEP. No need for KMSAN. The reasoning is solid. Time to call the `set-results` tool and move on. |
| 72/3 | 2026/09/26 21:09 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch series refactors DMA-BUF naming (allowing kernel exporters to set debug names) and integrates DMA-BUF backing into the VFIO PCI BAR mmap infrastructure, introducing a new VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE feature and granular locking with a new rwsem (dmabuf_lock).\n\nKMSAN vs KASAN Applicability:\n1. Uninitialized Memory / Info-leak Analysis:\n - All heap allocations for newly introduced structures (struct vfio_pci_dma_buf, phys_vec) use kzalloc_obj, ensuring all fields are zero-initialized.\n - The string handling in dma_buf_set_name operates on NULL-terminated, initialized strings produced by kasprintf() and strndup_user().\n - The new ioctl feature (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE) is an input-only command (VFIO_DEVICE_FEATURE_SET) that uses copy_from_user; no kernel data or padding is returned or copied to user space.\n - Local variables in the fault handler (vfio_pci_mmap_huge_fault, vfio_pci_dma_buf_find_pfn) such as pfn are explicitly initialized prior to being used or checked.\n\n2. Bug Detector Suitability:\n - The potential bugs introduced by these changes are concurrency/locking deadlocks (dmabuf_lock vs memory_lock vs DMA-BUF resv lock), object lifetime/racing teardown issues (DMABUF vs VFIO device vs VMA references), and potential use-after-free or invalid PFN insertion.\n - These issues are comprehensively covered by LOCKDEP and KASAN.\n - There are no uninitialized memory reads or info-leaks, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch series refactors DMA-BUF naming (allowing kernel exporters to set debug names) and integrates DMA-BUF backing into the VFIO PCI BAR mmap infrastructure, introducing a new VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE feature and granular locking with a new rwsem (dmabuf_lock).\n\nKMSAN vs KASAN Applicability:\n1. Uninitialized Memory / Info-leak Analysis:\n - All heap allocations for newly introduced structures (struct vfio_pci_dma_buf, phys_vec) use kzalloc_obj, ensuring all fields are zero-initialized.\n - The string handling in dma_buf_set_name operates on NULL-terminated, initialized strings produced by kasprintf() and strndup_user().\n - The new ioctl feature (VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE) is an input-only command (VFIO_DEVICE_FEATURE_SET) that uses copy_from_user; no kernel data or padding is returned or copied to user space.\n - Local variables in the fault handler (vfio_pci_mmap_huge_fault, vfio_pci_dma_buf_find_pfn) such as pfn are explicitly initialized prior to being used or checked.\n\n2. Bug Detector Suitability:\n - The potential bugs introduced by these changes are concurrency/locking deadlocks (dmabuf_lock vs memory_lock vs DMA-BUF resv lock), object lifetime/racing teardown issues (DMABUF vs VFIO device vs VMA references), and potential use-after-free or invalid PFN insertion.\n - These issues are comprehensively covered by LOCKDEP and KASAN.\n - There are no uninitialized memory reads or info-leaks, so a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|