| Seq | Timestamp | Type | Name | Duration |
|---|---|---|---|---|
| 0/0 | 2026/10/06 20:14 | flow | patch-triage |
10mArgs:null Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER",
"ZONE_DEVICE",
"MEMORY_HOTPLUG",
"MEMORY_HOTREMOVE"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_mmap",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_dma_buf_cleanup",
"vfio_pci_dma_buf_mmap"
],
"KMSANReasoning": "The patch refactors dma-buf naming by exporting `dma_buf_set_name()`, transitions VFIO PCI BAR mmaps to be backed by dma-buf allocations, and adds a `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl feature along with associated locking (`dmabuf_lock`) and fault-handling logic (`vfio_pci_dma_buf_find_pfn`).\n\nAnalysis of uninitialized memory vs other sanitizer applicability:\n1. No info leaks to userspace: The new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl only consumes input (`copy_from_user` for SET) and copies nothing to userspace. The `dma_buf_set_name` ioctl path reads from userspace via `strndup_user` and outputs no data to userspace.\n2. Memory allocations: All dynamically allocated structures (`priv`, `phys_vec`) are allocated using `kzalloc_obj`, zeroing all fields. String formatting via `kasprintf` ensures proper null-termination.\n3. Fault handling and address calculations: `vfio_pci_dma_buf_find_pfn()` uses overflow checks (`check_add_overflow`) and strictly writes `*out_pfn` on success, while the caller initializes `pfn = 0`. Mapped PFNs correspond to device MMIO BAR physical addresses, not kernel RAM.\n4. Bug profile: The changes primarily involve complex concurrency management, locking hierarchy (`memory_lock`, `dmabuf_lock`, and reservation locks), object lifetimes (`dma_buf`, `vfio_pci_core_device`, `vfile`), and refcounting (`kref`, `vfio_device_registration`). These issues (use-after-free, deadlocks, out-of-bounds accesses) are detected by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory risks or kernel information leaks that would exclusively be detected by KMSAN, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false,
"Reasoning": "The patch series modifies core DMA-BUF code (exporting and refactoring dma_buf_set_name and its ioctl handling) and heavily reworks VFIO PCI memory management and revocation. In particular, VFIO PCI BAR mmaps are now backed by DMA-BUFs, page fault handling is rewritten to navigate complex lock ordering across memory_lock, dmabuf_lock, and reservation locks, and a new feature ioctl VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE is introduced. These code paths are fully reachable in virtualized QEMU environments with emulated PCI devices and warrant targeted fuzzing.",
"WorthFuzzing": true
} |
| 1/1 | 2026/10/06 20:14 | action | read-patch-diff |
0mArgs:null Results: {
"PatchDiff": "commit 8be538d0a9753d071d10cdbcaec6469f392184a3\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate: Tue Oct 6 20:14:46 2026 +0000\n\n syz-cluster: applied patch under review\n\ndiff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c\nindex 4c9add51f9ef5..b99da4ff2ef62 100644\n--- a/drivers/dma-buf/dma-buf.c\n+++ b/drivers/dma-buf/dma-buf.c\n@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)\n \treturn events;\n }\n \n-/**\n- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.\n- * It could support changing the name of the dma-buf if the same\n- * piece of memory is used for multiple purpose between different devices.\n- *\n- * @dmabuf: [in] dmabuf buffer that will be renamed.\n- * @buf: [in] A piece of userspace memory that contains the name of\n- * the dma-buf.\n- *\n- * Returns 0 on success. If the dma-buf buffer is already attached to\n- * devices, return -EBUSY.\n- *\n- */\n-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)\n-{\n-\tchar *name = strndup_user(buf, DMA_BUF_NAME_LEN);\n-\n-\tif (IS_ERR(name))\n-\t\treturn PTR_ERR(name);\n-\n-\tspin_lock(\u0026dmabuf-\u003ename_lock);\n-\tkfree(dmabuf-\u003ename);\n-\tdmabuf-\u003ename = name;\n-\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n-\n-\treturn 0;\n-}\n-\n #if IS_ENABLED(CONFIG_SYNC_FILE)\n static long dma_buf_export_sync_file(struct dma_buf *dmabuf,\n \t\t\t\t void __user *user_data)\n@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,\n \t\treturn ret;\n \n \tcase DMA_BUF_SET_NAME_A:\n-\tcase DMA_BUF_SET_NAME_B:\n-\t\treturn dma_buf_set_name(dmabuf, (const char __user *)arg);\n+\tcase DMA_BUF_SET_NAME_B: {\n+\t\tchar *name = strndup_user((const char __user *)arg,\n+\t\t\t\t\t DMA_BUF_NAME_LEN);\n+\n+\t\tif (IS_ERR(name))\n+\t\t\treturn PTR_ERR(name);\n+\n+\t\tret = dma_buf_set_name(dmabuf, name);\n+\t\tif (ret)\n+\t\t\tkfree(name);\n+\n+\t\treturn ret;\n+\t}\n \n #if IS_ENABLED(CONFIG_SYNC_FILE)\n \tcase DMA_BUF_IOCTL_EXPORT_SYNC_FILE:\n@@ -874,6 +857,42 @@ void dma_buf_put(struct dma_buf *dmabuf)\n }\n EXPORT_SYMBOL_NS_GPL(dma_buf_put, \"DMA_BUF\");\n \n+/**\n+ * dma_buf_set_name - Set a dmabuf's name\n+ * @dmabuf: [in] dmabuf buffer that will be renamed.\n+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or\n+ * similar. This takes ownership of the allocation\n+ * on success, which will be kfree()d when the\n+ * dmabuf is released or a new name assigned.\n+ *\n+ * Intended to be used by the exporter to set a name for debug\n+ * purposes. This can also change an existing name if the same piece\n+ * of memory is used for multiple purposes over time. Importers\n+ * should NOT use this function to overwrite the name.\n+ *\n+ * Return: 0 on success, -EINVAL if the name is NULL, or -E2BIG if the\n+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.\n+ */\n+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\n+{\n+\tif (!name)\n+\t\treturn -EINVAL;\n+\n+\t/* dmabuffs_dname() won't use the string if the length\n+\t * (including terminator) exceeds DMA_BUF_NAME_LEN:\n+\t */\n+\tif (strlen(name) \u003e= DMA_BUF_NAME_LEN)\n+\t\treturn -E2BIG;\n+\n+\tspin_lock(\u0026dmabuf-\u003ename_lock);\n+\tkfree(dmabuf-\u003ename);\n+\tdmabuf-\u003ename = name;\n+\tspin_unlock(\u0026dmabuf-\u003ename_lock);\n+\n+\treturn 0;\n+}\n+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, \"DMA_BUF\");\n+\n static int dma_buf_wrap_sg_table(struct sg_table **sg_table)\n {\n \tstruct scatterlist *to_sg, *from_sg;\n@@ -1004,6 +1023,11 @@ dma_buf_pin_on_map(struct dma_buf_attachment *attach)\n * functions:\n *\n * - dma_buf_invalidate_mappings()\n+ *\n+ * 4. Exporters must not hold the dma-buf reservation lock when calling these\n+ * functions:\n+ *\n+ * - dma_buf_set_name()\n */\n \n /**\ndiff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig\nindex 296bf01e185ec..c6d6fb09af86c 100644\n--- a/drivers/vfio/pci/Kconfig\n+++ b/drivers/vfio/pci/Kconfig\n@@ -6,6 +6,7 @@ config VFIO_PCI_CORE\n \ttristate\n \tselect VFIO_VIRQFD\n \tselect IRQ_BYPASS_MANAGER\n+\tselect DMA_SHARED_BUFFER\n \n config VFIO_PCI_INTX\n \tdef_bool y if !S390\n@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM\n \t To enable s390x KVM vfio-pci extensions, say Y.\n \n config VFIO_PCI_DMABUF\n-\tdef_bool y if VFIO_PCI_CORE \u0026\u0026 PCI_P2PDMA \u0026\u0026 DMA_SHARED_BUFFER\n+\tdef_bool y if PCI_P2PDMA\n+\tdepends on VFIO_PCI_CORE\n \n source \"drivers/vfio/pci/mlx5/Kconfig\"\n \ndiff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile\nindex 6138f1bf241df..881452ea89be0 100644\n--- a/drivers/vfio/pci/Makefile\n+++ b/drivers/vfio/pci/Makefile\n@@ -1,8 +1,7 @@\n # SPDX-License-Identifier: GPL-2.0-only\n \n-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o\n+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o\n vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o\n-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o\n obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o\n \n vfio-pci-y := vfio_pci.o\ndiff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\nindex 86362ec424a50..14622556355eb 100644\n--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\n+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c\n@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)\n \tstruct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);\n \tstruct pci_dev *pdev = to_pci_dev(core_vdev-\u003edev);\n \tstruct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);\n+\tint ret;\n \n \thisi_acc_vdev-\u003evf_id = pci_iov_vf_id(pdev) + 1;\n \thisi_acc_vdev-\u003epf_qm = pf_qm;\n@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)\n \tcore_vdev-\u003emigration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;\n \tcore_vdev-\u003emig_ops = \u0026hisi_acc_vfio_pci_migrn_state_ops;\n \n-\treturn vfio_pci_core_init_dev(core_vdev);\n+\tret = vfio_pci_core_init_dev(core_vdev);\n+\tif (ret)\n+\t\treturn ret;\n+\t/*\n+\t * hisi_acc_vfio_pci_mmap() calls down to\n+\t * vfio_pci_core_mmap(), so BAR mappings are still\n+\t * DMABUF-backed. They don't require a zap on revoke, so opt\n+\t * out:\n+\t */\n+\thisi_acc_vdev-\u003ecore_device.zap_bars_on_revoke = false;\n+\n+\treturn 0;\n }\n \n static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {\ndiff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c\nindex 9914f3ac69aef..cef337f4e8f2e 100644\n--- a/drivers/vfio/pci/vfio_pci_config.c\n+++ b/drivers/vfio/pci/vfio_pci_config.c\n@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\tvirt_mem = !!(le16_to_cpu(*virt_cmd) \u0026 PCI_COMMAND_MEMORY);\n \t\tnew_mem = !!(new_cmd \u0026 PCI_COMMAND_MEMORY);\n \n-\t\tif (!new_mem) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t\t} else {\n+\t\tif (!new_mem)\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n+\t\telse\n \t\t\tdown_write(\u0026vdev-\u003ememory_lock);\n-\t\t}\n \n \t\t/*\n \t\t * If the user is writing mem/io enable (new_mem/io) and we\n@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t*virt_cmd |= cpu_to_le16(new_cmd \u0026 mask);\n \n \t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n \t}\n \n@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)\n static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,\n \t\t\t\t\t pci_power_t state)\n {\n-\tif (state \u003e= PCI_D3hot) {\n-\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t} else {\n+\tif (state \u003e= PCI_D3hot)\n+\t\tvfio_pci_lock_revoke_bars(vdev);\n+\telse\n \t\tdown_write(\u0026vdev-\u003ememory_lock);\n-\t}\n \n \tvfio_pci_set_power_state(vdev, state);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t\t\t\t\t \u0026cap);\n \n \t\tif (!ret \u0026\u0026 (cap \u0026 PCI_EXP_DEVCAP_FLR)) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n \t\t\tpci_try_reset_function(vdev-\u003epdev);\n \t\t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\t\tup_write(\u0026vdev-\u003ememory_lock);\n \t\t}\n \t}\n@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,\n \t\t\t\t\t\t\u0026cap);\n \n \t\tif (!ret \u0026\u0026 (cap \u0026 PCI_AF_CAP_FLR) \u0026\u0026 (cap \u0026 PCI_AF_CAP_TP)) {\n-\t\t\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\t\t\tvfio_pci_dma_buf_move(vdev, true);\n+\t\t\tvfio_pci_lock_revoke_bars(vdev);\n \t\t\tpci_try_reset_function(vdev-\u003epdev);\n \t\t\tif (__vfio_pci_memory_enabled(vdev))\n-\t\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\t\tup_write(\u0026vdev-\u003ememory_lock);\n \t\t}\n \t}\ndiff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c\nindex 6757054e9d875..68e582ad38463 100644\n--- a/drivers/vfio/pci/vfio_pci_core.c\n+++ b/drivers/vfio/pci/vfio_pci_core.c\n@@ -13,6 +13,8 @@\n #include \u003clinux/aperture.h\u003e\n #include \u003clinux/debugfs.h\u003e\n #include \u003clinux/device.h\u003e\n+#include \u003clinux/dma-buf.h\u003e\n+#include \u003clinux/dma-resv.h\u003e\n #include \u003clinux/eventfd.h\u003e\n #include \u003clinux/file.h\u003e\n #include \u003clinux/interrupt.h\u003e\n@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,\n \t * The vdev power related flags are protected with 'memory_lock'\n \t * semaphore.\n \t */\n-\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n-\tvfio_pci_dma_buf_move(vdev, true);\n+\tvfio_pci_lock_revoke_bars(vdev);\n \n \tif (vdev-\u003epm_runtime_engaged) {\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)\n \tdown_write(\u0026vdev-\u003ememory_lock);\n \t__vfio_pci_runtime_pm_exit(vdev);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)\n \t */\n \tdown_write(\u0026vdev-\u003ememory_lock);\n \tif (vdev-\u003epm_wake_eventfd_ctx) {\n-\t\teventfd_signal(vdev-\u003epm_wake_eventfd_ctx);\n+\t\tstruct eventfd_ctx *ctx = vdev-\u003epm_wake_eventfd_ctx;\n+\n+\t\tvdev-\u003epm_wake_eventfd_ctx = NULL;\n \t\t__vfio_pci_runtime_pm_exit(vdev);\n+\t\tif (__vfio_pci_memory_enabled(vdev))\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n+\t\teventfd_signal(ctx);\n+\t\teventfd_ctx_put(ctx);\n \t}\n \tup_write(\u0026vdev-\u003ememory_lock);\n \n@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)\n \t\tvdev-\u003ehas_vga = true;\n \n \tvfio_pci_core_map_bars(vdev);\n+\tvdev-\u003ebars_revoked = false;\n \n \treturn 0;\n \n@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,\n \treturn ret;\n }\n \n+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);\n+\n static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \t\t\t\tvoid __user *arg)\n {\n@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \tif (!vdev-\u003ereset_works)\n \t\treturn -EINVAL;\n \n-\tvfio_pci_zap_and_down_write_memory_lock(vdev);\n+\tdown_write(\u0026vdev-\u003ememory_lock);\n \n \t/*\n \t * This function can be invoked while the power state is non-D0. If\n@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,\n \t * have NoSoftRst-, the reset function can cause the PCI config space\n \t * reset without restoring the original state (saved locally in\n \t * 'vdev-\u003epm_save').\n+\t *\n+\t * The zap is done after making the device accessible in D0,\n+\t * because a DMABUF importer could access the device as part\n+\t * of its revocation cleanup.\n \t */\n \tvfio_pci_set_power_state(vdev, PCI_D0);\n \n-\tvfio_pci_dma_buf_move(vdev, true);\n+\tvfio_pci_revoke_bars(vdev);\n+\n \tret = pci_try_reset_function(vdev-\u003epdev);\n \tif (__vfio_pci_memory_enabled(vdev))\n-\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\tvfio_pci_unrevoke_bars(vdev);\n \tup_write(\u0026vdev-\u003ememory_lock);\n \n \treturn ret;\n@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,\n \t\treturn vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);\n \tcase VFIO_DEVICE_FEATURE_ZPCI_ERROR:\n \t\treturn vfio_pci_zdev_feature_err(device, flags, arg, argsz);\n+\tcase VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:\n+\t\treturn vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);\n \tdefault:\n \t\treturn -ENOTTY;\n \t}\n@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu\n }\n EXPORT_SYMBOL_GPL(vfio_pci_core_write);\n \n-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)\n+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)\n {\n-\tstruct vfio_device *core_vdev = \u0026vdev-\u003evdev;\n-\tloff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);\n-\tloff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);\n-\tloff_t len = end - start;\n+\tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n+\tvfio_pci_dma_buf_move(vdev, true);\n \n-\tunmap_mapping_range(core_vdev-\u003einode-\u003ei_mapping, start, len, true);\n+\t/*\n+\t * If a driver could possibly create BAR mappings in the\n+\t * vdev's address_space, do an additional zap on revoke. See\n+\t * vfio_pci_core_init_dev().\n+\t */\n+\tif (vdev-\u003ezap_bars_on_revoke) {\n+\t\tstruct vfio_device *core_vdev = \u0026vdev-\u003evdev;\n+\t\tloff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);\n+\t\tloff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);\n+\t\tloff_t len = end - start;\n+\n+\t\tunmap_mapping_range(core_vdev-\u003einode-\u003ei_mapping,\n+\t\t\t\t start, len, true);\n+\t}\n }\n \n-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)\n+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)\n {\n \tdown_write(\u0026vdev-\u003ememory_lock);\n-\tvfio_pci_zap_bars(vdev);\n+\tvfio_pci_revoke_bars(vdev);\n+}\n+\n+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)\n+{\n+\tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n+\tvfio_pci_dma_buf_move(vdev, false);\n }\n \n u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)\n@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n \n-static unsigned long vma_to_pfn(struct vm_area_struct *vma)\n-{\n-\tstruct vfio_pci_core_device *vdev = vma-\u003evm_private_data;\n-\tint index = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n-\tu64 pgoff;\n-\n-\tpgoff = vma-\u003evm_pgoff \u0026\n-\t\t((1U \u003c\u003c (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);\n-\n-\treturn (pci_resource_start(vdev-\u003epdev, index) \u003e\u003e PAGE_SHIFT) + pgoff;\n-}\n-\n vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\n \t\t\t\t struct vm_fault *vmf,\n \t\t\t\t unsigned long pfn,\n@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n \t\t\t\t\t unsigned int order)\n {\n \tstruct vm_area_struct *vma = vmf-\u003evma;\n-\tstruct vfio_pci_core_device *vdev = vma-\u003evm_private_data;\n-\tunsigned long addr = vmf-\u003eaddress \u0026 ~((PAGE_SIZE \u003c\u003c order) - 1);\n-\tunsigned long pgoff = linear_page_delta(vma, addr);\n-\tunsigned long pfn = vma_to_pfn(vma) + pgoff;\n-\tvm_fault_t ret = VM_FAULT_FALLBACK;\n-\n-\tif (is_aligned_for_order(vma, addr, pfn, order)) {\n-\t\tscoped_guard(rwsem_read, \u0026vdev-\u003ememory_lock)\n-\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);\n+\tstruct vfio_pci_dma_buf *priv = vma-\u003evm_private_data;\n+\tstruct vfio_pci_core_device *vdev;\n+\tunsigned long pfn = 0;\n+\tvm_fault_t ret = VM_FAULT_SIGBUS;\n+\n+\t/*\n+\t * The only thing this can rely on is that the DMABUF relating\n+\t * to the VMA's vm_file exists (priv).\n+\t *\n+\t * A DMABUF for a VFIO device fd mmap() holds a reference to\n+\t * the original VFIO device fd, but an explicitly-exported\n+\t * DMABUF does not. The original fd might have closed,\n+\t * meaning this fault can race with\n+\t * vfio_pci_dma_buf_cleanup(), meaning the buffer could have\n+\t * been revoked (in which case priv-\u003evdev might be NULL), and\n+\t * the VFIO device registration might have been dropped.\n+\t *\n+\t * With the goal of taking vdev locks in a world where vdev\n+\t * might not still exist:\n+\t *\n+\t * 1. Take the resv lock on the DMABUF:\n+\t * - If racing cleanup got in first, the buffer is revoked;\n+\t * stop/exit if so.\n+\t * - If we got in first, the buffer is not revoked so vdev is\n+\t * non-NULL, accessible, and cleanup _has not yet put the\n+\t * VFIO device registration_. So, the device refcount must\n+\t * be \u003e0.\n+\t *\n+\t * 2. Take vfio_device registration (refcount guaranteed \u003e0\n+\t * hereafter).\n+\t *\n+\t * 3. Unlock the DMABUF's resv lock:\n+\t * - A racing cleanup can now complete.\n+\t * - But, the device refcount \u003e0, meaning the vfio_device\n+\t * (and vfio_pci_core_device vdev) have not yet been\n+\t * freed. vdev is accessible, even if the DMABUF has been\n+\t * revoked or cleanup has happened, because\n+\t * vfio_unregister_group_dev() can't complete.\n+\t *\n+\t * 4. Take the vdev-\u003ememory_lock then vdev-\u003edmabuf_lock:\n+\t * - Either the DMABUF is usable, or has been cleaned up.\n+\t * - It's not necessary to also take the resv lock, because\n+\t * the status/vdev can't change while dmabuf_lock is held.\n+\t * - Test the DMABUF revocation status again: if it was\n+\t * revoked between 1 and 4, return a SIGBUS. Otherwise,\n+\t * return a PFN.\n+\t *\n+\t * 5. Unlock, done.\n+\t */\n+\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK) {\n+\t\tpr_debug_ratelimited(\"%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\\n\",\n+\t\t\t\t __func__, vmf-\u003eaddress, vma-\u003evm_pgoff);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\treturn VM_FAULT_SIGBUS;\n+\t}\n+\n+\t/* If the buffer isn't revoked, vdev is valid */\n+\tvdev = priv-\u003evdev;\n+\n+\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev)) {\n+\t\t/*\n+\t\t * If vdev != NULL (above), the registration should\n+\t\t * already be \u003e0 and so this try_get should never\n+\t\t * fail.\n+\t\t */\n+\t\tdev_warn_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\t \"%s: Unexpected registration failure\\n\",\n+\t\t\t\t __func__);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\treturn VM_FAULT_SIGBUS;\n+\t}\n+\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\n+\t/* memory_lock for vfio_pci_vmf_insert_pfn() */\n+\tdown_read(\u0026vdev-\u003ememory_lock);\n+\t/* Re-test revocation status under dmabuf_lock */\n+\tdown_read(\u0026vdev-\u003edmabuf_lock);\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_OK) {\n+\t\tint pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,\n+\t\t\t\t\t\t vmf-\u003eaddress,\n+\t\t\t\t\t\t order, \u0026pfn);\n+\n+\t\tif (pres == 0)\n+\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf,\n+\t\t\t\t\t\t pfn, order);\n+\t\telse if (pres == -ERANGE)\n+\t\t\tret = VM_FAULT_FALLBACK;\n \t}\n+\tup_read(\u0026vdev-\u003edmabuf_lock);\n+\tup_read(\u0026vdev-\u003ememory_lock);\n \n \tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n-\t\t\t \"%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\\n\",\n-\t\t\t __func__, order,\n-\t\t\t vma-\u003evm_pgoff \u003e\u003e\n-\t\t\t\t(VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),\n-\t\t\t pgoff, (unsigned int)ret);\n+\t\t\t \"%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\\n\",\n+\t\t\t __func__, order, pfn, vmf-\u003eaddress,\n+\t\t\t vma-\u003evm_pgoff, (unsigned int)ret);\n \n+\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n \treturn ret;\n }\n \n@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {\n #endif\n };\n \n+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)\n+{\n+\tvma-\u003evm_ops = \u0026vfio_pci_mmap_ops;\n+}\n+\n int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)\n {\n \tstruct vfio_pci_core_device *vdev =\n@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma\n \tunsigned int index;\n \tu64 phys_len, req_len, pgoff, req_start;\n \tvoid __iomem *bar_io;\n+\tint ret;\n \n \tindex = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n \n@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma\n \tif (IS_ERR(bar_io))\n \t\treturn PTR_ERR(bar_io);\n \n-\tvma-\u003evm_private_data = vdev;\n+\tret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,\n+\t\t\t\t\t pci_resource_start(pdev, index),\n+\t\t\t\t\t req_len, index);\n+\tif (ret)\n+\t\treturn ret;\n+\n \tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n \tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n \n@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)\n \t\treturn ret;\n \tINIT_LIST_HEAD(\u0026vdev-\u003edmabufs);\n \tinit_rwsem(\u0026vdev-\u003ememory_lock);\n+\tinit_rwsem(\u0026vdev-\u003edmabuf_lock);\n \txa_init(\u0026vdev-\u003ectx);\n \n+\t/*\n+\t * If a driver overrides .mmap, it has to be assumed that it\n+\t * might not use the DMABUF-backed core mmap; this flag\n+\t * enables a zap at revoke time. A driver can opt out by\n+\t * clearing this flag at init, if their .mmap override calls\n+\t * down to vfio_pci_core_mmap().\n+\t */\n+\tif (vdev-\u003evdev.ops-\u003emmap != vfio_pci_core_mmap)\n+\t\tvdev-\u003ezap_bars_on_revoke = true;\n+\n \treturn 0;\n }\n EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);\n@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \t\t}\n \n \t\t/*\n-\t\t * Take the memory write lock for each device and zap BAR\n-\t\t * mappings to prevent the user accessing the device while in\n-\t\t * reset. Locking multiple devices is prone to deadlock,\n+\t\t * Take the memory write lock for each device and\n+\t\t * zap/revoke BAR mappings to prevent the user (or\n+\t\t * peers) accessing the device while in reset.\n+\t\t * Locking multiple devices is prone to deadlock,\n \t\t * runaway and unwind if we hit contention.\n \t\t */\n \t\tif (!down_write_trylock(\u0026vdev-\u003ememory_lock)) {\n@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \t\t\tbreak;\n \t\t}\n \n-\t\tvfio_pci_dma_buf_move(vdev, true);\n-\t\tvfio_pci_zap_bars(vdev);\n+\t\tvfio_pci_revoke_bars(vdev);\n \t}\n \n \tif (!list_entry_is_head(vdev,\n@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,\n \tlist_for_each_entry_from_reverse(vdev, \u0026dev_set-\u003edevice_list,\n \t\t\t\t\t vdev.dev_set_list) {\n \t\tif (vdev-\u003evdev.open_count \u0026\u0026 __vfio_pci_memory_enabled(vdev))\n-\t\t\tvfio_pci_dma_buf_move(vdev, false);\n+\t\t\tvfio_pci_unrevoke_bars(vdev);\n \t\tup_write(\u0026vdev-\u003ememory_lock);\n \t}\n \ndiff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c\nindex c16f460c01d68..b57bfaefd9fae 100644\n--- a/drivers/vfio/pci/vfio_pci_dmabuf.c\n+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c\n@@ -3,25 +3,14 @@\n */\n #include \u003clinux/dma-buf-mapping.h\u003e\n #include \u003clinux/pci-p2pdma.h\u003e\n+#include \u003clinux/dma-buf.h\u003e\n #include \u003clinux/dma-resv.h\u003e\n \n #include \"vfio_pci_priv.h\"\n \n MODULE_IMPORT_NS(\"DMA_BUF\");\n \n-struct vfio_pci_dma_buf {\n-\tstruct dma_buf *dmabuf;\n-\tstruct vfio_pci_core_device *vdev;\n-\tstruct list_head dmabufs_elm;\n-\tsize_t size;\n-\tstruct phys_vec *phys_vec;\n-\tstruct p2pdma_provider *provider;\n-\tu32 nr_ranges;\n-\tstruct kref kref;\n-\tstruct completion comp;\n-\tu8 revoked : 1;\n-};\n-\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \t\t\t\t struct dma_buf_attachment *attachment)\n {\n@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \tif (!attachment-\u003epeer2peer)\n \t\treturn -EOPNOTSUPP;\n \n-\tif (priv-\u003erevoked)\n+\tif (READ_ONCE(priv-\u003estatus) != VFIO_PCI_DMABUF_OK)\n \t\treturn -ENODEV;\n \n \tif (!dma_buf_attach_revocable(attachment))\n@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n \treturn 0;\n }\n \n+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)\n+{\n+\tstruct vfio_pci_dma_buf *priv = dmabuf-\u003epriv;\n+\n+\t/*\n+\t * dma_buf_mmap_internal() has asserted that the VMA is\n+\t * contained within the DMABUF size before calling this.\n+\t *\n+\t * Also, if we observe that the buffer is revoked now then\n+\t * refuse the mmap(). This is a belt-and-braces early failure\n+\t * to ease debugging a revoked buffer being used. Userspace\n+\t * might also race an mmap() against an explicit revocation,\n+\t * or an action causing a revoke; race scenarios are still\n+\t * safe because the fault handler ultimately prevents access\n+\t * to a revoked buffer if it isn't caught here.\n+\t */\n+\tif (READ_ONCE(priv-\u003estatus) != VFIO_PCI_DMABUF_OK)\n+\t\treturn -ENODEV;\n+\t/*\n+\t * Make clear that anything with an offset adjustment is\n+\t * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()\n+\t * maths would underflow; this doesn't happen through the\n+\t * regular DMABUF export path used with this mmap(). A DMABUF\n+\t * implicitly created for BAR mmap could have adjust \u003e 0, but\n+\t * these can't currently be re-opened and mmap()ed again.\n+\t * Catch here in case that assumption ever changes.\n+\t */\n+\tif (priv-\u003evma_pgoff_adjust)\n+\t\treturn -EINVAL;\n+\tif ((vma-\u003evm_flags \u0026 VM_SHARED) == 0)\n+\t\treturn -EINVAL;\n+\n+\tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n+\tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n+\n+\t/* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */\n+\tvm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |\n+\t\t VM_DONTEXPAND | VM_DONTDUMP);\n+\tvma-\u003evm_private_data = priv;\n+\tvfio_pci_set_vma_ops(vma);\n+\n+\treturn 0;\n+}\n+#else\n+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n+\t\t\t\t struct dma_buf_attachment *attachment)\n+{\n+\t/*\n+\t * Explicit export can't occur without the DMABUF feature, but\n+\t * DMABUFs are implicitly created for BAR mappings. An\n+\t * .attach that fails prevents dma_buf_attach().\n+\t */\n+\treturn -EOPNOTSUPP;\n+}\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\n+\n static void vfio_pci_dma_buf_done(struct kref *kref)\n {\n \tstruct vfio_pci_dma_buf *priv =\n@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,\n \n \tdma_resv_assert_held(priv-\u003edmabuf-\u003eresv);\n \n-\tif (priv-\u003erevoked)\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK)\n \t\treturn ERR_PTR(-ENODEV);\n \n \tret = dma_buf_phys_vec_to_sgt(attachment, priv-\u003eprovider,\n@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)\n \t * The refcount prevents both.\n \t */\n \tif (priv-\u003evdev) {\n-\t\tdown_write(\u0026priv-\u003evdev-\u003ememory_lock);\n+\t\tdown_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n \t\tlist_del_init(\u0026priv-\u003edmabufs_elm);\n-\t\tup_write(\u0026priv-\u003evdev-\u003ememory_lock);\n+\t\tup_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n \t\tvfio_device_put_registration(\u0026priv-\u003evdev-\u003evdev);\n \t}\n+\tif (priv-\u003evfile)\n+\t\tfput(priv-\u003evfile);\n \tkfree(priv-\u003ephys_vec);\n \tkfree(priv);\n }\n \n static const struct dma_buf_ops vfio_pci_dmabuf_ops = {\n \t.attach = vfio_pci_dma_buf_attach,\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n+\t.mmap = vfio_pci_dma_buf_mmap,\n+#endif\n \t.map_dma_buf = vfio_pci_dma_buf_map,\n \t.unmap_dma_buf = vfio_pci_dma_buf_unmap,\n \t.release = vfio_pci_dma_buf_release,\n };\n \n+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,\n+\t\t\t struct vfio_pci_dma_buf *priv,\n+\t\t\t struct vm_area_struct *vma,\n+\t\t\t unsigned long fault_addr,\n+\t\t\t unsigned int order,\n+\t\t\t unsigned long *out_pfn)\n+{\n+\t/*\n+\t * Given a VMA (start, end, pgoffs) and a fault address,\n+\t * search the corresponding DMABUF's phys_vec[] to find the\n+\t * range representing the address's offset into the VMA, and\n+\t * its PFN. vdev must be the device that the DMABUF priv was\n+\t * exported from; vdev-\u003edmabuf_lock must be held, and priv\n+\t * must not be revoked.\n+\t *\n+\t * The phys_vec[] ranges represent contiguous spans of VAs\n+\t * upwards from the buffer offset 0; the actual PFNs might be\n+\t * in any order, overlap/alias, etc. Calculate an offset of\n+\t * the desired page given VMA start/pgoff and address, then\n+\t * search upwards from 0 to find which span contains it.\n+\t *\n+\t * On success, a valid PFN for a page sized by 'order' is\n+\t * returned into out_pfn.\n+\t *\n+\t * Failure occurs if:\n+\t * - A hugepage would cross the edge of the VMA,\n+\t * - A hugepage isn't entirely contained within a range\n+\t * (including where it straddles the boundary between\n+\t * ranges),\n+\t * - We find a range, but the final PFN isn't aligned to the\n+\t * requested order.\n+\t *\n+\t * Upon failure, -ERANGE is returned and the caller is\n+\t * expected to try again with a smaller order, which will\n+\t * eventually succeed.\n+\t *\n+\t * It's suboptimal if DMABUFs are created with neighbouring\n+\t * ranges that are physically contiguous, since hugepages\n+\t * can't straddle range boundaries. (The construction of the\n+\t * ranges should merge them in this case.)\n+\t *\n+\t * Finally, vma_pgoff_adjust is used with a DMABUF created for\n+\t * a VFIO BAR mmap: a BAR mapped with vm_pgoff \u003e 0 creates a\n+\t * DMABUF such that byte 0 of the VMA corresponds to byte 0 of\n+\t * the DMABUF and byte 'vm_pgoff \u003c\u003c PAGE_SHIFT' into the BAR.\n+\t * To avoid double-offsetting in this scenario, subtracting\n+\t * vma_pgoff_adjust from this (non-zero) vm_pgoff generates\n+\t * the effective offset. This also removes the VFIO region\n+\t * index encoded in vm_pgoff for VFIO BAR mmaps.\n+\t */\n+\n+\tconst unsigned long pagesize = PAGE_SIZE \u003c\u003c order;\n+\tunsigned long vma_off = (vma-\u003evm_pgoff - priv-\u003evma_pgoff_adjust) \u003c\u003c\n+\t\t\t\t PAGE_SHIFT;\n+\tunsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);\n+\tunsigned long rounded_page_end = rounded_page_addr + pagesize;\n+\tunsigned long fault_offset;\n+\tunsigned long fault_offset_end;\n+\tunsigned long range_start_offset = 0;\n+\tunsigned int i;\n+\tint ret;\n+\n+\tif (unlikely(!vdev))\n+\t\treturn -ENODEV;\n+\n+\t/* This prevents the dmabuf revocation state from changing under us */\n+\tlockdep_assert_held(\u0026vdev-\u003edmabuf_lock);\n+\n+\tif (unlikely(priv-\u003evdev != vdev || priv-\u003estatus != VFIO_PCI_DMABUF_OK))\n+\t\treturn -ENODEV;\n+\n+\tif (rounded_page_addr \u003c vma-\u003evm_start || rounded_page_end \u003e vma-\u003evm_end) {\n+\t\tif (order \u003e 0)\n+\t\t\treturn -ERANGE;\n+\n+\t\t/* A fault address outside of the VMA is absurd. */\n+\t\tdev_warn_ratelimited(\n+\t\t\t\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\"Fault addr 0x%lx outside VMA 0x%lx-0x%lx\\n\",\n+\t\t\tfault_addr, vma-\u003evm_start, vma-\u003evm_end);\n+\t\treturn -EFAULT;\n+\t}\n+\n+\t/*\n+\t * fault_offset[_end] is the span within the DMABUF\n+\t * corresponding to the faulting page:\n+\t */\n+\tif (unlikely(check_add_overflow(rounded_page_addr - vma-\u003evm_start,\n+\t\t\t\t\tvma_off, \u0026fault_offset) ||\n+\t\t check_add_overflow(fault_offset, pagesize,\n+\t\t\t\t\t\u0026fault_offset_end)))\n+\t\treturn -EFAULT;\n+\n+\t/*\n+\t * Iterate over ranges in the buffer, summing their lengths:\n+\t * range_start_offset represents the current range's starting\n+\t * offset in the buffer (from 0 upwards).\n+\t *\n+\t * A failure for order == 0 is unexpected, and triggers a\n+\t * fault/warn.\n+\t */\n+\tret = (order == 0) ? -EFAULT : -ERANGE;\n+\n+\tfor (i = 0; i \u003c priv-\u003enr_ranges; i++) {\n+\t\tsize_t range_len = priv-\u003ephys_vec[i].len;\n+\n+\t\t/* Early exit if range starts after the page end */\n+\t\tif (fault_offset_end \u003c= range_start_offset)\n+\t\t\tbreak;\n+\n+\t\tif (fault_offset \u003e= range_start_offset \u0026\u0026\n+\t\t fault_offset_end \u003c= range_start_offset + range_len) {\n+\t\t\t/*\n+\t\t\t * The faulting page is wholly contained\n+\t\t\t * within the span represented by this range,\n+\t\t\t * so validate PFN alignment for the order.\n+\t\t\t * The if() condition ensures the pfn\n+\t\t\t * arithmetic won't overflow.\n+\t\t\t */\n+\t\t\tunsigned long pfn =\n+\t\t\t\t((fault_offset - range_start_offset) +\n+\t\t\t\t priv-\u003ephys_vec[i].paddr) \u003e\u003e PAGE_SHIFT;\n+\n+\t\t\tif (IS_ALIGNED(pfn, 1 \u003c\u003c order)) {\n+\t\t\t\t*out_pfn = pfn;\n+\t\t\t\tret = 0;\n+\t\t\t}\n+\t\t\t/*\n+\t\t\t * Else order \u003e 0; ERANGE retries with smaller\n+\t\t\t * order\n+\t\t\t */\n+\t\t\tbreak;\n+\t\t}\n+\t\trange_start_offset += range_len;\n+\t}\n+\n+\tif (order == 0 \u0026\u0026 ret != 0)\n+\t\t/*\n+\t\t * The address fell outside of the span represented by\n+\t\t * the (concatenated) ranges. As setup of a mapping\n+\t\t * ensures that the VMA is \u003c= the total size of the\n+\t\t * ranges this should never happen. If it does, warn\n+\t\t * and SIGBUS.\n+\t\t */\n+\t\tdev_warn_ratelimited(\n+\t\t\t\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\"No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\\n\",\n+\t\t\tfault_addr, order, vma-\u003evm_start, vma-\u003evm_end,\n+\t\t\tvma-\u003evm_pgoff, priv-\u003enr_ranges, priv-\u003esize);\n+\n+\treturn ret;\n+}\n+\n+/*\n+ * Create a DMABUF corresponding to priv, add it to vdev-\u003edmabufs list\n+ * for tracking (meaning cleanup or revocation will zap it), and take\n+ * a vfio_device registration.\n+ */\n+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vfio_pci_dma_buf *priv, u32 flags)\n+{\n+\tDEFINE_DMA_BUF_EXPORT_INFO(exp_info);\n+\n+\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev))\n+\t\treturn -ENODEV;\n+\n+\texp_info.ops = \u0026vfio_pci_dmabuf_ops;\n+\texp_info.size = priv-\u003esize;\n+\texp_info.flags = flags;\n+\texp_info.priv = priv;\n+\n+\tpriv-\u003edmabuf = dma_buf_export(\u0026exp_info);\n+\tif (IS_ERR(priv-\u003edmabuf)) {\n+\t\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n+\t\treturn PTR_ERR(priv-\u003edmabuf);\n+\t}\n+\n+\tkref_init(\u0026priv-\u003ekref);\n+\tinit_completion(\u0026priv-\u003ecomp);\n+\n+\t/* dma_buf_put() now frees priv */\n+\tINIT_LIST_HEAD(\u0026priv-\u003edmabufs_elm);\n+\n+\t/*\n+\t * dmabuf_lock synchronises access (R) or updates (W) to the\n+\t * vdev-\u003edmabufs list and to bars_revoked (see below). The\n+\t * revocation state of DMABUF elements in the list is written\n+\t * holding both dmabuf_lock(W) and resv, and tested with\n+\t * either.\n+\t *\n+\t * (memory_lock, if held -\u003e) dmabuf_lock -\u003e resv\n+\t *\n+\t * NOTE: memory_lock is strictly avoided here, to avoid a\n+\t * dependency on memory_lock when mmap_lock is held, when\n+\t * mmap() leads to export. vfio-pci variant drivers are\n+\t * permitted to hold memory_lock across actions that might\n+\t * fault (such as user access); a deadlock could result when\n+\t * that fault path attempts to take mmap_lock (if held by an\n+\t * export waiting for memory_lock).\n+\t *\n+\t * vdev-\u003ebars_revoked tracks the BAR revocation status updated\n+\t * via vfio_pci_dma_buf_move(), so the initial DMABUF state\n+\t * follows the same criteria that later update the DMABUF\n+\t * state (BAR zap, etc.).\n+\t */\n+\tlockdep_assert_not_held(\u0026vdev-\u003ememory_lock);\n+\n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\tpriv-\u003estatus = vdev-\u003ebars_revoked ? VFIO_PCI_DMABUF_REVOKED :\n+\t\tVFIO_PCI_DMABUF_OK;\n+\tlist_add_tail(\u0026priv-\u003edmabufs_elm, \u0026vdev-\u003edmabufs);\n+\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n+\n+\treturn 0;\n+}\n+\n+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vm_area_struct *vma,\n+\t\t\t\t u64 phys_start, u64 req_len,\n+\t\t\t\t unsigned int res_index)\n+{\n+\tstruct vfio_pci_dma_buf *priv;\n+\tunsigned long vma_pgoff = vma-\u003evm_pgoff \u0026 (VFIO_PCI_OFFSET_MASK \u003e\u003e PAGE_SHIFT);\n+\tchar *bufname;\n+\tint ret;\n+\n+\tpriv = kzalloc_obj(*priv);\n+\tif (!priv)\n+\t\treturn -ENOMEM;\n+\n+\tpriv-\u003ephys_vec = kzalloc_obj(*priv-\u003ephys_vec);\n+\tif (!priv-\u003ephys_vec) {\n+\t\tret = -ENOMEM;\n+\t\tgoto err_free_priv;\n+\t}\n+\n+\t/*\n+\t * Debug name: The absolute maximum size of the name\n+\t * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.\n+\t */\n+\tbufname = kasprintf(GFP_KERNEL, \"vfio:%s/%x\",\n+\t\t\t pci_name(vdev-\u003epdev),\n+\t\t\t res_index);\n+\n+\tif (!bufname) {\n+\t\tret = -ENOMEM;\n+\t\tgoto err_free_phys;\n+\t}\n+\n+\t/*\n+\t * The DMABUF begins from the mmap()'s BAR offset, i.e. the\n+\t * start of the VMA corresponds to byte 0 of the DMABUF and\n+\t * byte (vma_pgoff \u003c\u003c PAGE_SHIFT) of the BAR.\n+\t *\n+\t * vfio_pci_dma_buf_find_pfn() reverses this offset using\n+\t * vma_pgoff_adjust, so that ultimately a fault's offset from\n+\t * the start of the _VMA_ has a consistent usage whether the\n+\t * VMA originates from an mmap() of the VFIO device here or a\n+\t * direct DMABUF mmap(). Note vma_pgoff_adjust also includes\n+\t * the encoded VFIO region index, which cancels out the index\n+\t * encoded in vm_pgoff.\n+\t */\n+\tpriv-\u003evdev = vdev;\n+\tpriv-\u003esize = req_len;\n+\tpriv-\u003enr_ranges = 1;\n+\tpriv-\u003evma_pgoff_adjust = vma-\u003evm_pgoff;\n+\n+\t/*\n+\t * The provider can be NULL _iff_ the DMABUF feature isn't\n+\t * supported, because it's only used by DMABUF import and\n+\t * attach is prohibited if the feature isn't present.\n+\t */\n+\tpriv-\u003eprovider = pcim_p2pdma_provider(vdev-\u003epdev, res_index);\n+\tif (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) \u0026\u0026 !priv-\u003eprovider) {\n+\t\tret = -EINVAL;\n+\t\tgoto err_free_name;\n+\t}\n+\n+\tpriv-\u003ephys_vec[0].paddr = phys_start + ((u64)vma_pgoff \u003c\u003c PAGE_SHIFT);\n+\tpriv-\u003ephys_vec[0].len = priv-\u003esize;\n+\n+\tret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);\n+\tif (ret)\n+\t\tgoto err_free_name;\n+\n+\tif (dma_buf_set_name(priv-\u003edmabuf, bufname)) {\n+\t\tdev_dbg_ratelimited(\u0026vdev-\u003epdev-\u003edev,\n+\t\t\t\t \"Failed to set map name '%s'\\n\",\n+\t\t\t\t bufname);\n+\t\tkfree(bufname);\n+\t}\n+\n+\t/*\n+\t * Ownership of the DMABUF file transfers to the VMA so that\n+\t * other users can locate the DMABUF via a VA. Ownership of\n+\t * the original VFIO device file being mmap()ed transfers to\n+\t * priv, and is put when the DMABUF is released. This\n+\t * intentionally does not use get_file()/vma_set_file()\n+\t * because the references are already held, and ownership\n+\t * moves.\n+\t */\n+\tpriv-\u003evfile = vma-\u003evm_file;\n+\tvma-\u003evm_file = priv-\u003edmabuf-\u003efile;\n+\tvma-\u003evm_private_data = priv;\n+\n+\treturn 0;\n+\n+err_free_name:\n+\tkfree(bufname);\n+err_free_phys:\n+\tkfree(priv-\u003ephys_vec);\n+err_free_priv:\n+\tkfree(priv);\n+\treturn ret;\n+}\n+\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n /*\n * This is a temporary \"private interconnect\" between VFIO DMABUF and iommufd.\n * It allows the two co-operating drivers to exchange the physical address of\n@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,\n \t\treturn -EOPNOTSUPP;\n \n \tpriv = attachment-\u003edmabuf-\u003epriv;\n-\tif (priv-\u003erevoked)\n+\tif (priv-\u003estatus != VFIO_PCI_DMABUF_OK)\n \t\treturn -ENODEV;\n \n \t/* More than one range to iommufd will require proper DMABUF support */\n@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n {\n \tstruct vfio_device_feature_dma_buf get_dma_buf = {};\n \tstruct vfio_region_dma_range *dma_ranges;\n-\tDEFINE_DMA_BUF_EXPORT_INFO(exp_info);\n \tstruct vfio_pci_dma_buf *priv;\n \tsize_t length;\n \tint ret;\n@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \tkfree(dma_ranges);\n \tdma_ranges = NULL;\n \n-\tif (!vfio_device_try_get_registration(\u0026vdev-\u003evdev)) {\n-\t\tret = -ENODEV;\n+\tret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);\n+\tif (ret)\n \t\tgoto err_free_phys;\n-\t}\n-\n-\texp_info.ops = \u0026vfio_pci_dmabuf_ops;\n-\texp_info.size = priv-\u003esize;\n-\texp_info.flags = get_dma_buf.open_flags;\n-\texp_info.priv = priv;\n-\n-\tpriv-\u003edmabuf = dma_buf_export(\u0026exp_info);\n-\tif (IS_ERR(priv-\u003edmabuf)) {\n-\t\tret = PTR_ERR(priv-\u003edmabuf);\n-\t\tgoto err_dev_put;\n-\t}\n-\n-\tkref_init(\u0026priv-\u003ekref);\n-\tinit_completion(\u0026priv-\u003ecomp);\n-\n-\t/* dma_buf_put() now frees priv */\n-\tINIT_LIST_HEAD(\u0026priv-\u003edmabufs_elm);\n-\tdown_write(\u0026vdev-\u003ememory_lock);\n-\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\tpriv-\u003erevoked = !__vfio_pci_memory_enabled(vdev);\n-\tlist_add_tail(\u0026priv-\u003edmabufs_elm, \u0026vdev-\u003edmabufs);\n-\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\tup_write(\u0026vdev-\u003ememory_lock);\n-\n \t/*\n \t * dma_buf_fd() consumes the reference, when the file closes the dmabuf\n \t * will be released.\n@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \n \treturn ret;\n \n-err_dev_put:\n-\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n err_free_phys:\n \tkfree(priv-\u003ephys_vec);\n err_free_priv:\n@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \tkfree(dma_ranges);\n \treturn ret;\n }\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\n+\n+/*\n+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives\n+ * the guarantee that all future map/attach attempts will fail no\n+ * matter what, whereas REVOKED can transition back to OK.\n+ */\n+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,\n+\t\t\t\t\tenum vfio_pci_dma_buf_status new_status)\n+{\n+\tbool was_revoked;\n+\n+\t/*\n+\t * Changes to the DMABUF's revocation status are synchronised\n+\t * using dmabuf_lock:\n+\t */\n+\tlockdep_assert_held_write(\u0026priv-\u003evdev-\u003edmabuf_lock);\n+\n+\t/* If DEAD, state can no longer change */\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_DEAD ||\n+\t priv-\u003estatus == new_status)\n+\t\treturn;\n+\n+\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n+\twas_revoked = (priv-\u003estatus == VFIO_PCI_DMABUF_REVOKED);\n+\n+\tif (new_status != VFIO_PCI_DMABUF_OK) {\n+\t\tpriv-\u003estatus = new_status;\n+\n+\t\tif (was_revoked) {\n+\t\t\t/*\n+\t\t\t * A REVOKED buffer is being marked DEAD.\n+\t\t\t * invalidate_mappings/unmap wait happened\n+\t\t\t * when it became REVOKED, don't wait again.\n+\t\t\t */\n+\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\t\treturn;\n+\t\t}\n+\t\tdma_buf_invalidate_mappings(priv-\u003edmabuf);\n+\t\tdma_resv_wait_timeout(priv-\u003edmabuf-\u003eresv,\n+\t\t\t\t DMA_RESV_USAGE_BOOKKEEP, false,\n+\t\t\t\t MAX_SCHEDULE_TIMEOUT);\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t\tkref_put(\u0026priv-\u003ekref, vfio_pci_dma_buf_done);\n+\t\twait_for_completion(\u0026priv-\u003ecomp);\n+\t\tunmap_mapping_range(priv-\u003edmabuf-\u003efile-\u003ef_mapping,\n+\t\t\t\t 0, 0, true);\n+\t\t/*\n+\t\t * Re-arm the registered kref reference and the\n+\t\t * completion so the post-revoke state matches the\n+\t\t * post-creation state. An un-revoke followed by a\n+\t\t * new mapping needs the kref to be non-zero before\n+\t\t * kref_get(), and vfio_pci_dma_buf_cleanup()\n+\t\t * delegates its drain back through this revoke\n+\t\t * path on a possibly-already-revoked dma-buf.\n+\t\t */\n+\t\tkref_init(\u0026priv-\u003ekref);\n+\t\treinit_completion(\u0026priv-\u003ecomp);\n+\t} else {\n+\t\tpriv-\u003estatus = VFIO_PCI_DMABUF_OK;\n+\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n+\t}\n+}\n \n void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)\n {\n@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)\n \n \tlockdep_assert_held_write(\u0026vdev-\u003ememory_lock);\n \n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tvdev-\u003ebars_revoked = revoked;\n \tlist_for_each_entry_safe(priv, tmp, \u0026vdev-\u003edmabufs, dmabufs_elm) {\n \t\tif (!get_file_active(\u0026priv-\u003edmabuf-\u003efile))\n \t\t\tcontinue;\n-\n-\t\tif (priv-\u003erevoked != revoked) {\n-\t\t\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\t\t\tif (revoked)\n-\t\t\t\tpriv-\u003erevoked = true;\n-\t\t\tdma_buf_invalidate_mappings(priv-\u003edmabuf);\n-\t\t\tdma_resv_wait_timeout(priv-\u003edmabuf-\u003eresv,\n-\t\t\t\t\t DMA_RESV_USAGE_BOOKKEEP, false,\n-\t\t\t\t\t MAX_SCHEDULE_TIMEOUT);\n-\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\t\t\tif (revoked) {\n-\t\t\t\tkref_put(\u0026priv-\u003ekref, vfio_pci_dma_buf_done);\n-\t\t\t\twait_for_completion(\u0026priv-\u003ecomp);\n-\t\t\t\t/*\n-\t\t\t\t * Re-arm the registered kref reference and the\n-\t\t\t\t * completion so the post-revoke state matches the\n-\t\t\t\t * post-creation state. An un-revoke followed by a\n-\t\t\t\t * new mapping needs the kref to be non-zero before\n-\t\t\t\t * kref_get(), and vfio_pci_dma_buf_cleanup()\n-\t\t\t\t * delegates its drain back through this revoke\n-\t\t\t\t * path on a possibly-already-revoked dma-buf.\n-\t\t\t\t */\n-\t\t\t\tkref_init(\u0026priv-\u003ekref);\n-\t\t\t\treinit_completion(\u0026priv-\u003ecomp);\n-\t\t\t} else {\n-\t\t\t\tdma_resv_lock(priv-\u003edmabuf-\u003eresv, NULL);\n-\t\t\t\tpriv-\u003erevoked = false;\n-\t\t\t\tdma_resv_unlock(priv-\u003edmabuf-\u003eresv);\n-\t\t\t}\n-\t\t}\n+\t\tvfio_pci_dma_buf_set_status(priv, revoked ?\n+\t\t\t\t\t VFIO_PCI_DMABUF_REVOKED :\n+\t\t\t\t\t VFIO_PCI_DMABUF_OK);\n \t\tfput(priv-\u003edmabuf-\u003efile);\n \t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n }\n \n void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n \t */\n \tvfio_pci_dma_buf_move(vdev, true);\n \n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n \tlist_for_each_entry_safe(priv, tmp, \u0026vdev-\u003edmabufs, dmabufs_elm) {\n \t\tif (!get_file_active(\u0026priv-\u003edmabuf-\u003efile))\n \t\t\tcontinue;\n \n \t\tlist_del_init(\u0026priv-\u003edmabufs_elm);\n-\t\tpriv-\u003evdev = NULL;\n+\t\tWRITE_ONCE(priv-\u003evdev, NULL);\n \t\tvfio_device_put_registration(\u0026vdev-\u003evdev);\n \t\tfput(priv-\u003edmabuf-\u003efile);\n \t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n \tup_write(\u0026vdev-\u003ememory_lock);\n }\n+\n+#ifdef CONFIG_VFIO_PCI_DMABUF\n+int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz)\n+{\n+\tstruct vfio_device_feature_dma_buf_revoke db_revoke;\n+\tstruct vfio_pci_dma_buf *priv;\n+\tstruct dma_buf *dmabuf;\n+\tint ret;\n+\n+\tif (!vdev-\u003epci_ops || !vdev-\u003epci_ops-\u003eget_dmabuf_phys)\n+\t\treturn -EOPNOTSUPP;\n+\n+\tret = vfio_check_feature(flags, argsz,\n+\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n+\t\t\t\t sizeof(db_revoke));\n+\tif (ret != 1)\n+\t\treturn ret;\n+\n+\tif (copy_from_user(\u0026db_revoke, arg, sizeof(db_revoke)))\n+\t\treturn -EFAULT;\n+\n+\tdmabuf = dma_buf_get(db_revoke.dmabuf_fd);\n+\tif (IS_ERR(dmabuf))\n+\t\treturn PTR_ERR(dmabuf);\n+\n+\tpriv = dmabuf-\u003epriv;\n+\t/*\n+\t * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_\n+\t * relates to the VFIO device it was provided with.\n+\t *\n+\t * If the DMABUF relates to this vdev then priv-\u003evdev is\n+\t * stable because this open fd prevents cleanup.\n+\t *\n+\t * If it relates to a different vdev, reading priv-\u003evdev might\n+\t * race with a concurrent cleanup on that device. But if so,\n+\t * it points to a non-matching vdev or NULL and is unusable\n+\t * either way.\n+\t */\n+\tif (dmabuf-\u003eops != \u0026vfio_pci_dmabuf_ops ||\n+\t READ_ONCE(priv-\u003evdev) != vdev) {\n+\t\tret = -ENODEV;\n+\t\tgoto out_put_buf;\n+\t}\n+\n+\t/*\n+\t * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()\n+\t * from getting it and then blocking all devices in the dev_set behind\n+\t * this revoke's drain.\n+\t */\n+\tdown_read(\u0026vdev-\u003ememory_lock);\n+\tdown_write(\u0026vdev-\u003edmabuf_lock);\n+\tif (priv-\u003estatus == VFIO_PCI_DMABUF_DEAD) {\n+\t\tret = -EBADFD;\n+\t} else {\n+\t\tvfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);\n+\t\tret = 0;\n+\t}\n+\tup_write(\u0026vdev-\u003edmabuf_lock);\n+\tup_read(\u0026vdev-\u003ememory_lock);\n+\n+out_put_buf:\n+\tdma_buf_put(dmabuf);\n+\n+\treturn ret;\n+}\n+#endif /* CONFIG_VFIO_PCI_DMABUF */\ndiff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h\nindex 4e7162234a2eb..ca12221af5553 100644\n--- a/drivers/vfio/pci/vfio_pci_priv.h\n+++ b/drivers/vfio/pci/vfio_pci_priv.h\n@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {\n \tbool\t\t\ttest_mem;\n };\n \n+enum vfio_pci_dma_buf_status {\n+\tVFIO_PCI_DMABUF_OK = 0,\n+\tVFIO_PCI_DMABUF_REVOKED = 1,\n+\tVFIO_PCI_DMABUF_DEAD = 2,\n+};\n+\n+struct vfio_pci_dma_buf {\n+\tstruct dma_buf *dmabuf;\n+\tstruct vfio_pci_core_device *vdev;\n+\tstruct list_head dmabufs_elm;\n+\tsize_t size;\n+\tstruct phys_vec *phys_vec;\n+\tstruct p2pdma_provider *provider;\n+\tstruct file *vfile;\n+\tu32 nr_ranges;\n+\tstruct kref kref;\n+\tstruct completion comp;\n+\tunsigned long vma_pgoff_adjust;\n+\tenum vfio_pci_dma_buf_status status;\n+};\n+\n bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);\n void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);\n \n@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);\n int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,\n \t\t\t pci_power_t state);\n \n-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);\n+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);\n+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);\n u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);\n void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,\n \t\t\t\t\tu16 cmd);\n@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)\n \treturn (pdev-\u003eclass \u003e\u003e 8) == PCI_CLASS_DISPLAY_VGA;\n }\n \n+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,\n+\t\t\t struct vfio_pci_dma_buf *priv,\n+\t\t\t struct vm_area_struct *vma,\n+\t\t\t unsigned long address,\n+\t\t\t unsigned int order,\n+\t\t\t unsigned long *out_pfn);\n+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n+\t\t\t\t struct vm_area_struct *vma,\n+\t\t\t\t u64 phys_start, u64 req_len,\n+\t\t\t\t unsigned int res_index);\n+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);\n+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);\n+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);\n+\n #ifdef CONFIG_VFIO_PCI_DMABUF\n int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n \t\t\t\t struct vfio_device_feature_dma_buf __user *arg,\n \t\t\t\t size_t argsz);\n-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);\n-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);\n+int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz);\n #else\n static inline int\n vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n {\n \treturn -ENOTTY;\n }\n-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n-{\n-}\n-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,\n-\t\t\t\t\t bool revoked)\n+static inline int vfio_pci_core_feature_dma_buf_revoke(\n+\tstruct vfio_pci_core_device *vdev, u32 flags,\n+\tstruct vfio_device_feature_dma_buf_revoke __user *arg,\n+\tsize_t argsz)\n {\n+\treturn -ENOTTY;\n }\n #endif\n \ndiff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h\nindex d15b2b31d3c91..0f88132c06546 100644\n--- a/include/linux/dma-buf.h\n+++ b/include/linux/dma-buf.h\n@@ -342,12 +342,14 @@ struct dma_buf {\n \t/**\n \t * @name:\n \t *\n-\t * Userspace-provided name. Default value is NULL. If not NULL,\n-\t * length cannot be longer than DMA_BUF_NAME_LEN, including NIL\n-\t * char. Useful for accounting and debugging. Read/Write accesses\n-\t * are protected by @name_lock\n-\t *\n-\t * See the IOCTLs DMA_BUF_SET_NAME or DMA_BUF_SET_NAME_A/B\n+\t * Exporter or userspace-provided name. Default value is\n+\t * NULL. If not NULL, length cannot be longer than\n+\t * DMA_BUF_NAME_LEN, including NIL char. Useful for accounting\n+\t * and debugging. Read/Write accesses are protected by\n+\t * @name_lock\n+\t *\n+\t * See dma_buf_set_name(), and the IOCTLs DMA_BUF_SET_NAME or\n+\t * DMA_BUF_SET_NAME_A/B\n \t */\n \tconst char *name;\n \n@@ -571,6 +573,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);\n struct dma_buf *dma_buf_get(int fd);\n void dma_buf_put(struct dma_buf *dmabuf);\n \n+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);\n+\n struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,\n \t\t\t\t\tenum dma_data_direction);\n void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,\ndiff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h\nindex 9a1674c152aa2..44891fdb7c76e 100644\n--- a/include/linux/vfio_pci_core.h\n+++ b/include/linux/vfio_pci_core.h\n@@ -129,11 +129,13 @@ struct vfio_pci_core_device {\n \tbool\t\t\tdisable_idle_d3:1;\n \tbool\t\t\tnointxmask:1;\n \tbool\t\t\tdisable_vga:1;\n+\tbool\t\t\tzap_bars_on_revoke:1;\n \t/* Flags modified at runtime - dedicated storage unit */\n \tbool\t\t\tneeds_reset;\n \tbool\t\t\tpm_intx_masked;\n \tbool\t\t\tpm_runtime_engaged;\n \tbool\t\t\tsriov_active;\n+\tbool\t\t\tbars_revoked;\n \tstruct pci_saved_state\t*pci_saved_state;\n \tstruct pci_saved_state\t*pm_save;\n \tint\t\t\tioeventfds_nr;\n@@ -148,6 +150,7 @@ struct vfio_pci_core_device {\n \tstruct vfio_pci_core_device\t*sriov_pf_core_dev;\n \tstruct notifier_block\tnb;\n \tstruct rw_semaphore\tmemory_lock;\n+\tstruct rw_semaphore\tdmabuf_lock;\n \tstruct list_head\tdmabufs;\n };\n \ndiff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h\nindex e41437fa17ad0..d3c6057983e09 100644\n--- a/include/uapi/linux/vfio.h\n+++ b/include/uapi/linux/vfio.h\n@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {\n \n #define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13\n \n+/**\n+ * Given a DMABUF fd previously exported from the same device by\n+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that\n+ * access to the corresponding DMABUF is immediately revoked. On\n+ * successful return, the buffer is no longer accessible through any\n+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future\n+ * mmap()s and map/attach requests from any new/existing importer.\n+ *\n+ * Return: 0 on success, -1 and errno is set on failure:\n+ *\n+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.\n+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.\n+ * ENODEV: The DMABUF was not exported from this device.\n+ * EBADFD: The DMABUF is already revoked by this feature.\n+ *\n+ * Additionally, common errors can occur: EFAULT accessing the struct,\n+ * or EINVAL requesting an unsupported feature op.\n+ */\n+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14\n+\n+struct vfio_device_feature_dma_buf_revoke {\n+\t__s32\tdmabuf_fd;\n+};\n+\n /* -------- API for Type1 VFIO IOMMU -------- */\n \n /**\n"
} |
| 2/1 | 2026/10/06 20:14 | agent | patch-evaluator |
2mModel:core Args: null Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER",
"ZONE_DEVICE",
"MEMORY_HOTPLUG",
"MEMORY_HOTREMOVE"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_mmap",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_dma_buf_cleanup",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch series modifies core DMA-BUF code (exporting and refactoring dma_buf_set_name and its ioctl handling) and heavily reworks VFIO PCI memory management and revocation. In particular, VFIO PCI BAR mmaps are now backed by DMA-BUFs, page fault handling is rewritten to navigate complex lock ordering across memory_lock, dmabuf_lock, and reservation locks, and a new feature ioctl VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE is introduced. These code paths are fully reachable in virtualized QEMU environments with emulated PCI devices and warrant targeted fuzzing.",
"WorthFuzzing": true
}Instruction: You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.
IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.
================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.
- REACHABILITY IS THE PRIMARY GATE:
Fuzzing can only discover bugs in code that can actually execute in standard virtualized
environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
regardless of whether it adds assertions or complex logic.
- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
Patch authors routinely label changes as "cleanups", "refactorings", or state
"No functional change intended". Do NOT take these claims at face value.
Code refactorings that rearrange logic, introduce helper functions, or alter state management
in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
If reachable executable code is modified or refactored, it MUST be fuzzed.
- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
Even if the author believes the invariant always holds, fuzzing is essential to verify whether
an unusual sequence of operations can violate it.
================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:
- Non-kernel and non-executable changes:
* Modifications to Documentation/, comments, or spelling fixes.
* User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
that do not affect the compiled kernel image (vmlinux) or kernel modules.
* Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
that do not alter control flow or data structures.
* Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
* Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
* Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
physical PCI hot-unplug or manual sysfs driver unbinding.
* Code paths exclusive to architectures other than the target architecture.
================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
* Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
* Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
* Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
* Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).
================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:
- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
Do NOT list generic, ubiquitous functions called by almost every program in the corpus
(including, but not limited to: general memory allocators and deallocators, page fault
and trap handlers, or core synchronization primitives; this is not an exhaustive list).
Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
which severely dilutes fuzzing effort away from the actual changes.
- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
List functions that specifically implement the logic being added or altered, or direct API entrypoints
for the subsystem feature under review.
- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).
================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 8be538d0a9753d071d10cdbcaec6469f392184a3
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 20:14:46 2026 +0000
syz-cluster: applied patch under review
diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
index 4c9add51f9ef5..b99da4ff2ef62 100644
--- a/drivers/dma-buf/dma-buf.c
+++ b/drivers/dma-buf/dma-buf.c
@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)
return events;
}
-/**
- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.
- * It could support changing the name of the dma-buf if the same
- * piece of memory is used for multiple purpose between different devices.
- *
- * @dmabuf: [in] dmabuf buffer that will be renamed.
- * @buf: [in] A piece of userspace memory that contains the name of
- * the dma-buf.
- *
- * Returns 0 on success. If the dma-buf buffer is already attached to
- * devices, return -EBUSY.
- *
- */
-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)
-{
- char *name = strndup_user(buf, DMA_BUF_NAME_LEN);
-
- if (IS_ERR(name))
- return PTR_ERR(name);
-
- spin_lock(&dmabuf->name_lock);
- kfree(dmabuf->name);
- dmabuf->name = name;
- spin_unlock(&dmabuf->name_lock);
-
- return 0;
-}
-
#if IS_ENABLED(CONFIG_SYNC_FILE)
static long dma_buf_export_sync_file(struct dma_buf *dmabuf,
void __user *user_data)
@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,
return ret;
case DMA_BUF_SET_NAME_A:
- case DMA_BUF_SET_NAME_B:
- return dma_buf_set_name(dmabuf, (const char __user *)arg);
+ case DMA_BUF_SET_NAME_B: {
+ char *name = strndup_user((const char __user *)arg,
+ DMA_BUF_NAME_LEN);
+
+ if (IS_ERR(name))
+ return PTR_ERR(name);
+
+ ret = dma_buf_set_name(dmabuf, name);
+ if (ret)
+ kfree(name);
+
+ return ret;
+ }
#if IS_ENABLED(CONFIG_SYNC_FILE)
case DMA_BUF_IOCTL_EXPORT_SYNC_FILE:
@@ -874,6 +857,42 @@ void dma_buf_put(struct dma_buf *dmabuf)
}
EXPORT_SYMBOL_NS_GPL(dma_buf_put, "DMA_BUF");
+/**
+ * dma_buf_set_name - Set a dmabuf's name
+ * @dmabuf: [in] dmabuf buffer that will be renamed.
+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or
+ * similar. This takes ownership of the allocation
+ * on success, which will be kfree()d when the
+ * dmabuf is released or a new name assigned.
+ *
+ * Intended to be used by the exporter to set a name for debug
+ * purposes. This can also change an existing name if the same piece
+ * of memory is used for multiple purposes over time. Importers
+ * should NOT use this function to overwrite the name.
+ *
+ * Return: 0 on success, -EINVAL if the name is NULL, or -E2BIG if the
+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.
+ */
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)
+{
+ if (!name)
+ return -EINVAL;
+
+ /* dmabuffs_dname() won't use the string if the length
+ * (including terminator) exceeds DMA_BUF_NAME_LEN:
+ */
+ if (strlen(name) >= DMA_BUF_NAME_LEN)
+ return -E2BIG;
+
+ spin_lock(&dmabuf->name_lock);
+ kfree(dmabuf->name);
+ dmabuf->name = name;
+ spin_unlock(&dmabuf->name_lock);
+
+ return 0;
+}
+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, "DMA_BUF");
+
static int dma_buf_wrap_sg_table(struct sg_table **sg_table)
{
struct scatterlist *to_sg, *from_sg;
@@ -1004,6 +1023,11 @@ dma_buf_pin_on_map(struct dma_buf_attachment *attach)
* functions:
*
* - dma_buf_invalidate_mappings()
+ *
+ * 4. Exporters must not hold the dma-buf reservation lock when calling these
+ * functions:
+ *
+ * - dma_buf_set_name()
*/
/**
diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig
index 296bf01e185ec..c6d6fb09af86c 100644
--- a/drivers/vfio/pci/Kconfig
+++ b/drivers/vfio/pci/Kconfig
@@ -6,6 +6,7 @@ config VFIO_PCI_CORE
tristate
select VFIO_VIRQFD
select IRQ_BYPASS_MANAGER
+ select DMA_SHARED_BUFFER
config VFIO_PCI_INTX
def_bool y if !S390
@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM
To enable s390x KVM vfio-pci extensions, say Y.
config VFIO_PCI_DMABUF
- def_bool y if VFIO_PCI_CORE && PCI_P2PDMA && DMA_SHARED_BUFFER
+ def_bool y if PCI_P2PDMA
+ depends on VFIO_PCI_CORE
source "drivers/vfio/pci/mlx5/Kconfig"
diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile
index 6138f1bf241df..881452ea89be0 100644
--- a/drivers/vfio/pci/Makefile
+++ b/drivers/vfio/pci/Makefile
@@ -1,8 +1,7 @@
# SPDX-License-Identifier: GPL-2.0-only
-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o
+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o
vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o
-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o
obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o
vfio-pci-y := vfio_pci.o
diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
index 86362ec424a50..14622556355eb 100644
--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);
struct pci_dev *pdev = to_pci_dev(core_vdev->dev);
struct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);
+ int ret;
hisi_acc_vdev->vf_id = pci_iov_vf_id(pdev) + 1;
hisi_acc_vdev->pf_qm = pf_qm;
@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
core_vdev->migration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;
core_vdev->mig_ops = &hisi_acc_vfio_pci_migrn_state_ops;
- return vfio_pci_core_init_dev(core_vdev);
+ ret = vfio_pci_core_init_dev(core_vdev);
+ if (ret)
+ return ret;
+ /*
+ * hisi_acc_vfio_pci_mmap() calls down to
+ * vfio_pci_core_mmap(), so BAR mappings are still
+ * DMABUF-backed. They don't require a zap on revoke, so opt
+ * out:
+ */
+ hisi_acc_vdev->core_device.zap_bars_on_revoke = false;
+
+ return 0;
}
static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index 9914f3ac69aef..cef337f4e8f2e 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
virt_mem = !!(le16_to_cpu(*virt_cmd) & PCI_COMMAND_MEMORY);
new_mem = !!(new_cmd & PCI_COMMAND_MEMORY);
- if (!new_mem) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (!new_mem)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
/*
* If the user is writing mem/io enable (new_mem/io) and we
@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
*virt_cmd |= cpu_to_le16(new_cmd & mask);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state)
{
- if (state >= PCI_D3hot) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (state >= PCI_D3hot)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
vfio_pci_set_power_state(vdev, state);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_EXP_DEVCAP_FLR)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_AF_CAP_FLR) && (cap & PCI_AF_CAP_TP)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d875..68e582ad38463 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -13,6 +13,8 @@
#include <linux/aperture.h>
#include <linux/debugfs.h>
#include <linux/device.h>
+#include <linux/dma-buf.h>
+#include <linux/dma-resv.h>
#include <linux/eventfd.h>
#include <linux/file.h>
#include <linux/interrupt.h>
@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
* The vdev power related flags are protected with 'memory_lock'
* semaphore.
*/
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
if (vdev->pm_runtime_engaged) {
up_write(&vdev->memory_lock);
@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)
down_write(&vdev->memory_lock);
__vfio_pci_runtime_pm_exit(vdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)
*/
down_write(&vdev->memory_lock);
if (vdev->pm_wake_eventfd_ctx) {
- eventfd_signal(vdev->pm_wake_eventfd_ctx);
+ struct eventfd_ctx *ctx = vdev->pm_wake_eventfd_ctx;
+
+ vdev->pm_wake_eventfd_ctx = NULL;
__vfio_pci_runtime_pm_exit(vdev);
+ if (__vfio_pci_memory_enabled(vdev))
+ vfio_pci_unrevoke_bars(vdev);
+ eventfd_signal(ctx);
+ eventfd_ctx_put(ctx);
}
up_write(&vdev->memory_lock);
@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
vdev->has_vga = true;
vfio_pci_core_map_bars(vdev);
+ vdev->bars_revoked = false;
return 0;
@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
return ret;
}
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);
+
static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
void __user *arg)
{
@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
if (!vdev->reset_works)
return -EINVAL;
- vfio_pci_zap_and_down_write_memory_lock(vdev);
+ down_write(&vdev->memory_lock);
/*
* This function can be invoked while the power state is non-D0. If
@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
* have NoSoftRst-, the reset function can cause the PCI config space
* reset without restoring the original state (saved locally in
* 'vdev->pm_save').
+ *
+ * The zap is done after making the device accessible in D0,
+ * because a DMABUF importer could access the device as part
+ * of its revocation cleanup.
*/
vfio_pci_set_power_state(vdev, PCI_D0);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_revoke_bars(vdev);
+
ret = pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
return ret;
@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
return vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);
case VFIO_DEVICE_FEATURE_ZPCI_ERROR:
return vfio_pci_zdev_feature_err(device, flags, arg, argsz);
+ case VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:
+ return vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);
default:
return -ENOTTY;
}
@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu
}
EXPORT_SYMBOL_GPL(vfio_pci_core_write);
-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)
{
- struct vfio_device *core_vdev = &vdev->vdev;
- loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
- loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
- loff_t len = end - start;
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, true);
- unmap_mapping_range(core_vdev->inode->i_mapping, start, len, true);
+ /*
+ * If a driver could possibly create BAR mappings in the
+ * vdev's address_space, do an additional zap on revoke. See
+ * vfio_pci_core_init_dev().
+ */
+ if (vdev->zap_bars_on_revoke) {
+ struct vfio_device *core_vdev = &vdev->vdev;
+ loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
+ loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
+ loff_t len = end - start;
+
+ unmap_mapping_range(core_vdev->inode->i_mapping,
+ start, len, true);
+ }
}
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)
{
down_write(&vdev->memory_lock);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
+}
+
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)
+{
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, false);
}
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)
@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c
up_write(&vdev->memory_lock);
}
-static unsigned long vma_to_pfn(struct vm_area_struct *vma)
-{
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- int index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
- u64 pgoff;
-
- pgoff = vma->vm_pgoff &
- ((1U << (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);
-
- return (pci_resource_start(vdev->pdev, index) >> PAGE_SHIFT) + pgoff;
-}
-
vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,
struct vm_fault *vmf,
unsigned long pfn,
@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,
unsigned int order)
{
struct vm_area_struct *vma = vmf->vma;
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- unsigned long addr = vmf->address & ~((PAGE_SIZE << order) - 1);
- unsigned long pgoff = linear_page_delta(vma, addr);
- unsigned long pfn = vma_to_pfn(vma) + pgoff;
- vm_fault_t ret = VM_FAULT_FALLBACK;
-
- if (is_aligned_for_order(vma, addr, pfn, order)) {
- scoped_guard(rwsem_read, &vdev->memory_lock)
- ret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);
+ struct vfio_pci_dma_buf *priv = vma->vm_private_data;
+ struct vfio_pci_core_device *vdev;
+ unsigned long pfn = 0;
+ vm_fault_t ret = VM_FAULT_SIGBUS;
+
+ /*
+ * The only thing this can rely on is that the DMABUF relating
+ * to the VMA's vm_file exists (priv).
+ *
+ * A DMABUF for a VFIO device fd mmap() holds a reference to
+ * the original VFIO device fd, but an explicitly-exported
+ * DMABUF does not. The original fd might have closed,
+ * meaning this fault can race with
+ * vfio_pci_dma_buf_cleanup(), meaning the buffer could have
+ * been revoked (in which case priv->vdev might be NULL), and
+ * the VFIO device registration might have been dropped.
+ *
+ * With the goal of taking vdev locks in a world where vdev
+ * might not still exist:
+ *
+ * 1. Take the resv lock on the DMABUF:
+ * - If racing cleanup got in first, the buffer is revoked;
+ * stop/exit if so.
+ * - If we got in first, the buffer is not revoked so vdev is
+ * non-NULL, accessible, and cleanup _has not yet put the
+ * VFIO device registration_. So, the device refcount must
+ * be >0.
+ *
+ * 2. Take vfio_device registration (refcount guaranteed >0
+ * hereafter).
+ *
+ * 3. Unlock the DMABUF's resv lock:
+ * - A racing cleanup can now complete.
+ * - But, the device refcount >0, meaning the vfio_device
+ * (and vfio_pci_core_device vdev) have not yet been
+ * freed. vdev is accessible, even if the DMABUF has been
+ * revoked or cleanup has happened, because
+ * vfio_unregister_group_dev() can't complete.
+ *
+ * 4. Take the vdev->memory_lock then vdev->dmabuf_lock:
+ * - Either the DMABUF is usable, or has been cleaned up.
+ * - It's not necessary to also take the resv lock, because
+ * the status/vdev can't change while dmabuf_lock is held.
+ * - Test the DMABUF revocation status again: if it was
+ * revoked between 1 and 4, return a SIGBUS. Otherwise,
+ * return a PFN.
+ *
+ * 5. Unlock, done.
+ */
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+
+ if (priv->status != VFIO_PCI_DMABUF_OK) {
+ pr_debug_ratelimited("%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\n",
+ __func__, vmf->address, vma->vm_pgoff);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+
+ /* If the buffer isn't revoked, vdev is valid */
+ vdev = priv->vdev;
+
+ if (!vfio_device_try_get_registration(&vdev->vdev)) {
+ /*
+ * If vdev != NULL (above), the registration should
+ * already be >0 and so this try_get should never
+ * fail.
+ */
+ dev_warn_ratelimited(&vdev->pdev->dev,
+ "%s: Unexpected registration failure\n",
+ __func__);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+ dma_resv_unlock(priv->dmabuf->resv);
+
+ /* memory_lock for vfio_pci_vmf_insert_pfn() */
+ down_read(&vdev->memory_lock);
+ /* Re-test revocation status under dmabuf_lock */
+ down_read(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_OK) {
+ int pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,
+ vmf->address,
+ order, &pfn);
+
+ if (pres == 0)
+ ret = vfio_pci_vmf_insert_pfn(vdev, vmf,
+ pfn, order);
+ else if (pres == -ERANGE)
+ ret = VM_FAULT_FALLBACK;
}
+ up_read(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
dev_dbg_ratelimited(&vdev->pdev->dev,
- "%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\n",
- __func__, order,
- vma->vm_pgoff >>
- (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),
- pgoff, (unsigned int)ret);
+ "%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\n",
+ __func__, order, pfn, vmf->address,
+ vma->vm_pgoff, (unsigned int)ret);
+ vfio_device_put_registration(&vdev->vdev);
return ret;
}
@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {
#endif
};
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)
+{
+ vma->vm_ops = &vfio_pci_mmap_ops;
+}
+
int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)
{
struct vfio_pci_core_device *vdev =
@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
unsigned int index;
u64 phys_len, req_len, pgoff, req_start;
void __iomem *bar_io;
+ int ret;
index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
if (IS_ERR(bar_io))
return PTR_ERR(bar_io);
- vma->vm_private_data = vdev;
+ ret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,
+ pci_resource_start(pdev, index),
+ req_len, index);
+ if (ret)
+ return ret;
+
vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
return ret;
INIT_LIST_HEAD(&vdev->dmabufs);
init_rwsem(&vdev->memory_lock);
+ init_rwsem(&vdev->dmabuf_lock);
xa_init(&vdev->ctx);
+ /*
+ * If a driver overrides .mmap, it has to be assumed that it
+ * might not use the DMABUF-backed core mmap; this flag
+ * enables a zap at revoke time. A driver can opt out by
+ * clearing this flag at init, if their .mmap override calls
+ * down to vfio_pci_core_mmap().
+ */
+ if (vdev->vdev.ops->mmap != vfio_pci_core_mmap)
+ vdev->zap_bars_on_revoke = true;
+
return 0;
}
EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);
@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
}
/*
- * Take the memory write lock for each device and zap BAR
- * mappings to prevent the user accessing the device while in
- * reset. Locking multiple devices is prone to deadlock,
+ * Take the memory write lock for each device and
+ * zap/revoke BAR mappings to prevent the user (or
+ * peers) accessing the device while in reset.
+ * Locking multiple devices is prone to deadlock,
* runaway and unwind if we hit contention.
*/
if (!down_write_trylock(&vdev->memory_lock)) {
@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
break;
}
- vfio_pci_dma_buf_move(vdev, true);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
}
if (!list_entry_is_head(vdev,
@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
list_for_each_entry_from_reverse(vdev, &dev_set->device_list,
vdev.dev_set_list) {
if (vdev->vdev.open_count && __vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d68..b57bfaefd9fae 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -3,25 +3,14 @@
*/
#include <linux/dma-buf-mapping.h>
#include <linux/pci-p2pdma.h>
+#include <linux/dma-buf.h>
#include <linux/dma-resv.h>
#include "vfio_pci_priv.h"
MODULE_IMPORT_NS("DMA_BUF");
-struct vfio_pci_dma_buf {
- struct dma_buf *dmabuf;
- struct vfio_pci_core_device *vdev;
- struct list_head dmabufs_elm;
- size_t size;
- struct phys_vec *phys_vec;
- struct p2pdma_provider *provider;
- u32 nr_ranges;
- struct kref kref;
- struct completion comp;
- u8 revoked : 1;
-};
-
+#ifdef CONFIG_VFIO_PCI_DMABUF
static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
struct dma_buf_attachment *attachment)
{
@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
if (!attachment->peer2peer)
return -EOPNOTSUPP;
- if (priv->revoked)
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
return -ENODEV;
if (!dma_buf_attach_revocable(attachment))
@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
return 0;
}
+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)
+{
+ struct vfio_pci_dma_buf *priv = dmabuf->priv;
+
+ /*
+ * dma_buf_mmap_internal() has asserted that the VMA is
+ * contained within the DMABUF size before calling this.
+ *
+ * Also, if we observe that the buffer is revoked now then
+ * refuse the mmap(). This is a belt-and-braces early failure
+ * to ease debugging a revoked buffer being used. Userspace
+ * might also race an mmap() against an explicit revocation,
+ * or an action causing a revoke; race scenarios are still
+ * safe because the fault handler ultimately prevents access
+ * to a revoked buffer if it isn't caught here.
+ */
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
+ return -ENODEV;
+ /*
+ * Make clear that anything with an offset adjustment is
+ * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()
+ * maths would underflow; this doesn't happen through the
+ * regular DMABUF export path used with this mmap(). A DMABUF
+ * implicitly created for BAR mmap could have adjust > 0, but
+ * these can't currently be re-opened and mmap()ed again.
+ * Catch here in case that assumption ever changes.
+ */
+ if (priv->vma_pgoff_adjust)
+ return -EINVAL;
+ if ((vma->vm_flags & VM_SHARED) == 0)
+ return -EINVAL;
+
+ vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
+ vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
+
+ /* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */
+ vm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |
+ VM_DONTEXPAND | VM_DONTDUMP);
+ vma->vm_private_data = priv;
+ vfio_pci_set_vma_ops(vma);
+
+ return 0;
+}
+#else
+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
+ struct dma_buf_attachment *attachment)
+{
+ /*
+ * Explicit export can't occur without the DMABUF feature, but
+ * DMABUFs are implicitly created for BAR mappings. An
+ * .attach that fails prevents dma_buf_attach().
+ */
+ return -EOPNOTSUPP;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
static void vfio_pci_dma_buf_done(struct kref *kref)
{
struct vfio_pci_dma_buf *priv =
@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,
dma_resv_assert_held(priv->dmabuf->resv);
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return ERR_PTR(-ENODEV);
ret = dma_buf_phys_vec_to_sgt(attachment, priv->provider,
@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)
* The refcount prevents both.
*/
if (priv->vdev) {
- down_write(&priv->vdev->memory_lock);
+ down_write(&priv->vdev->dmabuf_lock);
list_del_init(&priv->dmabufs_elm);
- up_write(&priv->vdev->memory_lock);
+ up_write(&priv->vdev->dmabuf_lock);
vfio_device_put_registration(&priv->vdev->vdev);
}
+ if (priv->vfile)
+ fput(priv->vfile);
kfree(priv->phys_vec);
kfree(priv);
}
static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
.attach = vfio_pci_dma_buf_attach,
+#ifdef CONFIG_VFIO_PCI_DMABUF
+ .mmap = vfio_pci_dma_buf_mmap,
+#endif
.map_dma_buf = vfio_pci_dma_buf_map,
.unmap_dma_buf = vfio_pci_dma_buf_unmap,
.release = vfio_pci_dma_buf_release,
};
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long fault_addr,
+ unsigned int order,
+ unsigned long *out_pfn)
+{
+ /*
+ * Given a VMA (start, end, pgoffs) and a fault address,
+ * search the corresponding DMABUF's phys_vec[] to find the
+ * range representing the address's offset into the VMA, and
+ * its PFN. vdev must be the device that the DMABUF priv was
+ * exported from; vdev->dmabuf_lock must be held, and priv
+ * must not be revoked.
+ *
+ * The phys_vec[] ranges represent contiguous spans of VAs
+ * upwards from the buffer offset 0; the actual PFNs might be
+ * in any order, overlap/alias, etc. Calculate an offset of
+ * the desired page given VMA start/pgoff and address, then
+ * search upwards from 0 to find which span contains it.
+ *
+ * On success, a valid PFN for a page sized by 'order' is
+ * returned into out_pfn.
+ *
+ * Failure occurs if:
+ * - A hugepage would cross the edge of the VMA,
+ * - A hugepage isn't entirely contained within a range
+ * (including where it straddles the boundary between
+ * ranges),
+ * - We find a range, but the final PFN isn't aligned to the
+ * requested order.
+ *
+ * Upon failure, -ERANGE is returned and the caller is
+ * expected to try again with a smaller order, which will
+ * eventually succeed.
+ *
+ * It's suboptimal if DMABUFs are created with neighbouring
+ * ranges that are physically contiguous, since hugepages
+ * can't straddle range boundaries. (The construction of the
+ * ranges should merge them in this case.)
+ *
+ * Finally, vma_pgoff_adjust is used with a DMABUF created for
+ * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
+ * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
+ * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
+ * To avoid double-offsetting in this scenario, subtracting
+ * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
+ * the effective offset. This also removes the VFIO region
+ * index encoded in vm_pgoff for VFIO BAR mmaps.
+ */
+
+ const unsigned long pagesize = PAGE_SIZE << order;
+ unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) <<
+ PAGE_SHIFT;
+ unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
+ unsigned long rounded_page_end = rounded_page_addr + pagesize;
+ unsigned long fault_offset;
+ unsigned long fault_offset_end;
+ unsigned long range_start_offset = 0;
+ unsigned int i;
+ int ret;
+
+ if (unlikely(!vdev))
+ return -ENODEV;
+
+ /* This prevents the dmabuf revocation state from changing under us */
+ lockdep_assert_held(&vdev->dmabuf_lock);
+
+ if (unlikely(priv->vdev != vdev || priv->status != VFIO_PCI_DMABUF_OK))
+ return -ENODEV;
+
+ if (rounded_page_addr < vma->vm_start || rounded_page_end > vma->vm_end) {
+ if (order > 0)
+ return -ERANGE;
+
+ /* A fault address outside of the VMA is absurd. */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "Fault addr 0x%lx outside VMA 0x%lx-0x%lx\n",
+ fault_addr, vma->vm_start, vma->vm_end);
+ return -EFAULT;
+ }
+
+ /*
+ * fault_offset[_end] is the span within the DMABUF
+ * corresponding to the faulting page:
+ */
+ if (unlikely(check_add_overflow(rounded_page_addr - vma->vm_start,
+ vma_off, &fault_offset) ||
+ check_add_overflow(fault_offset, pagesize,
+ &fault_offset_end)))
+ return -EFAULT;
+
+ /*
+ * Iterate over ranges in the buffer, summing their lengths:
+ * range_start_offset represents the current range's starting
+ * offset in the buffer (from 0 upwards).
+ *
+ * A failure for order == 0 is unexpected, and triggers a
+ * fault/warn.
+ */
+ ret = (order == 0) ? -EFAULT : -ERANGE;
+
+ for (i = 0; i < priv->nr_ranges; i++) {
+ size_t range_len = priv->phys_vec[i].len;
+
+ /* Early exit if range starts after the page end */
+ if (fault_offset_end <= range_start_offset)
+ break;
+
+ if (fault_offset >= range_start_offset &&
+ fault_offset_end <= range_start_offset + range_len) {
+ /*
+ * The faulting page is wholly contained
+ * within the span represented by this range,
+ * so validate PFN alignment for the order.
+ * The if() condition ensures the pfn
+ * arithmetic won't overflow.
+ */
+ unsigned long pfn =
+ ((fault_offset - range_start_offset) +
+ priv->phys_vec[i].paddr) >> PAGE_SHIFT;
+
+ if (IS_ALIGNED(pfn, 1 << order)) {
+ *out_pfn = pfn;
+ ret = 0;
+ }
+ /*
+ * Else order > 0; ERANGE retries with smaller
+ * order
+ */
+ break;
+ }
+ range_start_offset += range_len;
+ }
+
+ if (order == 0 && ret != 0)
+ /*
+ * The address fell outside of the span represented by
+ * the (concatenated) ranges. As setup of a mapping
+ * ensures that the VMA is <= the total size of the
+ * ranges this should never happen. If it does, warn
+ * and SIGBUS.
+ */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\n",
+ fault_addr, order, vma->vm_start, vma->vm_end,
+ vma->vm_pgoff, priv->nr_ranges, priv->size);
+
+ return ret;
+}
+
+/*
+ * Create a DMABUF corresponding to priv, add it to vdev->dmabufs list
+ * for tracking (meaning cleanup or revocation will zap it), and take
+ * a vfio_device registration.
+ */
+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv, u32 flags)
+{
+ DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
+
+ if (!vfio_device_try_get_registration(&vdev->vdev))
+ return -ENODEV;
+
+ exp_info.ops = &vfio_pci_dmabuf_ops;
+ exp_info.size = priv->size;
+ exp_info.flags = flags;
+ exp_info.priv = priv;
+
+ priv->dmabuf = dma_buf_export(&exp_info);
+ if (IS_ERR(priv->dmabuf)) {
+ vfio_device_put_registration(&vdev->vdev);
+ return PTR_ERR(priv->dmabuf);
+ }
+
+ kref_init(&priv->kref);
+ init_completion(&priv->comp);
+
+ /* dma_buf_put() now frees priv */
+ INIT_LIST_HEAD(&priv->dmabufs_elm);
+
+ /*
+ * dmabuf_lock synchronises access (R) or updates (W) to the
+ * vdev->dmabufs list and to bars_revoked (see below). The
+ * revocation state of DMABUF elements in the list is written
+ * holding both dmabuf_lock(W) and resv, and tested with
+ * either.
+ *
+ * (memory_lock, if held ->) dmabuf_lock -> resv
+ *
+ * NOTE: memory_lock is strictly avoided here, to avoid a
+ * dependency on memory_lock when mmap_lock is held, when
+ * mmap() leads to export. vfio-pci variant drivers are
+ * permitted to hold memory_lock across actions that might
+ * fault (such as user access); a deadlock could result when
+ * that fault path attempts to take mmap_lock (if held by an
+ * export waiting for memory_lock).
+ *
+ * vdev->bars_revoked tracks the BAR revocation status updated
+ * via vfio_pci_dma_buf_move(), so the initial DMABUF state
+ * follows the same criteria that later update the DMABUF
+ * state (BAR zap, etc.).
+ */
+ lockdep_assert_not_held(&vdev->memory_lock);
+
+ down_write(&vdev->dmabuf_lock);
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ priv->status = vdev->bars_revoked ? VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK;
+ list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
+ dma_resv_unlock(priv->dmabuf->resv);
+ up_write(&vdev->dmabuf_lock);
+
+ return 0;
+}
+
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index)
+{
+ struct vfio_pci_dma_buf *priv;
+ unsigned long vma_pgoff = vma->vm_pgoff & (VFIO_PCI_OFFSET_MASK >> PAGE_SHIFT);
+ char *bufname;
+ int ret;
+
+ priv = kzalloc_obj(*priv);
+ if (!priv)
+ return -ENOMEM;
+
+ priv->phys_vec = kzalloc_obj(*priv->phys_vec);
+ if (!priv->phys_vec) {
+ ret = -ENOMEM;
+ goto err_free_priv;
+ }
+
+ /*
+ * Debug name: The absolute maximum size of the name
+ * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.
+ */
+ bufname = kasprintf(GFP_KERNEL, "vfio:%s/%x",
+ pci_name(vdev->pdev),
+ res_index);
+
+ if (!bufname) {
+ ret = -ENOMEM;
+ goto err_free_phys;
+ }
+
+ /*
+ * The DMABUF begins from the mmap()'s BAR offset, i.e. the
+ * start of the VMA corresponds to byte 0 of the DMABUF and
+ * byte (vma_pgoff << PAGE_SHIFT) of the BAR.
+ *
+ * vfio_pci_dma_buf_find_pfn() reverses this offset using
+ * vma_pgoff_adjust, so that ultimately a fault's offset from
+ * the start of the _VMA_ has a consistent usage whether the
+ * VMA originates from an mmap() of the VFIO device here or a
+ * direct DMABUF mmap(). Note vma_pgoff_adjust also includes
+ * the encoded VFIO region index, which cancels out the index
+ * encoded in vm_pgoff.
+ */
+ priv->vdev = vdev;
+ priv->size = req_len;
+ priv->nr_ranges = 1;
+ priv->vma_pgoff_adjust = vma->vm_pgoff;
+
+ /*
+ * The provider can be NULL _iff_ the DMABUF feature isn't
+ * supported, because it's only used by DMABUF import and
+ * attach is prohibited if the feature isn't present.
+ */
+ priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
+ if (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) && !priv->provider) {
+ ret = -EINVAL;
+ goto err_free_name;
+ }
+
+ priv->phys_vec[0].paddr = phys_start + ((u64)vma_pgoff << PAGE_SHIFT);
+ priv->phys_vec[0].len = priv->size;
+
+ ret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);
+ if (ret)
+ goto err_free_name;
+
+ if (dma_buf_set_name(priv->dmabuf, bufname)) {
+ dev_dbg_ratelimited(&vdev->pdev->dev,
+ "Failed to set map name '%s'\n",
+ bufname);
+ kfree(bufname);
+ }
+
+ /*
+ * Ownership of the DMABUF file transfers to the VMA so that
+ * other users can locate the DMABUF via a VA. Ownership of
+ * the original VFIO device file being mmap()ed transfers to
+ * priv, and is put when the DMABUF is released. This
+ * intentionally does not use get_file()/vma_set_file()
+ * because the references are already held, and ownership
+ * moves.
+ */
+ priv->vfile = vma->vm_file;
+ vma->vm_file = priv->dmabuf->file;
+ vma->vm_private_data = priv;
+
+ return 0;
+
+err_free_name:
+ kfree(bufname);
+err_free_phys:
+ kfree(priv->phys_vec);
+err_free_priv:
+ kfree(priv);
+ return ret;
+}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
/*
* This is a temporary "private interconnect" between VFIO DMABUF and iommufd.
* It allows the two co-operating drivers to exchange the physical address of
@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
return -EOPNOTSUPP;
priv = attachment->dmabuf->priv;
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return -ENODEV;
/* More than one range to iommufd will require proper DMABUF support */
@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
struct vfio_device_feature_dma_buf get_dma_buf = {};
struct vfio_region_dma_range *dma_ranges;
- DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
struct vfio_pci_dma_buf *priv;
size_t length;
int ret;
@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
dma_ranges = NULL;
- if (!vfio_device_try_get_registration(&vdev->vdev)) {
- ret = -ENODEV;
+ ret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);
+ if (ret)
goto err_free_phys;
- }
-
- exp_info.ops = &vfio_pci_dmabuf_ops;
- exp_info.size = priv->size;
- exp_info.flags = get_dma_buf.open_flags;
- exp_info.priv = priv;
-
- priv->dmabuf = dma_buf_export(&exp_info);
- if (IS_ERR(priv->dmabuf)) {
- ret = PTR_ERR(priv->dmabuf);
- goto err_dev_put;
- }
-
- kref_init(&priv->kref);
- init_completion(&priv->comp);
-
- /* dma_buf_put() now frees priv */
- INIT_LIST_HEAD(&priv->dmabufs_elm);
- down_write(&vdev->memory_lock);
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = !__vfio_pci_memory_enabled(vdev);
- list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
- dma_resv_unlock(priv->dmabuf->resv);
- up_write(&vdev->memory_lock);
-
/*
* dma_buf_fd() consumes the reference, when the file closes the dmabuf
* will be released.
@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
return ret;
-err_dev_put:
- vfio_device_put_registration(&vdev->vdev);
err_free_phys:
kfree(priv->phys_vec);
err_free_priv:
@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
return ret;
}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
+/*
+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives
+ * the guarantee that all future map/attach attempts will fail no
+ * matter what, whereas REVOKED can transition back to OK.
+ */
+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,
+ enum vfio_pci_dma_buf_status new_status)
+{
+ bool was_revoked;
+
+ /*
+ * Changes to the DMABUF's revocation status are synchronised
+ * using dmabuf_lock:
+ */
+ lockdep_assert_held_write(&priv->vdev->dmabuf_lock);
+
+ /* If DEAD, state can no longer change */
+ if (priv->status == VFIO_PCI_DMABUF_DEAD ||
+ priv->status == new_status)
+ return;
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ was_revoked = (priv->status == VFIO_PCI_DMABUF_REVOKED);
+
+ if (new_status != VFIO_PCI_DMABUF_OK) {
+ priv->status = new_status;
+
+ if (was_revoked) {
+ /*
+ * A REVOKED buffer is being marked DEAD.
+ * invalidate_mappings/unmap wait happened
+ * when it became REVOKED, don't wait again.
+ */
+ dma_resv_unlock(priv->dmabuf->resv);
+ return;
+ }
+ dma_buf_invalidate_mappings(priv->dmabuf);
+ dma_resv_wait_timeout(priv->dmabuf->resv,
+ DMA_RESV_USAGE_BOOKKEEP, false,
+ MAX_SCHEDULE_TIMEOUT);
+ dma_resv_unlock(priv->dmabuf->resv);
+ kref_put(&priv->kref, vfio_pci_dma_buf_done);
+ wait_for_completion(&priv->comp);
+ unmap_mapping_range(priv->dmabuf->file->f_mapping,
+ 0, 0, true);
+ /*
+ * Re-arm the registered kref reference and the
+ * completion so the post-revoke state matches the
+ * post-creation state. An un-revoke followed by a
+ * new mapping needs the kref to be non-zero before
+ * kref_get(), and vfio_pci_dma_buf_cleanup()
+ * delegates its drain back through this revoke
+ * path on a possibly-already-revoked dma-buf.
+ */
+ kref_init(&priv->kref);
+ reinit_completion(&priv->comp);
+ } else {
+ priv->status = VFIO_PCI_DMABUF_OK;
+ dma_resv_unlock(priv->dmabuf->resv);
+ }
+}
void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
{
@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
lockdep_assert_held_write(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ vdev->bars_revoked = revoked;
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
-
- if (priv->revoked != revoked) {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- if (revoked)
- priv->revoked = true;
- dma_buf_invalidate_mappings(priv->dmabuf);
- dma_resv_wait_timeout(priv->dmabuf->resv,
- DMA_RESV_USAGE_BOOKKEEP, false,
- MAX_SCHEDULE_TIMEOUT);
- dma_resv_unlock(priv->dmabuf->resv);
- if (revoked) {
- kref_put(&priv->kref, vfio_pci_dma_buf_done);
- wait_for_completion(&priv->comp);
- /*
- * Re-arm the registered kref reference and the
- * completion so the post-revoke state matches the
- * post-creation state. An un-revoke followed by a
- * new mapping needs the kref to be non-zero before
- * kref_get(), and vfio_pci_dma_buf_cleanup()
- * delegates its drain back through this revoke
- * path on a possibly-already-revoked dma-buf.
- */
- kref_init(&priv->kref);
- reinit_completion(&priv->comp);
- } else {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = false;
- dma_resv_unlock(priv->dmabuf->resv);
- }
- }
+ vfio_pci_dma_buf_set_status(priv, revoked ?
+ VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
}
void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
*/
vfio_pci_dma_buf_move(vdev, true);
+ down_write(&vdev->dmabuf_lock);
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
list_del_init(&priv->dmabufs_elm);
- priv->vdev = NULL;
+ WRITE_ONCE(priv->vdev, NULL);
vfio_device_put_registration(&vdev->vdev);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
up_write(&vdev->memory_lock);
}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
+{
+ struct vfio_device_feature_dma_buf_revoke db_revoke;
+ struct vfio_pci_dma_buf *priv;
+ struct dma_buf *dmabuf;
+ int ret;
+
+ if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
+ return -EOPNOTSUPP;
+
+ ret = vfio_check_feature(flags, argsz,
+ VFIO_DEVICE_FEATURE_SET,
+ sizeof(db_revoke));
+ if (ret != 1)
+ return ret;
+
+ if (copy_from_user(&db_revoke, arg, sizeof(db_revoke)))
+ return -EFAULT;
+
+ dmabuf = dma_buf_get(db_revoke.dmabuf_fd);
+ if (IS_ERR(dmabuf))
+ return PTR_ERR(dmabuf);
+
+ priv = dmabuf->priv;
+ /*
+ * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_
+ * relates to the VFIO device it was provided with.
+ *
+ * If the DMABUF relates to this vdev then priv->vdev is
+ * stable because this open fd prevents cleanup.
+ *
+ * If it relates to a different vdev, reading priv->vdev might
+ * race with a concurrent cleanup on that device. But if so,
+ * it points to a non-matching vdev or NULL and is unusable
+ * either way.
+ */
+ if (dmabuf->ops != &vfio_pci_dmabuf_ops ||
+ READ_ONCE(priv->vdev) != vdev) {
+ ret = -ENODEV;
+ goto out_put_buf;
+ }
+
+ /*
+ * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()
+ * from getting it and then blocking all devices in the dev_set behind
+ * this revoke's drain.
+ */
+ down_read(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_DEAD) {
+ ret = -EBADFD;
+ } else {
+ vfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);
+ ret = 0;
+ }
+ up_write(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
+
+out_put_buf:
+ dma_buf_put(dmabuf);
+
+ return ret;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 4e7162234a2eb..ca12221af5553 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {
bool test_mem;
};
+enum vfio_pci_dma_buf_status {
+ VFIO_PCI_DMABUF_OK = 0,
+ VFIO_PCI_DMABUF_REVOKED = 1,
+ VFIO_PCI_DMABUF_DEAD = 2,
+};
+
+struct vfio_pci_dma_buf {
+ struct dma_buf *dmabuf;
+ struct vfio_pci_core_device *vdev;
+ struct list_head dmabufs_elm;
+ size_t size;
+ struct phys_vec *phys_vec;
+ struct p2pdma_provider *provider;
+ struct file *vfile;
+ u32 nr_ranges;
+ struct kref kref;
+ struct completion comp;
+ unsigned long vma_pgoff_adjust;
+ enum vfio_pci_dma_buf_status status;
+};
+
bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);
void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);
@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);
int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state);
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
u16 cmd);
@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)
return (pdev->class >> 8) == PCI_CLASS_DISPLAY_VGA;
}
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long address,
+ unsigned int order,
+ unsigned long *out_pfn);
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index);
+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);
+
#ifdef CONFIG_VFIO_PCI_DMABUF
int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
struct vfio_device_feature_dma_buf __user *arg,
size_t argsz);
-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz);
#else
static inline int
vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
return -ENOTTY;
}
-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
-{
-}
-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,
- bool revoked)
+static inline int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
{
+ return -ENOTTY;
}
#endif
diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
index d15b2b31d3c91..0f88132c06546 100644
--- a/include/linux/dma-buf.h
+++ b/include/linux/dma-buf.h
@@ -342,12 +342,14 @@ struct dma_buf {
/**
* @name:
*
- * Userspace-provided name. Default value is NULL. If not NULL,
- * length cannot be longer than DMA_BUF_NAME_LEN, including NIL
- * char. Useful for accounting and debugging. Read/Write accesses
- * are protected by @name_lock
- *
- * See the IOCTLs DMA_BUF_SET_NAME or DMA_BUF_SET_NAME_A/B
+ * Exporter or userspace-provided name. Default value is
+ * NULL. If not NULL, length cannot be longer than
+ * DMA_BUF_NAME_LEN, including NIL char. Useful for accounting
+ * and debugging. Read/Write accesses are protected by
+ * @name_lock
+ *
+ * See dma_buf_set_name(), and the IOCTLs DMA_BUF_SET_NAME or
+ * DMA_BUF_SET_NAME_A/B
*/
const char *name;
@@ -571,6 +573,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);
struct dma_buf *dma_buf_get(int fd);
void dma_buf_put(struct dma_buf *dmabuf);
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);
+
struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,
enum dma_data_direction);
void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa2..44891fdb7c76e 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -129,11 +129,13 @@ struct vfio_pci_core_device {
bool disable_idle_d3:1;
bool nointxmask:1;
bool disable_vga:1;
+ bool zap_bars_on_revoke:1;
/* Flags modified at runtime - dedicated storage unit */
bool needs_reset;
bool pm_intx_masked;
bool pm_runtime_engaged;
bool sriov_active;
+ bool bars_revoked;
struct pci_saved_state *pci_saved_state;
struct pci_saved_state *pm_save;
int ioeventfds_nr;
@@ -148,6 +150,7 @@ struct vfio_pci_core_device {
struct vfio_pci_core_device *sriov_pf_core_dev;
struct notifier_block nb;
struct rw_semaphore memory_lock;
+ struct rw_semaphore dmabuf_lock;
struct list_head dmabufs;
};
diff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h
index e41437fa17ad0..d3c6057983e09 100644
--- a/include/uapi/linux/vfio.h
+++ b/include/uapi/linux/vfio.h
@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {
#define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
+/**
+ * Given a DMABUF fd previously exported from the same device by
+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that
+ * access to the corresponding DMABUF is immediately revoked. On
+ * successful return, the buffer is no longer accessible through any
+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future
+ * mmap()s and map/attach requests from any new/existing importer.
+ *
+ * Return: 0 on success, -1 and errno is set on failure:
+ *
+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.
+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.
+ * ENODEV: The DMABUF was not exported from this device.
+ * EBADFD: The DMABUF is already revoked by this feature.
+ *
+ * Additionally, common errors can occur: EFAULT accessing the struct,
+ * or EINVAL requesting an unsupported feature op.
+ */
+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14
+
+struct vfio_device_feature_dma_buf_revoke {
+ __s32 dmabuf_fd;
+};
+
/* -------- API for Type1 VFIO IOMMU -------- */
/**
|
| 3/2 | 2026/10/06 20:14 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 24741 output: 34 thoughts: 2488 Thoughts: **VFIO/DMA-BUF Patch Analysis: Reachability Assessment** Okay, let's break this down. The patch series focuses on integrating DMA-BUF functionality within VFIO-PCI. My initial reaction is to evaluate reachability – can this code be tested in standard environments? First, I see changes to `dma_buf.c`. The `dma_buf_set_name` function is now exported and used in an ioctl handler (`DMA_BUF_SET_NAME_A`/`B`). Crucially, this means anyone with a DMA-BUF file descriptor can call `ioctl` to modify the name. That's good, `udmabuf`, `virtio-gpu`, and others. The core of the patch series implements a new `vfio_pci_dma_buf_mmap` and related functions. This leverages DMA-BUF backing for mmap operations within VFIO-PCI. This is interesting. The `vfio_pci_core_mmap_prep_dmabuf` will be a critical part to examine. The file `vfio_pci_dmabuf.c` has many implementations. This means there's a serious integration of DMA-BUF and VFIO-PCI. Now, reachability. The critical question is whether VFIO-PCI can be triggered in a QEMU/GCE environment, which I believe it can, due to generic PCI device assignment. The question is what standard emulated devices exist and which can be bound by vfio-pci. QEMU and syzkaller support is crucial. So the question is if the new functionality in vfio-pci is accessible from user space. I need to confirm whether syzkaller covers VFIO PCI or DMA-BUF. Based on the triage philosophy, I need to check whether the modified code is *structurally* reachable. Let's check the kernel configs to verify the VFIO_PCI option is enabled. We need to check the code paths exclusive to the target architecture. The `dma_buf_set_name` change is definitely accessible: anything creating a DMA-BUF and calling the relevant `ioctl` can trigger this code path. Okay, so dma-buf ioctls are reachable, and vfio-pci, in general, is reachable. The changes to `vfio_pci_mmap_huge_fault` and `vfio_pci_core_mmap` are most interesting, as they introduce `dma_buf` use. So, in summary: I think this looks good. The core DMA-BUF functions, at the lowest level, are accessible, so the VFIO-PCI DMABUF integration *should* be testable. |
| 4/2 | 2026/10/06 20:15 | tool | grepper |
0mArgs:{
"Expression": "CONFIG_VFIO_PCI",
"PathPrefixes": [
"drivers/vfio/pci"
]
}Results: {
"Output": "drivers/vfio/pci/Makefile=3=vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o\ndrivers/vfio/pci/Makefile:4:vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o\ndrivers/vfio/pci/Makefile:5:obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o\ndrivers/vfio/pci/Makefile-6-\ndrivers/vfio/pci/Makefile=7=vfio-pci-y := vfio_pci.o\ndrivers/vfio/pci/Makefile:8:vfio-pci-$(CONFIG_VFIO_PCI_IGD) += vfio_pci_igd.o\ndrivers/vfio/pci/Makefile:9:obj-$(CONFIG_VFIO_PCI) += vfio-pci.o\ndrivers/vfio/pci/Makefile-10-\n--\ndrivers/vfio/pci/vfio_pci.c=39=MODULE_PARM_DESC(nointxmask,\n--\ndrivers/vfio/pci/vfio_pci.c-41-\ndrivers/vfio/pci/vfio_pci.c:42:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci.c-43-static bool disable_vga;\n--\ndrivers/vfio/pci/vfio_pci.c=128=static int vfio_pci_init_dev(struct vfio_device *core_vdev)\n--\ndrivers/vfio/pci/vfio_pci.c-141-\tvdev-\u003edisable_idle_d3 = disable_idle_d3;\ndrivers/vfio/pci/vfio_pci.c:142:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci.c-143-\tvdev-\u003edisable_vga = disable_vga;\n--\ndrivers/vfio/pci/vfio_pci_config.c=1744=int vfio_config_init(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_config.c-1826-\ndrivers/vfio/pci/vfio_pci_config.c:1827:\tif (!IS_ENABLED(CONFIG_VFIO_PCI_INTX) || vdev-\u003enointx ||\ndrivers/vfio/pci/vfio_pci_config.c-1828-\t !vdev-\u003epdev-\u003eirq || vdev-\u003epdev-\u003eirq == IRQ_NOTCONNECTED)\n--\ndrivers/vfio/pci/vfio_pci_core.c=96=static inline bool vfio_vga_disabled(struct vfio_pci_core_device *vdev)\ndrivers/vfio/pci/vfio_pci_core.c-97-{\ndrivers/vfio/pci/vfio_pci_core.c:98:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_core.c-99-\treturn vdev-\u003edisable_vga;\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=11=MODULE_IMPORT_NS(\"DMA_BUF\");\ndrivers/vfio/pci/vfio_pci_dmabuf.c-12-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:13:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-14-static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=75=static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-84-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:85:#endif /* CONFIG_VFIO_PCI_DMABUF */\ndrivers/vfio/pci/vfio_pci_dmabuf.c-86-\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=149=static const struct dma_buf_ops vfio_pci_dmabuf_ops = {\ndrivers/vfio/pci/vfio_pci_dmabuf.c-150-\t.attach = vfio_pci_dma_buf_attach,\ndrivers/vfio/pci/vfio_pci_dmabuf.c:151:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-152-\t.mmap = vfio_pci_dma_buf_mmap,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=377=int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-433-\tpriv-\u003eprovider = pcim_p2pdma_provider(vdev-\u003epdev, res_index);\ndrivers/vfio/pci/vfio_pci_dmabuf.c:434:\tif (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) \u0026\u0026 !priv-\u003eprovider) {\ndrivers/vfio/pci/vfio_pci_dmabuf.c-435-\t\tret = -EINVAL;\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-476-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:477:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-478-/*\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=590=int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-675-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:676:#endif /* CONFIG_VFIO_PCI_DMABUF */\ndrivers/vfio/pci/vfio_pci_dmabuf.c-677-\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=760=void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-789-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:790:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_dmabuf.c-791-int vfio_pci_core_feature_dma_buf_revoke(\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-856-}\ndrivers/vfio/pci/vfio_pci_dmabuf.c:857:#endif /* CONFIG_VFIO_PCI_DMABUF */\n--\ndrivers/vfio/pci/vfio_pci_priv.h=65=ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_priv.h-67-\ndrivers/vfio/pci/vfio_pci_priv.h:68:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_priv.h-69-ssize_t vfio_pci_vga_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_priv.h=95=void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_priv.h-97-\ndrivers/vfio/pci/vfio_pci_priv.h:98:#ifdef CONFIG_VFIO_PCI_IGD\ndrivers/vfio/pci/vfio_pci_priv.h-99-bool vfio_pci_is_intel_display(struct pci_dev *pdev);\n--\ndrivers/vfio/pci/vfio_pci_priv.h=107=static inline int vfio_pci_igd_init(struct vfio_pci_core_device *vdev)\n--\ndrivers/vfio/pci/vfio_pci_priv.h-112-\ndrivers/vfio/pci/vfio_pci_priv.h:113:#ifdef CONFIG_VFIO_PCI_ZDEV_KVM\ndrivers/vfio/pci/vfio_pci_priv.h-114-int vfio_pci_info_zdev_add_caps(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_priv.h=160=void vfio_pci_set_vma_ops(struct vm_area_struct *vma);\ndrivers/vfio/pci/vfio_pci_priv.h-161-\ndrivers/vfio/pci/vfio_pci_priv.h:162:#ifdef CONFIG_VFIO_PCI_DMABUF\ndrivers/vfio/pci/vfio_pci_priv.h-163-int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_rdwr.c=201=ssize_t vfio_pci_bar_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n--\ndrivers/vfio/pci/vfio_pci_rdwr.c-280-\ndrivers/vfio/pci/vfio_pci_rdwr.c:281:#ifdef CONFIG_VFIO_PCI_VGA\ndrivers/vfio/pci/vfio_pci_rdwr.c-282-ssize_t vfio_pci_vga_rw(struct vfio_pci_core_device *vdev, char __user *buf,\n"
} |
| 5/2 | 2026/10/06 20:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 30654 output: 36 thoughts: 16 Thoughts: Okay, here's my summary, taking the role of the expert: **Diving into the VFIO PCI Configuration** Alright, let's get down to brass tacks. I need to understand something related to the VFIO PCI driver, so my first instinct, as always, is to navigate directly to the heart of the matter: `drivers/vfio/pci/Kconfig`. That's where all the configuration options for this critical component are defined. I need to understand what's enabled, disabled, how things are structured. No point in wandering around aimlessly. The Kconfig file is the definitive source of truth for all of this, and from there, I can get a clear picture of what's possible, what's supported, and where to look next if I need more detail. This is the starting point for any real understanding of the VFIO PCI configuration space. It's the key. |
| 6/2 | 2026/10/06 20:15 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/Kconfig",
"FirstLine": 1,
"LineCount": 70
}Results: {
"Contents": " 1:\t# SPDX-License-Identifier: GPL-2.0-only\n 2:\tmenu \"VFIO support for PCI devices\"\n 3:\t\tdepends on PCI\n 4:\t\n 5:\tconfig VFIO_PCI_CORE\n 6:\t\ttristate\n 7:\t\tselect VFIO_VIRQFD\n 8:\t\tselect IRQ_BYPASS_MANAGER\n 9:\t\tselect DMA_SHARED_BUFFER\n 10:\t\n 11:\tconfig VFIO_PCI_INTX\n 12:\t\tdef_bool y if !S390\n 13:\t\tdepends on VFIO_PCI_CORE\n 14:\t\n 15:\tconfig VFIO_PCI\n 16:\t\ttristate \"Generic VFIO support for any PCI device\"\n 17:\t\tselect VFIO_PCI_CORE\n 18:\t\thelp\n 19:\t\t Support for the generic PCI VFIO bus driver which can connect any\n 20:\t\t PCI device to the VFIO framework.\n 21:\t\n 22:\t\t If you don't know what to do here, say N.\n 23:\t\n 24:\tif VFIO_PCI\n 25:\tconfig VFIO_PCI_VGA\n 26:\t\tbool \"Generic VFIO PCI support for VGA devices\"\n 27:\t\tdepends on X86 \u0026\u0026 VGA_ARB\n 28:\t\thelp\n 29:\t\t Support for VGA extension to VFIO PCI. This exposes an additional\n 30:\t\t region on VGA devices for accessing legacy VGA addresses used by\n 31:\t\t BIOS and generic video drivers.\n 32:\t\n 33:\t\t If you don't know what to do here, say N.\n 34:\t\n 35:\tconfig VFIO_PCI_IGD\n 36:\t\tbool \"Generic VFIO PCI extensions for Intel graphics (GVT-d)\"\n 37:\t\tdepends on X86\n 38:\t\tdefault y\n 39:\t\thelp\n 40:\t\t Support for Intel IGD specific extensions to enable direct\n 41:\t\t assignment to virtual machines. This includes exposing an IGD\n 42:\t\t specific firmware table and read-only copies of the host bridge\n 43:\t\t and LPC bridge config space.\n 44:\t\n 45:\t\t To enable Intel IGD assignment through vfio-pci, say Y.\n 46:\tendif\n 47:\t\n 48:\tconfig VFIO_PCI_ZDEV_KVM\n 49:\t\tbool \"VFIO PCI extensions for s390x KVM passthrough\"\n 50:\t\tdepends on S390 \u0026\u0026 KVM\n 51:\t\tdefault y\n 52:\t\thelp\n 53:\t\t Support s390x-specific extensions to enable support for enhancements\n 54:\t\t to KVM passthrough capabilities, such as interpretive execution of\n 55:\t\t zPCI instructions.\n 56:\t\n 57:\t\t To enable s390x KVM vfio-pci extensions, say Y.\n 58:\t\n 59:\tconfig VFIO_PCI_DMABUF\n 60:\t\tdef_bool y if PCI_P2PDMA\n 61:\t\tdepends on VFIO_PCI_CORE\n 62:\t\n 63:\tsource \"drivers/vfio/pci/mlx5/Kconfig\"\n 64:\t\n 65:\tsource \"drivers/vfio/pci/ism/Kconfig\"\n 66:\t\n 67:\tsource \"drivers/vfio/pci/hisilicon/Kconfig\"\n 68:\t\n 69:\tsource \"drivers/vfio/pci/pds/Kconfig\"\n 70:\t\n"
} |
| 7/2 | 2026/10/06 20:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31829 output: 31 thoughts: 83 Thoughts: Okay, let's see what I can recall about VFIO_PCI and its relation to my work. **VFIO, QEMU, and Fuzzing - A Quick Re-cap** Right, so VFIO_PCI. That's the generic implementation, isn't it? It offers VFIO support for pretty much any PCI device, which is handy. Now, I'm thinking... does this integrate with QEMU? I should know this, but a quick refresh is always good. The possibilities it opens up for virtualization are significant, especially if I can leverage it. And this leads me to my next thought: does syzkaller target VFIO at all? I need to remember whether I've seen any work being done on fuzzing VFIO directly within the syzkaller framework. That would be an incredibly powerful avenue for vulnerability research. I need to quickly understand VFIO's core functionality, so I should probably jump into the code. I'll check `drivers/vfio/vfio_main.c` or a similar core VFIO file, just to get a refresher on what's going on under the hood and how it's implemented. That'll help me assess if there's any potential to integrate and fuzz that. |
| 8/2 | 2026/10/06 20:15 | tool | grepper |
0mError:no matches Args: {
"Expression": "class_create.*vfio",
"PathPrefixes": [
"drivers/vfio"
]
}Results: {
"error": "no matches"
} |
| 9/2 | 2026/10/06 20:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32225 output: 28 thoughts: 18 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience, presented in the first person: **Navigating the VFIO Landscape** Alright, so I'm thinking about the VFIO infrastructure today. Need to get a handle on how this all hangs together, so I should probably start digging around the drivers. My immediate focus? I need to get into the nitty-gritty of how the "miscdevice" or "devfs" aspects play out within the VFIO codebase. My instincts tell me there's something critical there related to device management. I need to nail down the precise integration point and how we're handling the creation and interaction with the device nodes. Specifically, I'm aiming for the drivers/vfio directory. That's the logical place to start this deep dive. Let's see what gems we can unearth. |
| 10/2 | 2026/10/06 20:15 | tool | grepper |
0mArgs:{
"Expression": "misc_register",
"PathPrefixes": [
"drivers/vfio"
]
}Results: {
"Output": "drivers/vfio/container.c=573=int __init vfio_container_init(void)\n--\ndrivers/vfio/container.c-579-\ndrivers/vfio/container.c:580:\tret = misc_register(\u0026vfio_dev);\ndrivers/vfio/container.c-581-\tif (ret) {\n"
} |
| 11/2 | 2026/10/06 20:15 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 32553 output: 32 thoughts: 2455 Thoughts: **A Deep Dive into a Potential Fuzzing Target** Okay, here's what's running through my head. I'm looking at this patch and immediately, my spidey senses are tingling. I see `/dev/vfio/vfio` created by `vfio_container_init`, and I know that interacts with PCI devices in QEMU. Syzkaller tests often touch `/dev/vfio/*` or sysfs to bind devices, so that's definitely a point of interest. The question in my mind, and I need to be clear about this, is whether or not `vfio-pci` is *structurally* unreachable hardware. The text references Section 2, which mentions vendor-specific PCIe switches, SmartNICs, or GPU drivers. I know, and this is important to get right, that `vfio-pci` is *not* a vendor-specific device driver requiring some physical ASIC. It's the core Linux kernel driver for PCI device passthrough! It can bind to virtually any PCI device – emulated ones like the e1000 or virtio-pci in QEMU, or even those test devices like edu or ivshmem. This is significant. Now, my attention snaps to `drivers/dma-buf/dma-buf.c`. This patch modifies core DMA-BUF code, specifically `dma_buf_set_name` and `dma_buf_ioctl`. `drivers/dma-buf/` is part of the core kernel, the driver framework for DMA buffer sharing. The critical question here is whether or not this code is *reachable*. And yes, it is. Anyone opening a dmabuf (through udmabuf, drm, virtio-gpu, etc.) can execute `ioctl(fd, DMA_BUF_SET_NAME_A, ...)` and `ioctl(fd, DMA_BUF_SET_NAME_B, ...)`! Let's dissect `dma_buf_ioctl`. I see the `DMA_BUF_SET_NAME_A` and `DMA_BUF_SET_NAME_B` cases, where userspace-provided data is copied into a buffer using `strndup_user`. `strndup_user` is then used to create a buffer. Then `dma_buf_set_name` is called, and on error it `kfree`s the buffer. That's a good start to analyze. Now, if I understand correctly, if `strlen(name)` in user space is >= `DMA_BUF_NAME_LEN` (let's say 32), then there's an `E2BIG` error, which, then `kfree`s the name! Also, `strndup_user(..., DMA_BUF_NAME_LEN)` will copy *at most* that many characters *including* the null terminator. I must know what `strndup_user` does in various scenarios. Ah, and I see now that `dma_buf_set_name` is also called from `vfio_pci_core_mmap_prep_dmabuf`, where a buffer name is created using `kasprintf`, with the format `vfio:%s/%x`. This is crucial because `vfio_pci_core_mmap_prep_dmabuf` is called every time a VFIO device BAR is mmap'ed through `vfio_pci_core_mmap`. And wait! This is a *major* change! Before this patch, `vfio_pci_core_mmap` did simple stuff with `vma->vm_private_data` and PFN insertion. But now, every BAR mmap in VFIO PCI is backed by a DMABUF (`vfio_pci_core_mmap_prep_dmabuf`). This exports a DMABUF, sets file on `vma->vm_file`, sets page fault handlers through DMABUF, handles revocation, hot reset, memory disable/enable in PCI config space. And there's a new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` feature ioctl! Locking is reworked with `dmabuf_lock` (a rw_semaphore), and the locking interactions between `resv`, `memory_lock`, and `dmabuf_lock`! The comment in `vfio_pci_mmap_huge_fault` further reinforces this significant reworking of memory management, fault handling, locking, revocation, and lifetime management in `vfio_pci` and DMABUF! This is where the magic happens! Now, to the core question: Is this `WorthFuzzing`? Let's check the criteria. It hits: * Core Sub |
| 12/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "DMA_SHARED_BUFFER",
"PathPrefixes": [
"drivers/dma-buf"
]
}Results: {
"Output": "drivers/dma-buf/Kconfig=4=config SYNC_FILE\n--\ndrivers/dma-buf/Kconfig-6-\tdefault n\ndrivers/dma-buf/Kconfig:7:\tselect DMA_SHARED_BUFFER\ndrivers/dma-buf/Kconfig-8-\thelp\n--\ndrivers/dma-buf/Kconfig=33=config UDMABUF\n--\ndrivers/dma-buf/Kconfig-35-\tdefault n\ndrivers/dma-buf/Kconfig:36:\tdepends on DMA_SHARED_BUFFER\ndrivers/dma-buf/Kconfig-37-\tdepends on MEMFD_CREATE || COMPILE_TEST\n--\ndrivers/dma-buf/Kconfig=43=config DMABUF_DEBUG\ndrivers/dma-buf/Kconfig-44-\tbool \"DMA-BUF debug checks\"\ndrivers/dma-buf/Kconfig:45:\tdepends on DMA_SHARED_BUFFER\ndrivers/dma-buf/Kconfig-46-\tdefault y if DEBUG\n--\ndrivers/dma-buf/Kconfig=52=config DMABUF_KUNIT_TEST\n--\ndrivers/dma-buf/Kconfig-54-\tdepends on KUNIT\ndrivers/dma-buf/Kconfig:55:\tselect DMA_SHARED_BUFFER\ndrivers/dma-buf/Kconfig-56-\tdefault KUNIT_ALL_TESTS\n--\ndrivers/dma-buf/Kconfig=60=menuconfig DMABUF_HEAPS\ndrivers/dma-buf/Kconfig-61-\tbool \"DMA-BUF Userland Memory Heaps\"\ndrivers/dma-buf/Kconfig:62:\tselect DMA_SHARED_BUFFER\ndrivers/dma-buf/Kconfig-63-\thelp\n"
} |
| 13/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36499 output: 33 thoughts: 15 Thoughts: Here's my take on those thoughts: **Checking the DMA-BUF Makefile** Okay, so we need to see how DMA-BUF is being handled at the makefile level. My thinking is to dive into `drivers/dma-buf/Makefile`. This will give me a clear picture of how the DMA-BUF framework is built, linked, and generally managed within the kernel's build system. Specifically, I'm interested in how it interacts with other drivers and subsystems. I need to understand the dependencies and compilation flags to make sure I'm not missing anything. I'll be looking for potential conflicts, build-time configurations, and where specific DMA-BUF components are defined. This is a crucial step to ensuring any changes or additions I make are properly integrated and don't introduce unexpected build errors or runtime issues. Knowing how the build system handles it all will be key to moving forward effectively. |
| 14/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "dma-buf\\.o",
"PathPrefixes": [
"drivers/dma-buf"
]
}Results: {
"Output": "drivers/dma-buf/Makefile-1-# SPDX-License-Identifier: GPL-2.0-only\ndrivers/dma-buf/Makefile:2:obj-y := dma-buf.o dma-fence.o dma-fence-array.o dma-fence-chain.o \\\ndrivers/dma-buf/Makefile-3-\t dma-fence-unwrap.o dma-resv.o dma-buf-mapping.o\n"
} |
| 15/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36854 output: 28 thoughts: 40 Thoughts: Here's my thought process, as if I were considering this myself: **Investigating the Build Configuration of dma-buf** Okay, so `dma-buf.o` is marked as `obj-y`. That typically means it's built directly into the kernel image. But hold on a second... is the whole `drivers/dma-buf` directory, and therefore everything within it, conditionally compiled? We need to verify that. Let's dive into the `drivers/Makefile` to see if there are any conditional assignments, `Kconfig` references, or other build logic that might influence whether that specific driver is enabled or disabled. That Makefile should provide the definitive answer. |
| 16/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "dma-buf",
"PathPrefixes": [
"drivers/Makefile"
]
}Results: {
"Output": "drivers/Makefile=78=obj-y\t\t\t\t+= dax/\ndrivers/Makefile:79:obj-$(CONFIG_DMA_SHARED_BUFFER) += dma-buf/\ndrivers/Makefile-80-obj-$(CONFIG_NUBUS)\t\t+= nubus/\n"
} |
| 17/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37143 output: 30 thoughts: 50 Thoughts: Okay, here's my interpretation of those thoughts, tailored for an expert audience, presented in the first person: **Kernel Configuration Dependencies: A Quick Investigation** Alright, so `CONFIG_DMA_SHARED_BUFFER` is the key to getting the `drivers/dma-buf/` subsystem built, that much is clear. My immediate next step is to drill down into the related configuration options. I'm going to quickly check the Kconfig files for Virtual Function I/O (VFIO), specifically `drivers/vfio/Kconfig` and `drivers/vfio/pci/Kconfig`. I'm expecting some dependency, or perhaps a direct inclusion of `dma-buf` features if VFIO is leveraging shared memory for device access or data transfer, as the names suggest. I need to understand if `DMA_SHARED_BUFFER` is a prerequisite for VFIO functionality, or if VFIO has its own nuances in how it handles DMA buffers. It's time to quickly trace the config dependencies to see how these subsystems interact. This also includes understanding if there is any shared implementation. Let's see if this requires an explicit enablement, or if there is a more subtle connection here. |
| 18/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "VFIO",
"PathPrefixes": [
"drivers/vfio/Kconfig"
]
}Results: {
"Output": "drivers/vfio/Kconfig-1-# SPDX-License-Identifier: GPL-2.0-only\ndrivers/vfio/Kconfig:2:menuconfig VFIO\ndrivers/vfio/Kconfig:3:\ttristate \"VFIO Non-Privileged userspace driver framework\"\ndrivers/vfio/Kconfig-4-\tselect IOMMU_API\n--\ndrivers/vfio/Kconfig-6-\tselect INTERVAL_TREE\ndrivers/vfio/Kconfig:7:\tselect VFIO_GROUP if SPAPR_TCE_IOMMU || IOMMUFD=n\ndrivers/vfio/Kconfig:8:\tselect VFIO_DEVICE_CDEV if !VFIO_GROUP\ndrivers/vfio/Kconfig:9:\tselect VFIO_CONTAINER if IOMMUFD=n\ndrivers/vfio/Kconfig-10-\thelp\ndrivers/vfio/Kconfig:11:\t VFIO provides a framework for secure userspace device drivers.\ndrivers/vfio/Kconfig-12-\t See Documentation/driver-api/vfio.rst for more details.\n--\ndrivers/vfio/Kconfig-15-\ndrivers/vfio/Kconfig:16:if VFIO\ndrivers/vfio/Kconfig:17:config VFIO_DEVICE_CDEV\ndrivers/vfio/Kconfig:18:\tbool \"Support for the VFIO cdev /dev/vfio/devices/vfioX\"\ndrivers/vfio/Kconfig-19-\tdepends on IOMMUFD \u0026\u0026 !SPAPR_TCE_IOMMU\ndrivers/vfio/Kconfig:20:\tdefault !VFIO_GROUP\ndrivers/vfio/Kconfig-21-\thelp\ndrivers/vfio/Kconfig:22:\t The VFIO device cdev is another way for userspace to get device\ndrivers/vfio/Kconfig-23-\t access. Userspace gets device fd by opening device cdev under\n--\ndrivers/vfio/Kconfig-28-\ndrivers/vfio/Kconfig:29:config VFIO_GROUP\ndrivers/vfio/Kconfig:30:\tbool \"Support for the VFIO group /dev/vfio/$group_id\"\ndrivers/vfio/Kconfig-31-\tdefault y\ndrivers/vfio/Kconfig-32-\thelp\ndrivers/vfio/Kconfig:33:\t VFIO group support provides the traditional model for accessing\ndrivers/vfio/Kconfig:34:\t devices through VFIO and is used by the majority of userspace\ndrivers/vfio/Kconfig:35:\t applications and drivers making use of VFIO.\ndrivers/vfio/Kconfig-36-\n--\ndrivers/vfio/Kconfig-38-\ndrivers/vfio/Kconfig:39:config VFIO_CONTAINER\ndrivers/vfio/Kconfig:40:\tbool \"Support for the VFIO container /dev/vfio/vfio\"\ndrivers/vfio/Kconfig:41:\tselect VFIO_IOMMU_TYPE1 if MMU \u0026\u0026 (X86 || S390 || ARM || ARM64)\ndrivers/vfio/Kconfig:42:\tdepends on VFIO_GROUP\ndrivers/vfio/Kconfig-43-\tdefault y\ndrivers/vfio/Kconfig-44-\thelp\ndrivers/vfio/Kconfig:45:\t The VFIO container is the classic interface to VFIO for establishing\ndrivers/vfio/Kconfig-46-\t IOMMU mappings. If N is selected here then IOMMUFD must be used to\n--\ndrivers/vfio/Kconfig-50-\ndrivers/vfio/Kconfig:51:if VFIO_CONTAINER\ndrivers/vfio/Kconfig:52:config VFIO_IOMMU_TYPE1\ndrivers/vfio/Kconfig-53-\ttristate\n--\ndrivers/vfio/Kconfig-55-\ndrivers/vfio/Kconfig:56:config VFIO_IOMMU_SPAPR_TCE\ndrivers/vfio/Kconfig-57-\ttristate\ndrivers/vfio/Kconfig-58-\tdepends on SPAPR_TCE_IOMMU\ndrivers/vfio/Kconfig:59:\tdefault VFIO\ndrivers/vfio/Kconfig-60-endif\ndrivers/vfio/Kconfig-61-\ndrivers/vfio/Kconfig:62:config VFIO_NOIOMMU\ndrivers/vfio/Kconfig:63:\tbool \"VFIO No-IOMMU support\"\ndrivers/vfio/Kconfig:64:\tdepends on VFIO_GROUP || (VFIO_DEVICE_CDEV \u0026\u0026 !GENERIC_ATOMIC64)\ndrivers/vfio/Kconfig:65:\tdepends on !VFIO_GROUP || VFIO_CONTAINER || IOMMUFD_VFIO_CONTAINER\ndrivers/vfio/Kconfig:66:\tselect IOMMUFD_NOIOMMU if VFIO_DEVICE_CDEV \u0026\u0026 !GENERIC_ATOMIC64\ndrivers/vfio/Kconfig-67-\thelp\ndrivers/vfio/Kconfig:68:\t VFIO is built on the ability to isolate devices using the IOMMU.\ndrivers/vfio/Kconfig-69-\t Only with an IOMMU can userspace access to DMA capable devices be\ndrivers/vfio/Kconfig:70:\t considered secure. VFIO No-IOMMU mode enables IOMMU groups for\ndrivers/vfio/Kconfig:71:\t devices without IOMMU backing for the purpose of re-using the VFIO\ndrivers/vfio/Kconfig-72-\t infrastructure in a non-secure mode. Use of this mode will result\n--\ndrivers/vfio/Kconfig-78-\ndrivers/vfio/Kconfig:79:config VFIO_VIRQFD\ndrivers/vfio/Kconfig-80-\tbool\n--\ndrivers/vfio/Kconfig-83-\ndrivers/vfio/Kconfig:84:config VFIO_DEBUGFS\ndrivers/vfio/Kconfig:85:\tbool \"Export VFIO internals in DebugFS\"\ndrivers/vfio/Kconfig-86-\tdepends on DEBUG_FS\ndrivers/vfio/Kconfig-87-\thelp\ndrivers/vfio/Kconfig:88:\t Allows exposure of VFIO device internals. This option enables\ndrivers/vfio/Kconfig:89:\t the use of debugfs by VFIO drivers as required. The device can\ndrivers/vfio/Kconfig:90:\t cause the VFIO code create a top-level debug/vfio directory\ndrivers/vfio/Kconfig-91-\t during initialization, and then populate a subdirectory with\n"
} |
| 19/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 39086 output: 32 thoughts: 563 Thoughts: **VFIO_PCI_DMABUF Deep Dive**
Okay, let's take a look at the `VFIO_PCI_DMABUF` configuration and how it plays with `PCI_P2PDMA`. I know this area, and I want to refresh my memory on the specific dependencies and compile-time behavior.
First, I see in `drivers/vfio/pci/Kconfig` that `VFIO_PCI_DMABUF` is defined as:
```kconfig
config VFIO_PCI_DMABUF
def_bool y if PCI_P2PDMA
depends on VFIO_PCI_CORE
```
So it's straightforward. `VFIO_PCI_DMABUF` defaults to 'y' if `PCI_P2PDMA` is enabled and also depends on `VFIO_PCI_CORE`. Now, in the `drivers/vfio/pci/Makefile`, I see:
```makefile
vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o
```
This tells me that if `VFIO_PCI` is built into the kernel (or as a module), then `VFIO_PCI_CORE` is selected, and therefore, `vfio_pci_dmabuf.o` is going to be compiled. Which means `VFIO_PCI_DMABUF` will be enabled because it is a direct result of `VFIO_PCI`.
Digging deeper into `drivers/vfio/pci/vfio_pci_dmabuf.c`, the `#ifdef CONFIG_VFIO_PCI_DMABUF` guards several functions: `vfio_pci_dma_buf_mmap`, `vfio_pci_dma_buf_iommufd_map`, `vfio_pci_core_feature_dma_buf`, and `vfio_pci_core_feature_dma_buf_revoke`. That makes sense. These functions are only included if the DMA-BUF functionality is enabled. And, further down, in `vfio_pci_core_mmap_prep_dmabuf`, I see that this code is looking for a provider for a `priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);` and there is a check in `vfio_pci_core_mmap_prep_dmabuf`:
```c
priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
if (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) && !priv->provider) {
ret = -EINVAL;
goto err_free_name;
}
```
So, it's checking that, if `VFIO_PCI_DMABUF` is enabled, the provider isn't null. I need to refresh my memory on `PCI_P2PDMA`. Let me check `drivers/pci/Kconfig` for the definition of `PCI_P2PDMA`.
|
| 20/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "config PCI_P2PDMA",
"PathPrefixes": [
"drivers/pci"
]
}Results: {
"Output": "drivers/pci/Kconfig=197=config PCIE_TPH\n--\ndrivers/pci/Kconfig-205-\ndrivers/pci/Kconfig:206:config PCI_P2PDMA\ndrivers/pci/Kconfig-207-\tbool \"PCI peer-to-peer transfer support\"\n"
} |
| 21/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40490 output: 35 thoughts: 20 Thoughts: Okay, let's dive into this. I'm going to take a look at the `drivers/pci/Kconfig` file. Specifically, I need to focus around line 206. I'm assuming that's where the interesting stuff is, or at least a relevant section for whatever I'm currently working on. I'll read that section thoroughly to understand the context and implications. I need to glean some information, likely related to configuration options within the PCI subsystem, but the exact focus will reveal itself as I read the surrounding code. My expertise in this area allows me to quickly identify relevant keywords, dependencies, and potential issues within this configuration file. |
| 22/2 | 2026/10/06 20:16 | tool | read-file |
0mArgs:{
"File": "drivers/pci/Kconfig",
"FirstLine": 205,
"LineCount": 20
}Results: {
"Contents": " 205:\t\n 206:\tconfig PCI_P2PDMA\n 207:\t\tbool \"PCI peer-to-peer transfer support\"\n 208:\t\tdepends on ZONE_DEVICE\n 209:\t\t#\n 210:\t\t# The need for the scatterlist DMA bus address flag means PCI P2PDMA\n 211:\t\t# requires 64bit\n 212:\t\t#\n 213:\t\tdepends on 64BIT\n 214:\t\tselect GENERIC_ALLOCATOR\n 215:\t\tselect NEED_SG_DMA_FLAGS\n 216:\t\thelp\n 217:\t\t Enables drivers to do PCI peer-to-peer transactions to and from\n 218:\t\t BARs that are exposed in other devices that are the part of\n 219:\t\t the hierarchy where peer-to-peer DMA is guaranteed by the PCI\n 220:\t\t specification to work (ie. anything below a single PCI bridge).\n 221:\t\n 222:\t\t Many PCIe root complexes do not support P2P transactions and\n 223:\t\t it's hard to tell which support it at all, so at this time,\n 224:\t\t P2P DMA transactions must be between devices behind the same root\n"
} |
| 23/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 41012 output: 29 thoughts: 2237 Thoughts: **DMA-BUF and VFIO PCI Deep Dive: Fuzzing Target Analysis** Okay, let's break this down systematically. I'm focusing on identifying the critical entry points and configuration dependencies for fuzzing the DMA-BUF integration within VFIO PCI, specifically with an eye toward generating effective test cases. First, I need to understand the dependencies. `PCI_P2PDMA` is key, depending on both `ZONE_DEVICE` and `64BIT`. `VFIO_PCI_DMABUF`, which is our primary target, is *dependent* on `PCI_P2PDMA` *and* `VFIO_PCI_CORE` (which itself is tied to `VFIO_PCI`, so that's a given). So, the enablement chain goes something like: `VFIO` -> `VFIO_PCI` -> `VFIO_PCI_CORE` -> `VFIO_PCI_DMABUF` + `PCI_P2PDMA`. `VFIO_PCI` also selects `DMA_SHARED_BUFFER`, which is also selected by `UDMABUF`. Now, the crucial part: identifying the *relevant* symbols. I'm avoiding generic allocation/deallocation or core synchronization primitives. I need the feature-specific logic. Let's list those: 1. `dma_buf_set_name`: This looks important as a direct entrypoint, it's modified and called during `dma_buf_ioctl` and `vfio_pci_core_mmap_prep_dmabuf`, so it's a good place to start. 2. `vfio_pci_core_mmap_prep_dmabuf`: A new exported function called by `vfio_pci_core_mmap`, definitely relevant. 3. `vfio_pci_dma_buf_find_pfn`: Core logic for address translation through the dmabuf phys_vec, this is likely to be a critical path. 4. `vfio_pci_core_feature_dma_buf_revoke`: The new feature ioctl entrypoint, part of a new interface that looks ripe for fuzzing. 5. `vfio_pci_dma_buf_move`: Handles dmabuf revocation/unrevocation. 6. `vfio_pci_dma_buf_cleanup`: DMABUF cleanup handler, which could be another fruitful entry point. 7. `vfio_pci_mmap_huge_fault`: The `vm_operations_struct` fault handler, rewritten to use the dmabuf, this is prime territory for memory mapping fuzzing. 8. `vfio_pci_core_mmap`: The entrypoint for mmap, updated to call `vfio_pci_core_mmap_prep_dmabuf`, which makes this a very important call. 9. `vfio_pci_lock_revoke_bars` and `vfio_pci_unrevoke_bars`: Revocation and Unrevocation, core functionality that could potentially be used for memory corruption. 10. `vfio_pci_dma_buf_mmap`: Specific mmap entrypoint related to the DMABUF functionality, important entrypoint. Okay, let's verify visibility, all of the non-static symbols are visible, and the static ones, `vfio_pci_mmap_huge_fault`, and `vfio_pci_dma_buf_mmap` are also visible. Now, checking the required configurations: 1. `VFIO` 2. `VFIO_PCI` 3. `VFIO_PCI_CORE` (selected by `VFIO_PCI`) 4. `VFIO_PCI_DMABUF` (depends on `PCI_P2PDMA` and `VFIO_PCI_CORE`) 5. `PCI_P2PDMA` 6. `DMA_SHARED_BUFFER` And let's drill down on `PCI_P2PDMA` further. That needs `ZONE_DEVICE`, `64BIT`, and also selects `GENERIC_ALLOCATOR` and `NEED_SG_DMA_FLAGS`, so those must be enabled, but those are not relevant. This gives us a solid set of targets and their prerequisites. I have a clear path to generating fuzzing targets that are as directed to this functionality as possible. |
| 24/2 | 2026/10/06 20:16 | tool | grepper |
0mArgs:{
"Expression": "config ZONE_DEVICE",
"PathPrefixes": [
"mm/Kconfig"
]
}Results: {
"Output": "mm/Kconfig=1220=config ZONE_DMA32\n--\nmm/Kconfig-1224-\nmm/Kconfig:1225:config ZONE_DEVICE\nmm/Kconfig-1226-\tbool \"Device memory (pmem, HMM, etc...) hotplug support\"\n"
} |
| 25/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44330 output: 33 thoughts: 18 Thoughts: Okay, let's dive into this. Here's my thought process on the matter: **Analyzing the Kconfig File for Insights** Alright, I need to get a handle on this particular aspect of the configuration process. My immediate thought is to zero in on `mm/Kconfig`. I'm going to start by focusing on the region around line 1225. That's usually a good starting point for understanding how a specific memory management feature is implemented or controlled. I expect this section likely deals with a crucial memory management variable. I'll read and analyze this section with a critical eye, looking for conditional statements, dependencies, and default values. I'm hoping to gain insights into how the specific feature is enabled, disabled, and configured. Based on the information present, I'll identify any potential issues that may be present. I also need to understand how the feature relates to any other configuration options or dependencies. Essentially, I'm expecting this area to give me the "why" and "how" behind a particular configuration setting. |
| 26/2 | 2026/10/06 20:16 | tool | read-file |
0mArgs:{
"File": "mm/Kconfig",
"FirstLine": 1224,
"LineCount": 15
}Results: {
"Contents": "1224:\t\n1225:\tconfig ZONE_DEVICE\n1226:\t\tbool \"Device memory (pmem, HMM, etc...) hotplug support\"\n1227:\t\tdepends on MEMORY_HOTREMOVE\n1228:\t\tselect XARRAY_MULTI\n1229:\t\n1230:\t\thelp\n1231:\t\t Device memory hotplug support allows for establishing pmem,\n1232:\t\t or other device driver discovered memory regions, in the\n1233:\t\t memmap. This allows pfn_to_page() lookups of otherwise\n1234:\t\t \"device-physical\" addresses which is needed for DAX, PCI_P2PDMA, and\n1235:\t\t DEVICE_PRIVATE features among others.\n1236:\t\n1237:\t\t Enabling this option will reduce the entropy of x86 KASLR memory\n1238:\t\t regions. For example - on a 46 bit system, the entropy goes down\n"
} |
| 27/2 | 2026/10/06 20:16 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44857 output: 31 thoughts: 31 Thoughts: Here's my take on this – a distilled thought process: **Dependencies and Kernel Configuration** Okay, so `ZONE_DEVICE` is flagging a dependency on `MEMORY_HOTREMOVE`. Right, that makes sense conceptually; if we want to dynamically allocate and deallocate zones, hot-removal capabilities seem crucial. Now, let me quickly confirm that dependency. I need to dive into the kernel configuration to be sure. I'll open up `mm/Kconfig` and double-check how `MEMORY_HOTREMOVE` is defined, and ideally, what other options might influence `ZONE_DEVICE`. Just want to verify the logic and potential impact of this relationship. |
| 28/2 | 2026/10/06 20:17 | tool | grepper |
0mArgs:{
"Expression": "config MEMORY_HOTREMOVE",
"PathPrefixes": [
"mm/Kconfig"
]
}Results: {
"Output": "mm/Kconfig=580=endchoice\nmm/Kconfig-581-\nmm/Kconfig:582:config MEMORY_HOTREMOVE\nmm/Kconfig-583-\tbool \"Allow for memory hot remove\"\n"
} |
| 29/2 | 2026/10/06 20:17 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45126 output: 32 thoughts: 17 Thoughts: Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person: **Analyzing the Kconfig for Memory Management** Alright, let's dive into the core configuration for memory management. I'm thinking we need to zero in on a specific section within the `mm/Kconfig` file. I recall that the area around line 580 has some particularly relevant settings, so that's where I'll begin. It likely touches on a critical aspect of memory configuration or a conditional build setting that impacts how we manage memory at runtime. I need to refresh my memory on what options are there so I can determine what is important right now. |
| 30/2 | 2026/10/06 20:17 | tool | read-file |
0mArgs:{
"File": "mm/Kconfig",
"FirstLine": 580,
"LineCount": 10
}Results: {
"Contents": " 580:\tendchoice\n 581:\t\n 582:\tconfig MEMORY_HOTREMOVE\n 583:\t\tbool \"Allow for memory hot remove\"\n 584:\t\tselect MIGRATION\n 585:\t\n 586:\tconfig MHP_MEMMAP_ON_MEMORY\n 587:\t\tdef_bool y\n 588:\t\tdepends on ARCH_MHP_MEMMAP_ON_MEMORY_ENABLE\n 589:\t\n"
} |
| 31/2 | 2026/10/06 20:17 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45453 output: 32 thoughts: 36 Thoughts: Okay, here's my interpretation of that thought process, from an expert's perspective: **Dependencies in Memory Management Configuration** Right, let's see, that `MEMORY_HOTREMOVE` configuration option... it hinges on `MEMORY_HOTPLUG`, doesn't it? Classic dependency, logically sound. Just to confirm my understanding and make sure there haven't been any recent, surprising changes in how they interact, I'll quickly check the relevant section in `mm/Kconfig`. Lines 540 through 580 are usually where these sorts of memory-related options are defined and their dependencies established. Quick scan should confirm my intuition and remind me of any nuances in their relationship. Gotta be sure I'm not missing anything crucial before moving on. |
| 32/2 | 2026/10/06 20:17 | tool | read-file |
0mArgs:{
"File": "mm/Kconfig",
"FirstLine": 540,
"LineCount": 40
}Results: {
"Contents": " 540:\t\t Select online_movable to generally disallow kernel usage of this memory.\n 541:\t\n 542:\t\t Example kernel usage would be page structs and page tables.\n 543:\t\n 544:\t\t See Documentation/admin-guide/mm/memory-hotplug.rst for more information.\n 545:\t\n 546:\tconfig MHP_DEFAULT_ONLINE_TYPE_OFFLINE\n 547:\t\tbool \"offline\"\n 548:\t\thelp\n 549:\t\t Hotplugged memory will not be onlined by default.\n 550:\t\t Choose this for systems with drivers and user policy that\n 551:\t\t handle onlining of hotplug memory policy.\n 552:\t\n 553:\tconfig MHP_DEFAULT_ONLINE_TYPE_ONLINE_AUTO\n 554:\t\tbool \"auto\"\n 555:\t\thelp\n 556:\t\t Select this if you want the kernel to automatically online\n 557:\t\t hotplugged memory into the zone it thinks is reasonable.\n 558:\t\t This memory may be utilized for kernel data.\n 559:\t\n 560:\tconfig MHP_DEFAULT_ONLINE_TYPE_ONLINE_KERNEL\n 561:\t\tbool \"kernel\"\n 562:\t\thelp\n 563:\t\t Select this if you want the kernel to automatically online\n 564:\t\t hotplugged memory into a zone capable of being used for kernel\n 565:\t\t data. This typically means ZONE_NORMAL.\n 566:\t\n 567:\tconfig MHP_DEFAULT_ONLINE_TYPE_ONLINE_MOVABLE\n 568:\t\tbool \"movable\"\n 569:\t\thelp\n 570:\t\t Select this if you want the kernel to automatically online\n 571:\t\t hotplug memory into ZONE_MOVABLE. This memory will generally\n 572:\t\t not be utilized for kernel data.\n 573:\t\n 574:\t\t This should only be used when the admin knows sufficient\n 575:\t\t ZONE_NORMAL memory is available to describe hotplug memory,\n 576:\t\t otherwise hotplug memory may fail to online. For example,\n 577:\t\t sufficient kernel-capable memory (ZONE_NORMAL) must be\n 578:\t\t available to allocate page structs to describe ZONE_MOVABLE.\n 579:\t\n"
} |
| 33/2 | 2026/10/06 20:17 | llm | patch-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 46296 output: 397 thoughts: 1847 Thoughts: **VFIO PCI DMA-BUF Integration: A Comprehensive Review** Okay, let's break this down. For the `VFIO_PCI_DMABUF` functionality, we need `PCI_P2PDMA`, which relies on `ZONE_DEVICE`, and by extension, `MEMORY_HOTREMOVE` and `MEMORY_HOTPLUG`. We're also dealing with `VFIO_PCI`, `VFIO`, `PCI_P2PDMA`, and `DMA_SHARED_BUFFER`. It's worth keeping `UDMABUF` in mind too, as it's great for testing `dma-buf` ioctls like `DMA_BUF_SET_NAME_A` and `B` without needing any physical hardware. Let's clarify the dependencies. `VFIO_PCI` selects `VFIO_PCI_CORE`, which in turn selects `DMA_SHARED_BUFFER`. And, as the Kconfig states, `PCI_P2PDMA` automatically enables `VFIO_PCI_DMABUF`. So, if both `PCI_P2PDMA` and `VFIO_PCI` are enabled, we're good to go. Remember that `VFIO_PCI` also requires `VFIO` and `PCI`. Additionally, we also need `ZONE_DEVICE`, and `MEMORY_HOTREMOVE` and `MEMORY_HOTPLUG` for `PCI_P2PDMA`. Looking at the file changes, we have modifications across `dma-buf.c`, `vfio/pci/Kconfig`, `vfio/pci/Makefile`, the HiSilicon-specific variant `hisi_acc_vfio_pci.c`, and several other core files including `vfio_pci_config.c`, `vfio_pci_core.c`, `vfio_pci_dmabuf.c`, and relevant header files. The `hisi_acc_vfio_pci.c` isn't critical here, it's just a specific variant. Now, let's pinpoint the configurations: `VFIO`, `VFIO_PCI`, `VFIO_PCI_CORE`, `PCI_P2PDMA`, `DMA_SHARED_BUFFER`, `ZONE_DEVICE`, `MEMORY_HOTPLUG`, and `MEMORY_HOTREMOVE`. Those are the critical ones. Let's validate our focus symbols to ensure we understand their functionality: `dma_buf_set_name`, `vfio_pci_core_mmap_prep_dmabuf`, `vfio_pci_dma_buf_find_pfn`, `vfio_pci_core_feature_dma_buf_revoke`, `vfio_pci_mmap_huge_fault`, `vfio_pci_core_mmap`, `vfio_pci_lock_revoke_bars`, `vfio_pci_unrevoke_bars`, `vfio_pci_dma_buf_move`, `vfio_pci_dma_buf_cleanup`, and `vfio_pci_dma_buf_mmap`. Good, each one does exist. The patch series is focusing on DMA-BUF handling within VFIO PCI, specifically refactoring DMA-BUF interactions and BAR mappings. The series exports `dma_buf_set_name` as a helper. Changes within the VFIO side introduce the `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl and also handles page faulting logic using complex synchronization with reservation locks, memory locks and `dmabuf_lock` for synchronization. All these can be tested in standard virtualized environments. These changes involve non-trivial concurrency and complex lifecycle/revocation state machines, so fuzzing would be very valuable. |
| 34/2 | 2026/10/06 20:17 | tool | set-results |
0mArgs:{
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER",
"ZONE_DEVICE",
"MEMORY_HOTPLUG",
"MEMORY_HOTREMOVE"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_mmap",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_dma_buf_cleanup",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch series modifies core DMA-BUF code (exporting and refactoring dma_buf_set_name and its ioctl handling) and heavily reworks VFIO PCI memory management and revocation. In particular, VFIO PCI BAR mmaps are now backed by DMA-BUFs, page fault handling is rewritten to navigate complex lock ordering across memory_lock, dmabuf_lock, and reservation locks, and a new feature ioctl VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE is introduced. These code paths are fully reachable in virtualized QEMU environments with emulated PCI devices and warrant targeted fuzzing.",
"WorthFuzzing": true
}Results: {
"EnableConfigs": [
"VFIO",
"VFIO_PCI",
"VFIO_PCI_CORE",
"VFIO_PCI_DMABUF",
"PCI_P2PDMA",
"DMA_SHARED_BUFFER",
"ZONE_DEVICE",
"MEMORY_HOTPLUG",
"MEMORY_HOTREMOVE"
],
"FocusSymbols": [
"dma_buf_set_name",
"vfio_pci_core_mmap_prep_dmabuf",
"vfio_pci_dma_buf_find_pfn",
"vfio_pci_core_feature_dma_buf_revoke",
"vfio_pci_mmap_huge_fault",
"vfio_pci_core_mmap",
"vfio_pci_lock_revoke_bars",
"vfio_pci_unrevoke_bars",
"vfio_pci_dma_buf_move",
"vfio_pci_dma_buf_cleanup",
"vfio_pci_dma_buf_mmap"
],
"Reasoning": "The patch series modifies core DMA-BUF code (exporting and refactoring dma_buf_set_name and its ioctl handling) and heavily reworks VFIO PCI memory management and revocation. In particular, VFIO PCI BAR mmaps are now backed by DMA-BUFs, page fault handling is rewritten to navigate complex lock ordering across memory_lock, dmabuf_lock, and reservation locks, and a new feature ioctl VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE is introduced. These code paths are fully reachable in virtualized QEMU environments with emulated PCI devices and warrant targeted fuzzing.",
"WorthFuzzing": true
} |
| 35/1 | 2026/10/06 20:17 | action | If |
8mArgs:{
"WorthFuzzing": true
}Results: null |
| 36/2 | 2026/10/06 20:17 | agent | kmsan-evaluator |
8mModel:core Args: null Results: {
"KMSANReasoning": "The patch refactors dma-buf naming by exporting `dma_buf_set_name()`, transitions VFIO PCI BAR mmaps to be backed by dma-buf allocations, and adds a `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl feature along with associated locking (`dmabuf_lock`) and fault-handling logic (`vfio_pci_dma_buf_find_pfn`).\n\nAnalysis of uninitialized memory vs other sanitizer applicability:\n1. No info leaks to userspace: The new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl only consumes input (`copy_from_user` for SET) and copies nothing to userspace. The `dma_buf_set_name` ioctl path reads from userspace via `strndup_user` and outputs no data to userspace.\n2. Memory allocations: All dynamically allocated structures (`priv`, `phys_vec`) are allocated using `kzalloc_obj`, zeroing all fields. String formatting via `kasprintf` ensures proper null-termination.\n3. Fault handling and address calculations: `vfio_pci_dma_buf_find_pfn()` uses overflow checks (`check_add_overflow`) and strictly writes `*out_pfn` on success, while the caller initializes `pfn = 0`. Mapped PFNs correspond to device MMIO BAR physical addresses, not kernel RAM.\n4. Bug profile: The changes primarily involve complex concurrency management, locking hierarchy (`memory_lock`, `dmabuf_lock`, and reservation locks), object lifetimes (`dma_buf`, `vfio_pci_core_device`, `vfile`), and refcounting (`kref`, `vfio_device_registration`). These issues (use-after-free, deadlocks, out-of-bounds accesses) are detected by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory risks or kernel information leaks that would exclusively be detected by KMSAN, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Instruction: You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.
CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
a comprehensive suite of debugging tools and sanitizers, including KASAN
(out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
(locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
or page allocations) and kernel-to-user memory info-leaks.
Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.
Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.
Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
attributes, ioctl output arguments, socket options, or BPF buffers) where fields
or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
uninitialized bytes of existing buffers.
Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
dereferences, locking deadlocks, or use-after-free bugs (these are already caught
by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.
Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.
Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.
Prompt: Target architecture: amd64
For your convenience, here is the diff of the changes:
commit 8be538d0a9753d071d10cdbcaec6469f392184a3
Author: syz-cluster <triage@syzkaller.com>
Date: Tue Oct 6 20:14:46 2026 +0000
syz-cluster: applied patch under review
diff --git a/drivers/dma-buf/dma-buf.c b/drivers/dma-buf/dma-buf.c
index 4c9add51f9ef5..b99da4ff2ef62 100644
--- a/drivers/dma-buf/dma-buf.c
+++ b/drivers/dma-buf/dma-buf.c
@@ -404,34 +404,6 @@ static __poll_t dma_buf_poll(struct file *file, poll_table *poll)
return events;
}
-/**
- * dma_buf_set_name - Set a name to a specific dma_buf to track the usage.
- * It could support changing the name of the dma-buf if the same
- * piece of memory is used for multiple purpose between different devices.
- *
- * @dmabuf: [in] dmabuf buffer that will be renamed.
- * @buf: [in] A piece of userspace memory that contains the name of
- * the dma-buf.
- *
- * Returns 0 on success. If the dma-buf buffer is already attached to
- * devices, return -EBUSY.
- *
- */
-static long dma_buf_set_name(struct dma_buf *dmabuf, const char __user *buf)
-{
- char *name = strndup_user(buf, DMA_BUF_NAME_LEN);
-
- if (IS_ERR(name))
- return PTR_ERR(name);
-
- spin_lock(&dmabuf->name_lock);
- kfree(dmabuf->name);
- dmabuf->name = name;
- spin_unlock(&dmabuf->name_lock);
-
- return 0;
-}
-
#if IS_ENABLED(CONFIG_SYNC_FILE)
static long dma_buf_export_sync_file(struct dma_buf *dmabuf,
void __user *user_data)
@@ -577,8 +549,19 @@ static long dma_buf_ioctl(struct file *file,
return ret;
case DMA_BUF_SET_NAME_A:
- case DMA_BUF_SET_NAME_B:
- return dma_buf_set_name(dmabuf, (const char __user *)arg);
+ case DMA_BUF_SET_NAME_B: {
+ char *name = strndup_user((const char __user *)arg,
+ DMA_BUF_NAME_LEN);
+
+ if (IS_ERR(name))
+ return PTR_ERR(name);
+
+ ret = dma_buf_set_name(dmabuf, name);
+ if (ret)
+ kfree(name);
+
+ return ret;
+ }
#if IS_ENABLED(CONFIG_SYNC_FILE)
case DMA_BUF_IOCTL_EXPORT_SYNC_FILE:
@@ -874,6 +857,42 @@ void dma_buf_put(struct dma_buf *dmabuf)
}
EXPORT_SYMBOL_NS_GPL(dma_buf_put, "DMA_BUF");
+/**
+ * dma_buf_set_name - Set a dmabuf's name
+ * @dmabuf: [in] dmabuf buffer that will be renamed.
+ * @name: [in] The name of the dma-buf, allocated with kmalloc() or
+ * similar. This takes ownership of the allocation
+ * on success, which will be kfree()d when the
+ * dmabuf is released or a new name assigned.
+ *
+ * Intended to be used by the exporter to set a name for debug
+ * purposes. This can also change an existing name if the same piece
+ * of memory is used for multiple purposes over time. Importers
+ * should NOT use this function to overwrite the name.
+ *
+ * Return: 0 on success, -EINVAL if the name is NULL, or -E2BIG if the
+ * name (including terminator) exceeds DMA_BUF_NAME_LEN.
+ */
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name)
+{
+ if (!name)
+ return -EINVAL;
+
+ /* dmabuffs_dname() won't use the string if the length
+ * (including terminator) exceeds DMA_BUF_NAME_LEN:
+ */
+ if (strlen(name) >= DMA_BUF_NAME_LEN)
+ return -E2BIG;
+
+ spin_lock(&dmabuf->name_lock);
+ kfree(dmabuf->name);
+ dmabuf->name = name;
+ spin_unlock(&dmabuf->name_lock);
+
+ return 0;
+}
+EXPORT_SYMBOL_NS_GPL(dma_buf_set_name, "DMA_BUF");
+
static int dma_buf_wrap_sg_table(struct sg_table **sg_table)
{
struct scatterlist *to_sg, *from_sg;
@@ -1004,6 +1023,11 @@ dma_buf_pin_on_map(struct dma_buf_attachment *attach)
* functions:
*
* - dma_buf_invalidate_mappings()
+ *
+ * 4. Exporters must not hold the dma-buf reservation lock when calling these
+ * functions:
+ *
+ * - dma_buf_set_name()
*/
/**
diff --git a/drivers/vfio/pci/Kconfig b/drivers/vfio/pci/Kconfig
index 296bf01e185ec..c6d6fb09af86c 100644
--- a/drivers/vfio/pci/Kconfig
+++ b/drivers/vfio/pci/Kconfig
@@ -6,6 +6,7 @@ config VFIO_PCI_CORE
tristate
select VFIO_VIRQFD
select IRQ_BYPASS_MANAGER
+ select DMA_SHARED_BUFFER
config VFIO_PCI_INTX
def_bool y if !S390
@@ -56,7 +57,8 @@ config VFIO_PCI_ZDEV_KVM
To enable s390x KVM vfio-pci extensions, say Y.
config VFIO_PCI_DMABUF
- def_bool y if VFIO_PCI_CORE && PCI_P2PDMA && DMA_SHARED_BUFFER
+ def_bool y if PCI_P2PDMA
+ depends on VFIO_PCI_CORE
source "drivers/vfio/pci/mlx5/Kconfig"
diff --git a/drivers/vfio/pci/Makefile b/drivers/vfio/pci/Makefile
index 6138f1bf241df..881452ea89be0 100644
--- a/drivers/vfio/pci/Makefile
+++ b/drivers/vfio/pci/Makefile
@@ -1,8 +1,7 @@
# SPDX-License-Identifier: GPL-2.0-only
-vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o
+vfio-pci-core-y := vfio_pci_core.o vfio_pci_intrs.o vfio_pci_rdwr.o vfio_pci_config.o vfio_pci_dmabuf.o
vfio-pci-core-$(CONFIG_VFIO_PCI_ZDEV_KVM) += vfio_pci_zdev.o
-vfio-pci-core-$(CONFIG_VFIO_PCI_DMABUF) += vfio_pci_dmabuf.o
obj-$(CONFIG_VFIO_PCI_CORE) += vfio-pci-core.o
vfio-pci-y := vfio_pci.o
diff --git a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
index 86362ec424a50..14622556355eb 100644
--- a/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
+++ b/drivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c
@@ -1564,6 +1564,7 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
struct hisi_acc_vf_core_device *hisi_acc_vdev = hisi_acc_get_vf_dev(core_vdev);
struct pci_dev *pdev = to_pci_dev(core_vdev->dev);
struct hisi_qm *pf_qm = hisi_acc_get_pf_qm(pdev);
+ int ret;
hisi_acc_vdev->vf_id = pci_iov_vf_id(pdev) + 1;
hisi_acc_vdev->pf_qm = pf_qm;
@@ -1575,7 +1576,18 @@ static int hisi_acc_vfio_pci_migrn_init_dev(struct vfio_device *core_vdev)
core_vdev->migration_flags = VFIO_MIGRATION_STOP_COPY | VFIO_MIGRATION_PRE_COPY;
core_vdev->mig_ops = &hisi_acc_vfio_pci_migrn_state_ops;
- return vfio_pci_core_init_dev(core_vdev);
+ ret = vfio_pci_core_init_dev(core_vdev);
+ if (ret)
+ return ret;
+ /*
+ * hisi_acc_vfio_pci_mmap() calls down to
+ * vfio_pci_core_mmap(), so BAR mappings are still
+ * DMABUF-backed. They don't require a zap on revoke, so opt
+ * out:
+ */
+ hisi_acc_vdev->core_device.zap_bars_on_revoke = false;
+
+ return 0;
}
static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {
diff --git a/drivers/vfio/pci/vfio_pci_config.c b/drivers/vfio/pci/vfio_pci_config.c
index 9914f3ac69aef..cef337f4e8f2e 100644
--- a/drivers/vfio/pci/vfio_pci_config.c
+++ b/drivers/vfio/pci/vfio_pci_config.c
@@ -590,12 +590,10 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
virt_mem = !!(le16_to_cpu(*virt_cmd) & PCI_COMMAND_MEMORY);
new_mem = !!(new_cmd & PCI_COMMAND_MEMORY);
- if (!new_mem) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (!new_mem)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
/*
* If the user is writing mem/io enable (new_mem/io) and we
@@ -631,7 +629,7 @@ static int vfio_basic_config_write(struct vfio_pci_core_device *vdev, int pos,
*virt_cmd |= cpu_to_le16(new_cmd & mask);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -712,16 +710,14 @@ static int __init init_pci_cap_basic_perm(struct perm_bits *perm)
static void vfio_lock_and_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state)
{
- if (state >= PCI_D3hot) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
- } else {
+ if (state >= PCI_D3hot)
+ vfio_pci_lock_revoke_bars(vdev);
+ else
down_write(&vdev->memory_lock);
- }
vfio_pci_set_power_state(vdev, state);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -908,11 +904,10 @@ static int vfio_exp_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_EXP_DEVCAP_FLR)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
@@ -993,11 +988,10 @@ static int vfio_af_config_write(struct vfio_pci_core_device *vdev, int pos,
&cap);
if (!ret && (cap & PCI_AF_CAP_FLR) && (cap & PCI_AF_CAP_TP)) {
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
}
diff --git a/drivers/vfio/pci/vfio_pci_core.c b/drivers/vfio/pci/vfio_pci_core.c
index 6757054e9d875..68e582ad38463 100644
--- a/drivers/vfio/pci/vfio_pci_core.c
+++ b/drivers/vfio/pci/vfio_pci_core.c
@@ -13,6 +13,8 @@
#include <linux/aperture.h>
#include <linux/debugfs.h>
#include <linux/device.h>
+#include <linux/dma-buf.h>
+#include <linux/dma-resv.h>
#include <linux/eventfd.h>
#include <linux/file.h>
#include <linux/interrupt.h>
@@ -376,8 +378,7 @@ static int vfio_pci_runtime_pm_entry(struct vfio_pci_core_device *vdev,
* The vdev power related flags are protected with 'memory_lock'
* semaphore.
*/
- vfio_pci_zap_and_down_write_memory_lock(vdev);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_lock_revoke_bars(vdev);
if (vdev->pm_runtime_engaged) {
up_write(&vdev->memory_lock);
@@ -463,7 +464,7 @@ static void vfio_pci_runtime_pm_exit(struct vfio_pci_core_device *vdev)
down_write(&vdev->memory_lock);
__vfio_pci_runtime_pm_exit(vdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
@@ -527,8 +528,14 @@ static int vfio_pci_core_runtime_resume(struct device *dev)
*/
down_write(&vdev->memory_lock);
if (vdev->pm_wake_eventfd_ctx) {
- eventfd_signal(vdev->pm_wake_eventfd_ctx);
+ struct eventfd_ctx *ctx = vdev->pm_wake_eventfd_ctx;
+
+ vdev->pm_wake_eventfd_ctx = NULL;
__vfio_pci_runtime_pm_exit(vdev);
+ if (__vfio_pci_memory_enabled(vdev))
+ vfio_pci_unrevoke_bars(vdev);
+ eventfd_signal(ctx);
+ eventfd_ctx_put(ctx);
}
up_write(&vdev->memory_lock);
@@ -663,6 +670,7 @@ int vfio_pci_core_enable(struct vfio_pci_core_device *vdev)
vdev->has_vga = true;
vfio_pci_core_map_bars(vdev);
+ vdev->bars_revoked = false;
return 0;
@@ -1312,6 +1320,8 @@ static int vfio_pci_ioctl_set_irqs(struct vfio_pci_core_device *vdev,
return ret;
}
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev);
+
static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
void __user *arg)
{
@@ -1320,7 +1330,7 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
if (!vdev->reset_works)
return -EINVAL;
- vfio_pci_zap_and_down_write_memory_lock(vdev);
+ down_write(&vdev->memory_lock);
/*
* This function can be invoked while the power state is non-D0. If
@@ -1330,13 +1340,18 @@ static int vfio_pci_ioctl_reset(struct vfio_pci_core_device *vdev,
* have NoSoftRst-, the reset function can cause the PCI config space
* reset without restoring the original state (saved locally in
* 'vdev->pm_save').
+ *
+ * The zap is done after making the device accessible in D0,
+ * because a DMABUF importer could access the device as part
+ * of its revocation cleanup.
*/
vfio_pci_set_power_state(vdev, PCI_D0);
- vfio_pci_dma_buf_move(vdev, true);
+ vfio_pci_revoke_bars(vdev);
+
ret = pci_try_reset_function(vdev->pdev);
if (__vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
return ret;
@@ -1627,6 +1642,8 @@ int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,
return vfio_pci_core_feature_dma_buf(vdev, flags, arg, argsz);
case VFIO_DEVICE_FEATURE_ZPCI_ERROR:
return vfio_pci_zdev_feature_err(device, flags, arg, argsz);
+ case VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE:
+ return vfio_pci_core_feature_dma_buf_revoke(vdev, flags, arg, argsz);
default:
return -ENOTTY;
}
@@ -1706,20 +1723,37 @@ ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *bu
}
EXPORT_SYMBOL_GPL(vfio_pci_core_write);
-static void vfio_pci_zap_bars(struct vfio_pci_core_device *vdev)
+static void vfio_pci_revoke_bars(struct vfio_pci_core_device *vdev)
{
- struct vfio_device *core_vdev = &vdev->vdev;
- loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
- loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
- loff_t len = end - start;
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, true);
- unmap_mapping_range(core_vdev->inode->i_mapping, start, len, true);
+ /*
+ * If a driver could possibly create BAR mappings in the
+ * vdev's address_space, do an additional zap on revoke. See
+ * vfio_pci_core_init_dev().
+ */
+ if (vdev->zap_bars_on_revoke) {
+ struct vfio_device *core_vdev = &vdev->vdev;
+ loff_t start = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_BAR0_REGION_INDEX);
+ loff_t end = VFIO_PCI_INDEX_TO_OFFSET(VFIO_PCI_ROM_REGION_INDEX);
+ loff_t len = end - start;
+
+ unmap_mapping_range(core_vdev->inode->i_mapping,
+ start, len, true);
+ }
}
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev)
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev)
{
down_write(&vdev->memory_lock);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
+}
+
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev)
+{
+ lockdep_assert_held_write(&vdev->memory_lock);
+ vfio_pci_dma_buf_move(vdev, false);
}
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev)
@@ -1741,18 +1775,6 @@ void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 c
up_write(&vdev->memory_lock);
}
-static unsigned long vma_to_pfn(struct vm_area_struct *vma)
-{
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- int index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
- u64 pgoff;
-
- pgoff = vma->vm_pgoff &
- ((1U << (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);
-
- return (pci_resource_start(vdev->pdev, index) >> PAGE_SHIFT) + pgoff;
-}
-
vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,
struct vm_fault *vmf,
unsigned long pfn,
@@ -1780,24 +1802,106 @@ static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,
unsigned int order)
{
struct vm_area_struct *vma = vmf->vma;
- struct vfio_pci_core_device *vdev = vma->vm_private_data;
- unsigned long addr = vmf->address & ~((PAGE_SIZE << order) - 1);
- unsigned long pgoff = linear_page_delta(vma, addr);
- unsigned long pfn = vma_to_pfn(vma) + pgoff;
- vm_fault_t ret = VM_FAULT_FALLBACK;
-
- if (is_aligned_for_order(vma, addr, pfn, order)) {
- scoped_guard(rwsem_read, &vdev->memory_lock)
- ret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);
+ struct vfio_pci_dma_buf *priv = vma->vm_private_data;
+ struct vfio_pci_core_device *vdev;
+ unsigned long pfn = 0;
+ vm_fault_t ret = VM_FAULT_SIGBUS;
+
+ /*
+ * The only thing this can rely on is that the DMABUF relating
+ * to the VMA's vm_file exists (priv).
+ *
+ * A DMABUF for a VFIO device fd mmap() holds a reference to
+ * the original VFIO device fd, but an explicitly-exported
+ * DMABUF does not. The original fd might have closed,
+ * meaning this fault can race with
+ * vfio_pci_dma_buf_cleanup(), meaning the buffer could have
+ * been revoked (in which case priv->vdev might be NULL), and
+ * the VFIO device registration might have been dropped.
+ *
+ * With the goal of taking vdev locks in a world where vdev
+ * might not still exist:
+ *
+ * 1. Take the resv lock on the DMABUF:
+ * - If racing cleanup got in first, the buffer is revoked;
+ * stop/exit if so.
+ * - If we got in first, the buffer is not revoked so vdev is
+ * non-NULL, accessible, and cleanup _has not yet put the
+ * VFIO device registration_. So, the device refcount must
+ * be >0.
+ *
+ * 2. Take vfio_device registration (refcount guaranteed >0
+ * hereafter).
+ *
+ * 3. Unlock the DMABUF's resv lock:
+ * - A racing cleanup can now complete.
+ * - But, the device refcount >0, meaning the vfio_device
+ * (and vfio_pci_core_device vdev) have not yet been
+ * freed. vdev is accessible, even if the DMABUF has been
+ * revoked or cleanup has happened, because
+ * vfio_unregister_group_dev() can't complete.
+ *
+ * 4. Take the vdev->memory_lock then vdev->dmabuf_lock:
+ * - Either the DMABUF is usable, or has been cleaned up.
+ * - It's not necessary to also take the resv lock, because
+ * the status/vdev can't change while dmabuf_lock is held.
+ * - Test the DMABUF revocation status again: if it was
+ * revoked between 1 and 4, return a SIGBUS. Otherwise,
+ * return a PFN.
+ *
+ * 5. Unlock, done.
+ */
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+
+ if (priv->status != VFIO_PCI_DMABUF_OK) {
+ pr_debug_ratelimited("%s VA 0x%lx, pgoff 0x%lx: DMABUF revoked/cleaned up\n",
+ __func__, vmf->address, vma->vm_pgoff);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+
+ /* If the buffer isn't revoked, vdev is valid */
+ vdev = priv->vdev;
+
+ if (!vfio_device_try_get_registration(&vdev->vdev)) {
+ /*
+ * If vdev != NULL (above), the registration should
+ * already be >0 and so this try_get should never
+ * fail.
+ */
+ dev_warn_ratelimited(&vdev->pdev->dev,
+ "%s: Unexpected registration failure\n",
+ __func__);
+ dma_resv_unlock(priv->dmabuf->resv);
+ return VM_FAULT_SIGBUS;
+ }
+ dma_resv_unlock(priv->dmabuf->resv);
+
+ /* memory_lock for vfio_pci_vmf_insert_pfn() */
+ down_read(&vdev->memory_lock);
+ /* Re-test revocation status under dmabuf_lock */
+ down_read(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_OK) {
+ int pres = vfio_pci_dma_buf_find_pfn(vdev, priv, vma,
+ vmf->address,
+ order, &pfn);
+
+ if (pres == 0)
+ ret = vfio_pci_vmf_insert_pfn(vdev, vmf,
+ pfn, order);
+ else if (pres == -ERANGE)
+ ret = VM_FAULT_FALLBACK;
}
+ up_read(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
dev_dbg_ratelimited(&vdev->pdev->dev,
- "%s(,order = %d) BAR %ld page offset 0x%lx: 0x%x\n",
- __func__, order,
- vma->vm_pgoff >>
- (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT),
- pgoff, (unsigned int)ret);
+ "%s(order = %d) PFN 0x%lx, VA 0x%lx, pgoff 0x%lx: 0x%x\n",
+ __func__, order, pfn, vmf->address,
+ vma->vm_pgoff, (unsigned int)ret);
+ vfio_device_put_registration(&vdev->vdev);
return ret;
}
@@ -1813,6 +1917,11 @@ static const struct vm_operations_struct vfio_pci_mmap_ops = {
#endif
};
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma)
+{
+ vma->vm_ops = &vfio_pci_mmap_ops;
+}
+
int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)
{
struct vfio_pci_core_device *vdev =
@@ -1821,6 +1930,7 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
unsigned int index;
u64 phys_len, req_len, pgoff, req_start;
void __iomem *bar_io;
+ int ret;
index = vma->vm_pgoff >> (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);
@@ -1860,7 +1970,12 @@ int vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma
if (IS_ERR(bar_io))
return PTR_ERR(bar_io);
- vma->vm_private_data = vdev;
+ ret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,
+ pci_resource_start(pdev, index),
+ req_len, index);
+ if (ret)
+ return ret;
+
vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
@@ -2197,8 +2312,19 @@ int vfio_pci_core_init_dev(struct vfio_device *core_vdev)
return ret;
INIT_LIST_HEAD(&vdev->dmabufs);
init_rwsem(&vdev->memory_lock);
+ init_rwsem(&vdev->dmabuf_lock);
xa_init(&vdev->ctx);
+ /*
+ * If a driver overrides .mmap, it has to be assumed that it
+ * might not use the DMABUF-backed core mmap; this flag
+ * enables a zap at revoke time. A driver can opt out by
+ * clearing this flag at init, if their .mmap override calls
+ * down to vfio_pci_core_mmap().
+ */
+ if (vdev->vdev.ops->mmap != vfio_pci_core_mmap)
+ vdev->zap_bars_on_revoke = true;
+
return 0;
}
EXPORT_SYMBOL_GPL(vfio_pci_core_init_dev);
@@ -2566,9 +2692,10 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
}
/*
- * Take the memory write lock for each device and zap BAR
- * mappings to prevent the user accessing the device while in
- * reset. Locking multiple devices is prone to deadlock,
+ * Take the memory write lock for each device and
+ * zap/revoke BAR mappings to prevent the user (or
+ * peers) accessing the device while in reset.
+ * Locking multiple devices is prone to deadlock,
* runaway and unwind if we hit contention.
*/
if (!down_write_trylock(&vdev->memory_lock)) {
@@ -2576,8 +2703,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
break;
}
- vfio_pci_dma_buf_move(vdev, true);
- vfio_pci_zap_bars(vdev);
+ vfio_pci_revoke_bars(vdev);
}
if (!list_entry_is_head(vdev,
@@ -2607,7 +2733,7 @@ static int vfio_pci_dev_set_hot_reset(struct vfio_device_set *dev_set,
list_for_each_entry_from_reverse(vdev, &dev_set->device_list,
vdev.dev_set_list) {
if (vdev->vdev.open_count && __vfio_pci_memory_enabled(vdev))
- vfio_pci_dma_buf_move(vdev, false);
+ vfio_pci_unrevoke_bars(vdev);
up_write(&vdev->memory_lock);
}
diff --git a/drivers/vfio/pci/vfio_pci_dmabuf.c b/drivers/vfio/pci/vfio_pci_dmabuf.c
index c16f460c01d68..b57bfaefd9fae 100644
--- a/drivers/vfio/pci/vfio_pci_dmabuf.c
+++ b/drivers/vfio/pci/vfio_pci_dmabuf.c
@@ -3,25 +3,14 @@
*/
#include <linux/dma-buf-mapping.h>
#include <linux/pci-p2pdma.h>
+#include <linux/dma-buf.h>
#include <linux/dma-resv.h>
#include "vfio_pci_priv.h"
MODULE_IMPORT_NS("DMA_BUF");
-struct vfio_pci_dma_buf {
- struct dma_buf *dmabuf;
- struct vfio_pci_core_device *vdev;
- struct list_head dmabufs_elm;
- size_t size;
- struct phys_vec *phys_vec;
- struct p2pdma_provider *provider;
- u32 nr_ranges;
- struct kref kref;
- struct completion comp;
- u8 revoked : 1;
-};
-
+#ifdef CONFIG_VFIO_PCI_DMABUF
static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
struct dma_buf_attachment *attachment)
{
@@ -30,7 +19,7 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
if (!attachment->peer2peer)
return -EOPNOTSUPP;
- if (priv->revoked)
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
return -ENODEV;
if (!dma_buf_attach_revocable(attachment))
@@ -39,6 +28,62 @@ static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
return 0;
}
+static int vfio_pci_dma_buf_mmap(struct dma_buf *dmabuf, struct vm_area_struct *vma)
+{
+ struct vfio_pci_dma_buf *priv = dmabuf->priv;
+
+ /*
+ * dma_buf_mmap_internal() has asserted that the VMA is
+ * contained within the DMABUF size before calling this.
+ *
+ * Also, if we observe that the buffer is revoked now then
+ * refuse the mmap(). This is a belt-and-braces early failure
+ * to ease debugging a revoked buffer being used. Userspace
+ * might also race an mmap() against an explicit revocation,
+ * or an action causing a revoke; race scenarios are still
+ * safe because the fault handler ultimately prevents access
+ * to a revoked buffer if it isn't caught here.
+ */
+ if (READ_ONCE(priv->status) != VFIO_PCI_DMABUF_OK)
+ return -ENODEV;
+ /*
+ * Make clear that anything with an offset adjustment is
+ * explicitly unsupported, as vfio_pci_dma_buf_find_pfn()
+ * maths would underflow; this doesn't happen through the
+ * regular DMABUF export path used with this mmap(). A DMABUF
+ * implicitly created for BAR mmap could have adjust > 0, but
+ * these can't currently be re-opened and mmap()ed again.
+ * Catch here in case that assumption ever changes.
+ */
+ if (priv->vma_pgoff_adjust)
+ return -EINVAL;
+ if ((vma->vm_flags & VM_SHARED) == 0)
+ return -EINVAL;
+
+ vma->vm_page_prot = pgprot_noncached(vma->vm_page_prot);
+ vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot);
+
+ /* See comments in vfio_pci_core_mmap() re VM_ALLOW_ANY_UNCACHED. */
+ vm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |
+ VM_DONTEXPAND | VM_DONTDUMP);
+ vma->vm_private_data = priv;
+ vfio_pci_set_vma_ops(vma);
+
+ return 0;
+}
+#else
+static int vfio_pci_dma_buf_attach(struct dma_buf *dmabuf,
+ struct dma_buf_attachment *attachment)
+{
+ /*
+ * Explicit export can't occur without the DMABUF feature, but
+ * DMABUFs are implicitly created for BAR mappings. An
+ * .attach that fails prevents dma_buf_attach().
+ */
+ return -EOPNOTSUPP;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
static void vfio_pci_dma_buf_done(struct kref *kref)
{
struct vfio_pci_dma_buf *priv =
@@ -56,7 +101,7 @@ vfio_pci_dma_buf_map(struct dma_buf_attachment *attachment,
dma_resv_assert_held(priv->dmabuf->resv);
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return ERR_PTR(-ENODEV);
ret = dma_buf_phys_vec_to_sgt(attachment, priv->provider,
@@ -90,22 +135,346 @@ static void vfio_pci_dma_buf_release(struct dma_buf *dmabuf)
* The refcount prevents both.
*/
if (priv->vdev) {
- down_write(&priv->vdev->memory_lock);
+ down_write(&priv->vdev->dmabuf_lock);
list_del_init(&priv->dmabufs_elm);
- up_write(&priv->vdev->memory_lock);
+ up_write(&priv->vdev->dmabuf_lock);
vfio_device_put_registration(&priv->vdev->vdev);
}
+ if (priv->vfile)
+ fput(priv->vfile);
kfree(priv->phys_vec);
kfree(priv);
}
static const struct dma_buf_ops vfio_pci_dmabuf_ops = {
.attach = vfio_pci_dma_buf_attach,
+#ifdef CONFIG_VFIO_PCI_DMABUF
+ .mmap = vfio_pci_dma_buf_mmap,
+#endif
.map_dma_buf = vfio_pci_dma_buf_map,
.unmap_dma_buf = vfio_pci_dma_buf_unmap,
.release = vfio_pci_dma_buf_release,
};
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long fault_addr,
+ unsigned int order,
+ unsigned long *out_pfn)
+{
+ /*
+ * Given a VMA (start, end, pgoffs) and a fault address,
+ * search the corresponding DMABUF's phys_vec[] to find the
+ * range representing the address's offset into the VMA, and
+ * its PFN. vdev must be the device that the DMABUF priv was
+ * exported from; vdev->dmabuf_lock must be held, and priv
+ * must not be revoked.
+ *
+ * The phys_vec[] ranges represent contiguous spans of VAs
+ * upwards from the buffer offset 0; the actual PFNs might be
+ * in any order, overlap/alias, etc. Calculate an offset of
+ * the desired page given VMA start/pgoff and address, then
+ * search upwards from 0 to find which span contains it.
+ *
+ * On success, a valid PFN for a page sized by 'order' is
+ * returned into out_pfn.
+ *
+ * Failure occurs if:
+ * - A hugepage would cross the edge of the VMA,
+ * - A hugepage isn't entirely contained within a range
+ * (including where it straddles the boundary between
+ * ranges),
+ * - We find a range, but the final PFN isn't aligned to the
+ * requested order.
+ *
+ * Upon failure, -ERANGE is returned and the caller is
+ * expected to try again with a smaller order, which will
+ * eventually succeed.
+ *
+ * It's suboptimal if DMABUFs are created with neighbouring
+ * ranges that are physically contiguous, since hugepages
+ * can't straddle range boundaries. (The construction of the
+ * ranges should merge them in this case.)
+ *
+ * Finally, vma_pgoff_adjust is used with a DMABUF created for
+ * a VFIO BAR mmap: a BAR mapped with vm_pgoff > 0 creates a
+ * DMABUF such that byte 0 of the VMA corresponds to byte 0 of
+ * the DMABUF and byte 'vm_pgoff << PAGE_SHIFT' into the BAR.
+ * To avoid double-offsetting in this scenario, subtracting
+ * vma_pgoff_adjust from this (non-zero) vm_pgoff generates
+ * the effective offset. This also removes the VFIO region
+ * index encoded in vm_pgoff for VFIO BAR mmaps.
+ */
+
+ const unsigned long pagesize = PAGE_SIZE << order;
+ unsigned long vma_off = (vma->vm_pgoff - priv->vma_pgoff_adjust) <<
+ PAGE_SHIFT;
+ unsigned long rounded_page_addr = ALIGN_DOWN(fault_addr, pagesize);
+ unsigned long rounded_page_end = rounded_page_addr + pagesize;
+ unsigned long fault_offset;
+ unsigned long fault_offset_end;
+ unsigned long range_start_offset = 0;
+ unsigned int i;
+ int ret;
+
+ if (unlikely(!vdev))
+ return -ENODEV;
+
+ /* This prevents the dmabuf revocation state from changing under us */
+ lockdep_assert_held(&vdev->dmabuf_lock);
+
+ if (unlikely(priv->vdev != vdev || priv->status != VFIO_PCI_DMABUF_OK))
+ return -ENODEV;
+
+ if (rounded_page_addr < vma->vm_start || rounded_page_end > vma->vm_end) {
+ if (order > 0)
+ return -ERANGE;
+
+ /* A fault address outside of the VMA is absurd. */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "Fault addr 0x%lx outside VMA 0x%lx-0x%lx\n",
+ fault_addr, vma->vm_start, vma->vm_end);
+ return -EFAULT;
+ }
+
+ /*
+ * fault_offset[_end] is the span within the DMABUF
+ * corresponding to the faulting page:
+ */
+ if (unlikely(check_add_overflow(rounded_page_addr - vma->vm_start,
+ vma_off, &fault_offset) ||
+ check_add_overflow(fault_offset, pagesize,
+ &fault_offset_end)))
+ return -EFAULT;
+
+ /*
+ * Iterate over ranges in the buffer, summing their lengths:
+ * range_start_offset represents the current range's starting
+ * offset in the buffer (from 0 upwards).
+ *
+ * A failure for order == 0 is unexpected, and triggers a
+ * fault/warn.
+ */
+ ret = (order == 0) ? -EFAULT : -ERANGE;
+
+ for (i = 0; i < priv->nr_ranges; i++) {
+ size_t range_len = priv->phys_vec[i].len;
+
+ /* Early exit if range starts after the page end */
+ if (fault_offset_end <= range_start_offset)
+ break;
+
+ if (fault_offset >= range_start_offset &&
+ fault_offset_end <= range_start_offset + range_len) {
+ /*
+ * The faulting page is wholly contained
+ * within the span represented by this range,
+ * so validate PFN alignment for the order.
+ * The if() condition ensures the pfn
+ * arithmetic won't overflow.
+ */
+ unsigned long pfn =
+ ((fault_offset - range_start_offset) +
+ priv->phys_vec[i].paddr) >> PAGE_SHIFT;
+
+ if (IS_ALIGNED(pfn, 1 << order)) {
+ *out_pfn = pfn;
+ ret = 0;
+ }
+ /*
+ * Else order > 0; ERANGE retries with smaller
+ * order
+ */
+ break;
+ }
+ range_start_offset += range_len;
+ }
+
+ if (order == 0 && ret != 0)
+ /*
+ * The address fell outside of the span represented by
+ * the (concatenated) ranges. As setup of a mapping
+ * ensures that the VMA is <= the total size of the
+ * ranges this should never happen. If it does, warn
+ * and SIGBUS.
+ */
+ dev_warn_ratelimited(
+ &vdev->pdev->dev,
+ "No range for addr 0x%lx, order %d: VMA 0x%lx-0x%lx pgoff 0x%lx, %u ranges, size 0x%zx\n",
+ fault_addr, order, vma->vm_start, vma->vm_end,
+ vma->vm_pgoff, priv->nr_ranges, priv->size);
+
+ return ret;
+}
+
+/*
+ * Create a DMABUF corresponding to priv, add it to vdev->dmabufs list
+ * for tracking (meaning cleanup or revocation will zap it), and take
+ * a vfio_device registration.
+ */
+static int vfio_pci_dmabuf_export(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv, u32 flags)
+{
+ DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
+
+ if (!vfio_device_try_get_registration(&vdev->vdev))
+ return -ENODEV;
+
+ exp_info.ops = &vfio_pci_dmabuf_ops;
+ exp_info.size = priv->size;
+ exp_info.flags = flags;
+ exp_info.priv = priv;
+
+ priv->dmabuf = dma_buf_export(&exp_info);
+ if (IS_ERR(priv->dmabuf)) {
+ vfio_device_put_registration(&vdev->vdev);
+ return PTR_ERR(priv->dmabuf);
+ }
+
+ kref_init(&priv->kref);
+ init_completion(&priv->comp);
+
+ /* dma_buf_put() now frees priv */
+ INIT_LIST_HEAD(&priv->dmabufs_elm);
+
+ /*
+ * dmabuf_lock synchronises access (R) or updates (W) to the
+ * vdev->dmabufs list and to bars_revoked (see below). The
+ * revocation state of DMABUF elements in the list is written
+ * holding both dmabuf_lock(W) and resv, and tested with
+ * either.
+ *
+ * (memory_lock, if held ->) dmabuf_lock -> resv
+ *
+ * NOTE: memory_lock is strictly avoided here, to avoid a
+ * dependency on memory_lock when mmap_lock is held, when
+ * mmap() leads to export. vfio-pci variant drivers are
+ * permitted to hold memory_lock across actions that might
+ * fault (such as user access); a deadlock could result when
+ * that fault path attempts to take mmap_lock (if held by an
+ * export waiting for memory_lock).
+ *
+ * vdev->bars_revoked tracks the BAR revocation status updated
+ * via vfio_pci_dma_buf_move(), so the initial DMABUF state
+ * follows the same criteria that later update the DMABUF
+ * state (BAR zap, etc.).
+ */
+ lockdep_assert_not_held(&vdev->memory_lock);
+
+ down_write(&vdev->dmabuf_lock);
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ priv->status = vdev->bars_revoked ? VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK;
+ list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
+ dma_resv_unlock(priv->dmabuf->resv);
+ up_write(&vdev->dmabuf_lock);
+
+ return 0;
+}
+
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index)
+{
+ struct vfio_pci_dma_buf *priv;
+ unsigned long vma_pgoff = vma->vm_pgoff & (VFIO_PCI_OFFSET_MASK >> PAGE_SHIFT);
+ char *bufname;
+ int ret;
+
+ priv = kzalloc_obj(*priv);
+ if (!priv)
+ return -ENOMEM;
+
+ priv->phys_vec = kzalloc_obj(*priv->phys_vec);
+ if (!priv->phys_vec) {
+ ret = -ENOMEM;
+ goto err_free_priv;
+ }
+
+ /*
+ * Debug name: The absolute maximum size of the name
+ * ('vfio:ffffffff:ff:1f.7/5') fits within DMA_BUF_NAME_LEN.
+ */
+ bufname = kasprintf(GFP_KERNEL, "vfio:%s/%x",
+ pci_name(vdev->pdev),
+ res_index);
+
+ if (!bufname) {
+ ret = -ENOMEM;
+ goto err_free_phys;
+ }
+
+ /*
+ * The DMABUF begins from the mmap()'s BAR offset, i.e. the
+ * start of the VMA corresponds to byte 0 of the DMABUF and
+ * byte (vma_pgoff << PAGE_SHIFT) of the BAR.
+ *
+ * vfio_pci_dma_buf_find_pfn() reverses this offset using
+ * vma_pgoff_adjust, so that ultimately a fault's offset from
+ * the start of the _VMA_ has a consistent usage whether the
+ * VMA originates from an mmap() of the VFIO device here or a
+ * direct DMABUF mmap(). Note vma_pgoff_adjust also includes
+ * the encoded VFIO region index, which cancels out the index
+ * encoded in vm_pgoff.
+ */
+ priv->vdev = vdev;
+ priv->size = req_len;
+ priv->nr_ranges = 1;
+ priv->vma_pgoff_adjust = vma->vm_pgoff;
+
+ /*
+ * The provider can be NULL _iff_ the DMABUF feature isn't
+ * supported, because it's only used by DMABUF import and
+ * attach is prohibited if the feature isn't present.
+ */
+ priv->provider = pcim_p2pdma_provider(vdev->pdev, res_index);
+ if (IS_ENABLED(CONFIG_VFIO_PCI_DMABUF) && !priv->provider) {
+ ret = -EINVAL;
+ goto err_free_name;
+ }
+
+ priv->phys_vec[0].paddr = phys_start + ((u64)vma_pgoff << PAGE_SHIFT);
+ priv->phys_vec[0].len = priv->size;
+
+ ret = vfio_pci_dmabuf_export(vdev, priv, O_RDWR);
+ if (ret)
+ goto err_free_name;
+
+ if (dma_buf_set_name(priv->dmabuf, bufname)) {
+ dev_dbg_ratelimited(&vdev->pdev->dev,
+ "Failed to set map name '%s'\n",
+ bufname);
+ kfree(bufname);
+ }
+
+ /*
+ * Ownership of the DMABUF file transfers to the VMA so that
+ * other users can locate the DMABUF via a VA. Ownership of
+ * the original VFIO device file being mmap()ed transfers to
+ * priv, and is put when the DMABUF is released. This
+ * intentionally does not use get_file()/vma_set_file()
+ * because the references are already held, and ownership
+ * moves.
+ */
+ priv->vfile = vma->vm_file;
+ vma->vm_file = priv->dmabuf->file;
+ vma->vm_private_data = priv;
+
+ return 0;
+
+err_free_name:
+ kfree(bufname);
+err_free_phys:
+ kfree(priv->phys_vec);
+err_free_priv:
+ kfree(priv);
+ return ret;
+}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
/*
* This is a temporary "private interconnect" between VFIO DMABUF and iommufd.
* It allows the two co-operating drivers to exchange the physical address of
@@ -128,7 +497,7 @@ int vfio_pci_dma_buf_iommufd_map(struct dma_buf_attachment *attachment,
return -EOPNOTSUPP;
priv = attachment->dmabuf->priv;
- if (priv->revoked)
+ if (priv->status != VFIO_PCI_DMABUF_OK)
return -ENODEV;
/* More than one range to iommufd will require proper DMABUF support */
@@ -224,7 +593,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
struct vfio_device_feature_dma_buf get_dma_buf = {};
struct vfio_region_dma_range *dma_ranges;
- DEFINE_DMA_BUF_EXPORT_INFO(exp_info);
struct vfio_pci_dma_buf *priv;
size_t length;
int ret;
@@ -284,34 +652,9 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
dma_ranges = NULL;
- if (!vfio_device_try_get_registration(&vdev->vdev)) {
- ret = -ENODEV;
+ ret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);
+ if (ret)
goto err_free_phys;
- }
-
- exp_info.ops = &vfio_pci_dmabuf_ops;
- exp_info.size = priv->size;
- exp_info.flags = get_dma_buf.open_flags;
- exp_info.priv = priv;
-
- priv->dmabuf = dma_buf_export(&exp_info);
- if (IS_ERR(priv->dmabuf)) {
- ret = PTR_ERR(priv->dmabuf);
- goto err_dev_put;
- }
-
- kref_init(&priv->kref);
- init_completion(&priv->comp);
-
- /* dma_buf_put() now frees priv */
- INIT_LIST_HEAD(&priv->dmabufs_elm);
- down_write(&vdev->memory_lock);
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = !__vfio_pci_memory_enabled(vdev);
- list_add_tail(&priv->dmabufs_elm, &vdev->dmabufs);
- dma_resv_unlock(priv->dmabuf->resv);
- up_write(&vdev->memory_lock);
-
/*
* dma_buf_fd() consumes the reference, when the file closes the dmabuf
* will be released.
@@ -322,8 +665,6 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
return ret;
-err_dev_put:
- vfio_device_put_registration(&vdev->vdev);
err_free_phys:
kfree(priv->phys_vec);
err_free_priv:
@@ -332,6 +673,69 @@ int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
kfree(dma_ranges);
return ret;
}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
+
+/*
+ * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives
+ * the guarantee that all future map/attach attempts will fail no
+ * matter what, whereas REVOKED can transition back to OK.
+ */
+static void vfio_pci_dma_buf_set_status(struct vfio_pci_dma_buf *priv,
+ enum vfio_pci_dma_buf_status new_status)
+{
+ bool was_revoked;
+
+ /*
+ * Changes to the DMABUF's revocation status are synchronised
+ * using dmabuf_lock:
+ */
+ lockdep_assert_held_write(&priv->vdev->dmabuf_lock);
+
+ /* If DEAD, state can no longer change */
+ if (priv->status == VFIO_PCI_DMABUF_DEAD ||
+ priv->status == new_status)
+ return;
+
+ dma_resv_lock(priv->dmabuf->resv, NULL);
+ was_revoked = (priv->status == VFIO_PCI_DMABUF_REVOKED);
+
+ if (new_status != VFIO_PCI_DMABUF_OK) {
+ priv->status = new_status;
+
+ if (was_revoked) {
+ /*
+ * A REVOKED buffer is being marked DEAD.
+ * invalidate_mappings/unmap wait happened
+ * when it became REVOKED, don't wait again.
+ */
+ dma_resv_unlock(priv->dmabuf->resv);
+ return;
+ }
+ dma_buf_invalidate_mappings(priv->dmabuf);
+ dma_resv_wait_timeout(priv->dmabuf->resv,
+ DMA_RESV_USAGE_BOOKKEEP, false,
+ MAX_SCHEDULE_TIMEOUT);
+ dma_resv_unlock(priv->dmabuf->resv);
+ kref_put(&priv->kref, vfio_pci_dma_buf_done);
+ wait_for_completion(&priv->comp);
+ unmap_mapping_range(priv->dmabuf->file->f_mapping,
+ 0, 0, true);
+ /*
+ * Re-arm the registered kref reference and the
+ * completion so the post-revoke state matches the
+ * post-creation state. An un-revoke followed by a
+ * new mapping needs the kref to be non-zero before
+ * kref_get(), and vfio_pci_dma_buf_cleanup()
+ * delegates its drain back through this revoke
+ * path on a possibly-already-revoked dma-buf.
+ */
+ kref_init(&priv->kref);
+ reinit_completion(&priv->comp);
+ } else {
+ priv->status = VFIO_PCI_DMABUF_OK;
+ dma_resv_unlock(priv->dmabuf->resv);
+ }
+}
void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
{
@@ -340,41 +744,17 @@ void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked)
lockdep_assert_held_write(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ vdev->bars_revoked = revoked;
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
-
- if (priv->revoked != revoked) {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- if (revoked)
- priv->revoked = true;
- dma_buf_invalidate_mappings(priv->dmabuf);
- dma_resv_wait_timeout(priv->dmabuf->resv,
- DMA_RESV_USAGE_BOOKKEEP, false,
- MAX_SCHEDULE_TIMEOUT);
- dma_resv_unlock(priv->dmabuf->resv);
- if (revoked) {
- kref_put(&priv->kref, vfio_pci_dma_buf_done);
- wait_for_completion(&priv->comp);
- /*
- * Re-arm the registered kref reference and the
- * completion so the post-revoke state matches the
- * post-creation state. An un-revoke followed by a
- * new mapping needs the kref to be non-zero before
- * kref_get(), and vfio_pci_dma_buf_cleanup()
- * delegates its drain back through this revoke
- * path on a possibly-already-revoked dma-buf.
- */
- kref_init(&priv->kref);
- reinit_completion(&priv->comp);
- } else {
- dma_resv_lock(priv->dmabuf->resv, NULL);
- priv->revoked = false;
- dma_resv_unlock(priv->dmabuf->resv);
- }
- }
+ vfio_pci_dma_buf_set_status(priv, revoked ?
+ VFIO_PCI_DMABUF_REVOKED :
+ VFIO_PCI_DMABUF_OK);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
}
void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
@@ -393,14 +773,85 @@ void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
*/
vfio_pci_dma_buf_move(vdev, true);
+ down_write(&vdev->dmabuf_lock);
list_for_each_entry_safe(priv, tmp, &vdev->dmabufs, dmabufs_elm) {
if (!get_file_active(&priv->dmabuf->file))
continue;
list_del_init(&priv->dmabufs_elm);
- priv->vdev = NULL;
+ WRITE_ONCE(priv->vdev, NULL);
vfio_device_put_registration(&vdev->vdev);
fput(priv->dmabuf->file);
}
+ up_write(&vdev->dmabuf_lock);
up_write(&vdev->memory_lock);
}
+
+#ifdef CONFIG_VFIO_PCI_DMABUF
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
+{
+ struct vfio_device_feature_dma_buf_revoke db_revoke;
+ struct vfio_pci_dma_buf *priv;
+ struct dma_buf *dmabuf;
+ int ret;
+
+ if (!vdev->pci_ops || !vdev->pci_ops->get_dmabuf_phys)
+ return -EOPNOTSUPP;
+
+ ret = vfio_check_feature(flags, argsz,
+ VFIO_DEVICE_FEATURE_SET,
+ sizeof(db_revoke));
+ if (ret != 1)
+ return ret;
+
+ if (copy_from_user(&db_revoke, arg, sizeof(db_revoke)))
+ return -EFAULT;
+
+ dmabuf = dma_buf_get(db_revoke.dmabuf_fd);
+ if (IS_ERR(dmabuf))
+ return PTR_ERR(dmabuf);
+
+ priv = dmabuf->priv;
+ /*
+ * Sanity-check the DMABUF is really a vfio_pci_dma_buf _and_
+ * relates to the VFIO device it was provided with.
+ *
+ * If the DMABUF relates to this vdev then priv->vdev is
+ * stable because this open fd prevents cleanup.
+ *
+ * If it relates to a different vdev, reading priv->vdev might
+ * race with a concurrent cleanup on that device. But if so,
+ * it points to a non-matching vdev or NULL and is unusable
+ * either way.
+ */
+ if (dmabuf->ops != &vfio_pci_dmabuf_ops ||
+ READ_ONCE(priv->vdev) != vdev) {
+ ret = -ENODEV;
+ goto out_put_buf;
+ }
+
+ /*
+ * memory_lock(R) is taken to stop vfio_pci_dev_set_hot_reset()
+ * from getting it and then blocking all devices in the dev_set behind
+ * this revoke's drain.
+ */
+ down_read(&vdev->memory_lock);
+ down_write(&vdev->dmabuf_lock);
+ if (priv->status == VFIO_PCI_DMABUF_DEAD) {
+ ret = -EBADFD;
+ } else {
+ vfio_pci_dma_buf_set_status(priv, VFIO_PCI_DMABUF_DEAD);
+ ret = 0;
+ }
+ up_write(&vdev->dmabuf_lock);
+ up_read(&vdev->memory_lock);
+
+out_put_buf:
+ dma_buf_put(dmabuf);
+
+ return ret;
+}
+#endif /* CONFIG_VFIO_PCI_DMABUF */
diff --git a/drivers/vfio/pci/vfio_pci_priv.h b/drivers/vfio/pci/vfio_pci_priv.h
index 4e7162234a2eb..ca12221af5553 100644
--- a/drivers/vfio/pci/vfio_pci_priv.h
+++ b/drivers/vfio/pci/vfio_pci_priv.h
@@ -23,6 +23,27 @@ struct vfio_pci_ioeventfd {
bool test_mem;
};
+enum vfio_pci_dma_buf_status {
+ VFIO_PCI_DMABUF_OK = 0,
+ VFIO_PCI_DMABUF_REVOKED = 1,
+ VFIO_PCI_DMABUF_DEAD = 2,
+};
+
+struct vfio_pci_dma_buf {
+ struct dma_buf *dmabuf;
+ struct vfio_pci_core_device *vdev;
+ struct list_head dmabufs_elm;
+ size_t size;
+ struct phys_vec *phys_vec;
+ struct p2pdma_provider *provider;
+ struct file *vfile;
+ u32 nr_ranges;
+ struct kref kref;
+ struct completion comp;
+ unsigned long vma_pgoff_adjust;
+ enum vfio_pci_dma_buf_status status;
+};
+
bool vfio_pci_intx_mask(struct vfio_pci_core_device *vdev);
void vfio_pci_intx_unmask(struct vfio_pci_core_device *vdev);
@@ -68,7 +89,8 @@ void vfio_config_free(struct vfio_pci_core_device *vdev);
int vfio_pci_set_power_state(struct vfio_pci_core_device *vdev,
pci_power_t state);
-void vfio_pci_zap_and_down_write_memory_lock(struct vfio_pci_core_device *vdev);
+void vfio_pci_lock_revoke_bars(struct vfio_pci_core_device *vdev);
+void vfio_pci_unrevoke_bars(struct vfio_pci_core_device *vdev);
u16 vfio_pci_memory_lock_and_enable(struct vfio_pci_core_device *vdev);
void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev,
u16 cmd);
@@ -123,12 +145,28 @@ static inline bool vfio_pci_is_vga(struct pci_dev *pdev)
return (pdev->class >> 8) == PCI_CLASS_DISPLAY_VGA;
}
+int vfio_pci_dma_buf_find_pfn(struct vfio_pci_core_device *vdev,
+ struct vfio_pci_dma_buf *priv,
+ struct vm_area_struct *vma,
+ unsigned long address,
+ unsigned int order,
+ unsigned long *out_pfn);
+int vfio_pci_core_mmap_prep_dmabuf(struct vfio_pci_core_device *vdev,
+ struct vm_area_struct *vma,
+ u64 phys_start, u64 req_len,
+ unsigned int res_index);
+void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
+void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+void vfio_pci_set_vma_ops(struct vm_area_struct *vma);
+
#ifdef CONFIG_VFIO_PCI_DMABUF
int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
struct vfio_device_feature_dma_buf __user *arg,
size_t argsz);
-void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev);
-void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev, bool revoked);
+int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz);
#else
static inline int
vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
@@ -137,12 +175,12 @@ vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,
{
return -ENOTTY;
}
-static inline void vfio_pci_dma_buf_cleanup(struct vfio_pci_core_device *vdev)
-{
-}
-static inline void vfio_pci_dma_buf_move(struct vfio_pci_core_device *vdev,
- bool revoked)
+static inline int vfio_pci_core_feature_dma_buf_revoke(
+ struct vfio_pci_core_device *vdev, u32 flags,
+ struct vfio_device_feature_dma_buf_revoke __user *arg,
+ size_t argsz)
{
+ return -ENOTTY;
}
#endif
diff --git a/include/linux/dma-buf.h b/include/linux/dma-buf.h
index d15b2b31d3c91..0f88132c06546 100644
--- a/include/linux/dma-buf.h
+++ b/include/linux/dma-buf.h
@@ -342,12 +342,14 @@ struct dma_buf {
/**
* @name:
*
- * Userspace-provided name. Default value is NULL. If not NULL,
- * length cannot be longer than DMA_BUF_NAME_LEN, including NIL
- * char. Useful for accounting and debugging. Read/Write accesses
- * are protected by @name_lock
- *
- * See the IOCTLs DMA_BUF_SET_NAME or DMA_BUF_SET_NAME_A/B
+ * Exporter or userspace-provided name. Default value is
+ * NULL. If not NULL, length cannot be longer than
+ * DMA_BUF_NAME_LEN, including NIL char. Useful for accounting
+ * and debugging. Read/Write accesses are protected by
+ * @name_lock
+ *
+ * See dma_buf_set_name(), and the IOCTLs DMA_BUF_SET_NAME or
+ * DMA_BUF_SET_NAME_A/B
*/
const char *name;
@@ -571,6 +573,8 @@ void dma_buf_fd_install(struct dma_buf *dmabuf, int fd);
struct dma_buf *dma_buf_get(int fd);
void dma_buf_put(struct dma_buf *dmabuf);
+int dma_buf_set_name(struct dma_buf *dmabuf, char *name);
+
struct sg_table *dma_buf_map_attachment(struct dma_buf_attachment *,
enum dma_data_direction);
void dma_buf_unmap_attachment(struct dma_buf_attachment *, struct sg_table *,
diff --git a/include/linux/vfio_pci_core.h b/include/linux/vfio_pci_core.h
index 9a1674c152aa2..44891fdb7c76e 100644
--- a/include/linux/vfio_pci_core.h
+++ b/include/linux/vfio_pci_core.h
@@ -129,11 +129,13 @@ struct vfio_pci_core_device {
bool disable_idle_d3:1;
bool nointxmask:1;
bool disable_vga:1;
+ bool zap_bars_on_revoke:1;
/* Flags modified at runtime - dedicated storage unit */
bool needs_reset;
bool pm_intx_masked;
bool pm_runtime_engaged;
bool sriov_active;
+ bool bars_revoked;
struct pci_saved_state *pci_saved_state;
struct pci_saved_state *pm_save;
int ioeventfds_nr;
@@ -148,6 +150,7 @@ struct vfio_pci_core_device {
struct vfio_pci_core_device *sriov_pf_core_dev;
struct notifier_block nb;
struct rw_semaphore memory_lock;
+ struct rw_semaphore dmabuf_lock;
struct list_head dmabufs;
};
diff --git a/include/uapi/linux/vfio.h b/include/uapi/linux/vfio.h
index e41437fa17ad0..d3c6057983e09 100644
--- a/include/uapi/linux/vfio.h
+++ b/include/uapi/linux/vfio.h
@@ -1555,6 +1555,30 @@ struct vfio_device_feature_zpci_err {
#define VFIO_DEVICE_FEATURE_ZPCI_ERROR 13
+/**
+ * Given a DMABUF fd previously exported from the same device by
+ * VFIO_DEVICE_FEATURE_DMA_BUF, a SET of this feature requests that
+ * access to the corresponding DMABUF is immediately revoked. On
+ * successful return, the buffer is no longer accessible through any
+ * VMA or DMABUF import. Thereafter, VFIO also refuses all future
+ * mmap()s and map/attach requests from any new/existing importer.
+ *
+ * Return: 0 on success, -1 and errno is set on failure:
+ *
+ * EBADF, EINVAL: dmabuf_fd is not a DMABUF fd.
+ * EOPNOTSUPP: The VFIO device does not support DMABUF export.
+ * ENODEV: The DMABUF was not exported from this device.
+ * EBADFD: The DMABUF is already revoked by this feature.
+ *
+ * Additionally, common errors can occur: EFAULT accessing the struct,
+ * or EINVAL requesting an unsupported feature op.
+ */
+#define VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE 14
+
+struct vfio_device_feature_dma_buf_revoke {
+ __s32 dmabuf_fd;
+};
+
/* -------- API for Type1 VFIO IOMMU -------- */
/**
|
| 37/3 | 2026/10/06 20:17 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 23953 output: 20 thoughts: 1227 Thoughts: **KMSAN Fuzzing Session Feasibility Assessment**
Okay, let's take a look at this patch series and see if it warrants a dedicated KMSAN fuzzing session. My initial thought is that it's worth the effort to investigate.
First, I need to understand what's actually being changed. Let's break down the patches:
* **`dma-buf.c`:** This is interesting. The refactoring of `dma_buf_set_name` with `strndup_user` and the new `dma_buf_set_name` function is a candidate for fuzzing as it involves user-provided input. Handling `strlen(name) >= DMA_BUF_NAME_LEN` and the freeing/setting of `dmabuf->name` could expose vulnerabilities, especially if there are any race conditions.
* **`vfio/pci/Kconfig` & `Makefile`:** These patches enable DMA buffer support within VFIO-PCI. This indicates a potential new attack surface by making DMA buffers first class citizens in the vfio-pci subsystem.
* **`hisi_acc_vfio_pci.c`:** This seems less critical in terms of fuzzing focus. It's just setting a boolean that disables something for this specific hardware, but it might be worth verifying that setting the bool does not open any potential issues.
* **`vfio_pci_config.c`:** Swapping out functions related to bar revocation and the move, as well as the changes in `runtime_resume` are good. This is a point of attack given it involves multiple events interacting.
* **`vfio_pci_core.c`:** *This* is where things get interesting. The integration of dmabuf backed mmaps using `vfio_pci_mmap_huge_fault` and `vfio_pci_core_mmap` now calling `vfio_pci_core_mmap_prep_dmabuf`. `init_rwsem` and the `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl are very good spots for potential bugs. I should really pay special attention to memory management when investigating this one. The fact that the `mmap` is supporting DMA buffers suggests there might be interaction between user-space and kernel DMA with associated concurrency and memory safety issues.
* I'll need to focus on how `vfio_pci_core_mmap_prep_dmabuf` is handling the allocation and initialization of `struct vfio_pci_dma_buf`. The `kzalloc_obj` calls, `kasprintf`, and dmabuf allocation are all potential areas of memory corruption and race conditions.
* **`vfio_pci_dmabuf.c`:** This adds several mmap, find, export, and cleanup handlers specifically for DMA buffers. The addition of `vfio_pci_dma_buf_set_status`, `vfio_pci_dma_buf_move` and `vfio_pci_dma_buf_cleanup` are red flags. The `vfio_pci_core_feature_dma_buf_revoke` call and the `kzalloc_obj`, `kasprintf`, and allocation within `vfio_pci_core_mmap_prep_dmabuf` are all important.
* The `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl is especially worth scrutinizing. I need to figure out if it supports GET or only SET, and what `vfio_check_feature` does. The `copy_from_user` call (which is `SET` only) is important to explore since this is an IOCTL that takes user provided input.
* **`dma-buf.h`, `vfio_pci_core.h`, `vfio.h`:** Declarations and definitions - these need to be reviewed to see what's being introduced and used.
Overall, it looks like there are plenty of interesting spots here. I need to dive deeper into the code, especially around the ioctls, memory management, and how DMA buffers interact with user-space.
I'll check `vfio_check_feature` in the codebase to understand its behavior. This should be an interesting start for a KMSAN session.
|
| 38/3 | 2026/10/06 20:17 | tool | grepper |
0mArgs:{
"Expression": "vfio_check_feature"
}Results: {
"Output": "drivers/vfio/cdx/main.c=84=static int vfio_cdx_bm_ctrl(struct vfio_device *core_vdev, u32 flags,\n--\ndrivers/vfio/cdx/main.c-97-\ndrivers/vfio/cdx/main.c:98:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/cdx/main.c-99-\t\t\t\t sizeof(ops));\n--\ndrivers/vfio/pci/vfio_pci_core.c=396=static int vfio_pci_core_pm_entry(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_core.c-400-\ndrivers/vfio/pci/vfio_pci_core.c:401:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/pci/vfio_pci_core.c-402-\tif (ret != 1)\n--\ndrivers/vfio/pci/vfio_pci_core.c=414=static int vfio_pci_core_pm_entry_with_wakeup(\n--\ndrivers/vfio/pci/vfio_pci_core.c-422-\ndrivers/vfio/pci/vfio_pci_core.c:423:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/pci/vfio_pci_core.c-424-\t\t\t\t sizeof(entry));\n--\ndrivers/vfio/pci/vfio_pci_core.c=471=static int vfio_pci_core_pm_exit(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_core.c-475-\ndrivers/vfio/pci/vfio_pci_core.c:476:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/pci/vfio_pci_core.c-477-\tif (ret != 1)\n--\ndrivers/vfio/pci/vfio_pci_core.c=1598=static int vfio_pci_core_feature_token(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1610-\t */\ndrivers/vfio/pci/vfio_pci_core.c:1611:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET,\ndrivers/vfio/pci/vfio_pci_core.c-1612-\t\t\t\t sizeof(uuid));\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=590=int vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-602-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:603:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/pci/vfio_pci_dmabuf.c-604-\t\t\t\t sizeof(get_dma_buf));\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c=791=int vfio_pci_core_feature_dma_buf_revoke(\n--\ndrivers/vfio/pci/vfio_pci_dmabuf.c-803-\ndrivers/vfio/pci/vfio_pci_dmabuf.c:804:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/pci/vfio_pci_dmabuf.c-805-\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n--\ndrivers/vfio/pci/vfio_pci_zdev.c=145=int vfio_pci_zdev_feature_err(struct vfio_device *device, u32 flags,\n--\ndrivers/vfio/pci/vfio_pci_zdev.c-158-\ndrivers/vfio/pci/vfio_pci_zdev.c:159:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/pci/vfio_pci_zdev.c-160-\t\t\t\t sizeof(err));\n--\ndrivers/vfio/vfio_main.c=931=vfio_ioctl_device_feature_mig_device_state(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-943-\ndrivers/vfio/vfio_main.c:944:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-945-\t\t\t\t VFIO_DEVICE_FEATURE_SET |\n--\ndrivers/vfio/vfio_main.c=981=vfio_ioctl_device_feature_migration_data_size(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-991-\ndrivers/vfio/vfio_main.c:992:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/vfio_main.c-993-\t\t\t\t sizeof(data_size));\n--\ndrivers/vfio/vfio_main.c=1009=vfio_ioctl_device_feature_migration_precopy_info_v2(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1016-\ndrivers/vfio/vfio_main.c:1017:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_SET, 0);\ndrivers/vfio/vfio_main.c-1018-\tif (ret != 1)\n--\ndrivers/vfio/vfio_main.c=1025=static int vfio_ioctl_device_feature_migration(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1036-\ndrivers/vfio/vfio_main.c:1037:\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\ndrivers/vfio/vfio_main.c-1038-\t\t\t\t sizeof(mig));\n--\ndrivers/vfio/vfio_main.c=1108=vfio_ioctl_device_feature_logging_start(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1126-\ndrivers/vfio/vfio_main.c:1127:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1128-\t\t\t\t VFIO_DEVICE_FEATURE_SET,\n--\ndrivers/vfio/vfio_main.c=1192=vfio_ioctl_device_feature_logging_stop(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1200-\ndrivers/vfio/vfio_main.c:1201:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1202-\t\t\t\t VFIO_DEVICE_FEATURE_SET, 0);\n--\ndrivers/vfio/vfio_main.c=1219=vfio_ioctl_device_feature_logging_report(struct vfio_device *device,\n--\ndrivers/vfio/vfio_main.c-1233-\ndrivers/vfio/vfio_main.c:1234:\tret = vfio_check_feature(flags, argsz,\ndrivers/vfio/vfio_main.c-1235-\t\t\t\t VFIO_DEVICE_FEATURE_GET,\n--\ninclude/linux/vfio.h=250=struct vfio_log_ops {\n--\ninclude/linux/vfio.h-259-/**\ninclude/linux/vfio.h:260: * vfio_check_feature - Validate user input for the VFIO_DEVICE_FEATURE ioctl\ninclude/linux/vfio.h-261- * @flags: Arg from the device_feature op\n--\ninclude/linux/vfio.h-271- */\ninclude/linux/vfio.h:272:static inline int vfio_check_feature(u32 flags, size_t argsz, u32 supported_ops,\ninclude/linux/vfio.h-273-\t\t\t\t size_t minsz)\n"
} |
| 39/3 | 2026/10/06 20:17 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 28452 output: 36 thoughts: 33 Thoughts: Here's my analysis of the provided text, rendered in the first person, assuming an expert user: **Diving into vfio_check_feature** Right, let's get down to brass tacks. I need to understand the nuances of `vfio_check_feature`, and the best place to start is, naturally, the source code itself. I'm going to jump into `include/linux/vfio.h` and focus specifically around line 272. That's where the heart of `vfio_check_feature` likely lies. I need to see how it's implemented. From there, I can start to understand its function and behaviour. I will start by seeing how the checks are performed and any related data structures. With that information, I can get a better handle on the potential edge cases and implications. |
| 40/3 | 2026/10/06 20:18 | tool | read-file |
0mArgs:{
"File": "include/linux/vfio.h",
"FirstLine": 265,
"LineCount": 40
}Results: {
"Contents": " 265:\t * @minsz: Minimum data size the driver accepts\n 266:\t *\n 267:\t * For use in a driver's device_feature op. Checks that the inputs to the\n 268:\t * VFIO_DEVICE_FEATURE ioctl are correct for the driver's feature. Returns 1 if\n 269:\t * the driver should execute the get or set, otherwise the relevant\n 270:\t * value should be returned.\n 271:\t */\n 272:\tstatic inline int vfio_check_feature(u32 flags, size_t argsz, u32 supported_ops,\n 273:\t\t\t\t\t size_t minsz)\n 274:\t{\n 275:\t\tif ((flags \u0026 (VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_SET)) \u0026\n 276:\t\t ~supported_ops)\n 277:\t\t\treturn -EINVAL;\n 278:\t\tif (flags \u0026 VFIO_DEVICE_FEATURE_PROBE)\n 279:\t\t\treturn 0;\n 280:\t\t/* Without PROBE one of GET or SET must be requested */\n 281:\t\tif (!(flags \u0026 (VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_SET)))\n 282:\t\t\treturn -EINVAL;\n 283:\t\tif (argsz \u003c minsz)\n 284:\t\t\treturn -EINVAL;\n 285:\t\treturn 1;\n 286:\t}\n 287:\t\n 288:\t/**\n 289:\t * vfio_check_precopy_ioctl - Validate user input for the VFIO_MIG_GET_PRECOPY_INFO ioctl\n 290:\t * @vdev: The vfio device\n 291:\t * @cmd: Cmd from the ioctl\n 292:\t * @arg: Arg from the ioctl\n 293:\t * @info: Driver pointer to hold the userspace input to the ioctl\n 294:\t *\n 295:\t * For use in a driver's get_precopy_info. Checks that the inputs to the\n 296:\t * VFIO_MIG_GET_PRECOPY_INFO ioctl are correct.\n 297:\t\n 298:\t * Returns 0 on success, otherwise errno.\n 299:\t */\n 300:\t\n 301:\tstatic inline int\n 302:\tvfio_check_precopy_ioctl(struct vfio_device *vdev, unsigned int cmd,\n 303:\t\t\t\t unsigned long arg, struct vfio_precopy_info *info)\n 304:\t{\n"
} |
| 41/3 | 2026/10/06 20:18 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 29381 output: 19 thoughts: 696 Thoughts: **Analyzing Potential Information Leakage in DMA Buffer Handling** Okay, let's dive into this `vfio_pci_core_feature_dma_buf_revoke` function. My initial impression is that it seems reasonably safe, at least from a quick glance. I see that `vfio_check_feature` handles the probing, and it's properly limited to the `SET` operation, returning appropriate errors for `PROBE` and `GET`. The code copies data *from* userspace, but crucially, it doesn't copy anything *to* userspace. So, it seems unlikely there's an immediate info leak here related to this specific function. Now, the plot thickens with the `DMA_BUF_SET_NAME` and `dma_buf_ioctl` functions. I notice that the code copies a name string from userspace using `strndup_user` and then calls `dma_buf_set_name`. The `dma_buf_set_name` function itself looks... interesting. It checks for a null pointer, enforces `DMA_BUF_NAME_LEN` length, and *crucially* updates `dmabuf->name`. I'm immediately raising a red flag. Where is this `dmabuf->name` actually *read*? Is it used in a way that could lead to an information leak? The length check prevents a large string from being written, but that still may not fully address any potential issues. I need to investigate further. Time to trace all the occurrences of `dmabuf->name` within the kernel source code. I need to understand where and how this string is used. This will determine whether any user-provided string that is set here can be read back or cause other problems, maybe a race condition. |
| 42/3 | 2026/10/06 20:18 | tool | grepper |
0mArgs:{
"Expression": "dmabuf-\u003ename"
}Results: {
"Output": "drivers/dma-buf/dma-buf.c-39-/*\ndrivers/dma-buf/dma-buf.c:40: * dmabuf-\u003ename must be accessed with holding dmabuf-\u003ename_lock.\ndrivers/dma-buf/dma-buf.c-41- * we need to take the lock around the tracepoint call itself where\n--\ndrivers/dma-buf/dma-buf.c-50-\t\tif (IS_ENABLED(CONFIG_LOCKDEP) || FUNC##_enabled()) {\t\\\ndrivers/dma-buf/dma-buf.c:51:\t\t\tguard(spinlock)(\u0026dmabuf-\u003ename_lock);\t\t\\\ndrivers/dma-buf/dma-buf.c-52-\t\t\tFUNC(__VA_ARGS__);\t\t\t\t\\\n--\ndrivers/dma-buf/dma-buf.c=150=static char *dmabuffs_dname(struct dentry *dentry, char *buffer, int buflen)\n--\ndrivers/dma-buf/dma-buf.c-156-\tdmabuf = dentry-\u003ed_fsdata;\ndrivers/dma-buf/dma-buf.c:157:\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:158:\tif (dmabuf-\u003ename)\ndrivers/dma-buf/dma-buf.c:159:\t\tret = strscpy(name, dmabuf-\u003ename, sizeof(name));\ndrivers/dma-buf/dma-buf.c:160:\tspin_unlock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c-161-\n--\ndrivers/dma-buf/dma-buf.c=166=static void dma_buf_release(struct dentry *dentry)\n--\ndrivers/dma-buf/dma-buf.c-189-\tmodule_put(dmabuf-\u003eowner);\ndrivers/dma-buf/dma-buf.c:190:\tkfree(dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c-191-\tkfree(dmabuf);\n--\ndrivers/dma-buf/dma-buf.c=578=static void dma_buf_show_fdinfo(struct seq_file *m, struct file *file)\n--\ndrivers/dma-buf/dma-buf.c-585-\tseq_printf(m, \"exp_name:\\t%s\\n\", dmabuf-\u003eexp_name);\ndrivers/dma-buf/dma-buf.c:586:\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:587:\tif (dmabuf-\u003ename)\ndrivers/dma-buf/dma-buf.c:588:\t\tseq_printf(m, \"name:\\t%s\\n\", dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c:589:\tspin_unlock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c-590-}\n--\ndrivers/dma-buf/dma-buf.c=691=struct dma_buf *dma_buf_export(const struct dma_buf_export_info *exp_info)\n--\ndrivers/dma-buf/dma-buf.c-732-\tdmabuf-\u003eowner = exp_info-\u003eowner;\ndrivers/dma-buf/dma-buf.c:733:\tspin_lock_init(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c-734-\tinit_waitqueue_head(\u0026dmabuf-\u003epoll);\n--\ndrivers/dma-buf/dma-buf.c=876=int dma_buf_set_name(struct dma_buf *dmabuf, char *name)\n--\ndrivers/dma-buf/dma-buf.c-886-\ndrivers/dma-buf/dma-buf.c:887:\tspin_lock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c:888:\tkfree(dmabuf-\u003ename);\ndrivers/dma-buf/dma-buf.c:889:\tdmabuf-\u003ename = name;\ndrivers/dma-buf/dma-buf.c:890:\tspin_unlock(\u0026dmabuf-\u003ename_lock);\ndrivers/dma-buf/dma-buf.c-891-\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c=31=int dmabuf_collector(struct bpf_iter__dmabuf *ctx)\n--\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-44-\t bpf_core_read(\u0026size, sizeof(size), \u0026dmabuf-\u003esize) ||\ntools/testing/selftests/bpf/progs/dmabuf_iter.c:45:\t bpf_core_read(\u0026pname, sizeof(pname), \u0026dmabuf-\u003ename) ||\ntools/testing/selftests/bpf/progs/dmabuf_iter.c-46-\t bpf_core_read(\u0026exporter, sizeof(exporter), \u0026dmabuf-\u003eexp_name))\n"
} |
| 43/3 | 2026/10/06 20:18 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 31832 output: 27 thoughts: 1087 Thoughts: **Analyzing DMA-BUF and VFIO Interactions** Okay, so I'm looking at how DMA buffers are named and managed within the VFIO-PCI context. My initial thought process is to ensure that the naming scheme and memory management are sound, particularly in error scenarios. First, I see that the `dmabuf->name` field is critical, used both for reporting and potentially exposed to userspace. The name itself comes from `seq_printf` or `dmabuffs_dname`. Also, user-provided names, read via `strndup_user`, have a size limit using `DMA_BUF_NAME_LEN`. Good, that prevents arbitrary allocation sizes. Now, focusing on the `vfio_pci_core_mmap_prep_dmabuf` function, I notice it allocates a `vfio_pci_dma_buf` structure and a `phys_vec` within that using `kzalloc_obj`, which is nice for initialization. The debug name, constructed using `kasprintf`, looks reasonably sized; the comment confirms it fits within `DMA_BUF_NAME_LEN`. Here's the critical bit: the code calls `dma_buf_set_name` to assign the name to the DMA buffer. If that *fails* (returns non-zero), then the code *also* calls `kfree(bufname)`. Okay, so `dma_buf_set_name` should handle memory ownership. Let me quickly verify that. Yes, looking at the source of `dma_buf_set_name`, it does indeed take ownership of the buffer name, freeing any previous name and setting the new one. On destruction of the DMA buffer (`dma_buf_release`), the `name` is freed. Excellent, memory management seems sane here. The use of `kzalloc_obj` ensures no accidental information leakage. Now, a completely different concern: `vfio_pci_dma_buf_find_pfn` and the `vfio_pci_mmap_huge_fault` function. I see a `pfn` (page frame number) is retrieved and used in `vfio_pci_vmf_insert_pfn` in the fault handler. The `pfn` is initialized to 0. If `vfio_pci_dma_buf_find_pfn` succeeds (returns 0), `pfn` is set to the correct physical address. This seems fine, assuming `vfio_pci_dma_buf_find_pfn` can always initialize the `pfn`. Then `vfio_pci_vmf_insert_pfn` is used to insert the PFN into the VMA. I now need to examine the implementation of `vfio_pci_vmf_insert_pfn` to see how it inserts the PFN and confirm no issues there. |
| 44/3 | 2026/10/06 20:19 | tool | grepper |
0mArgs:{
"Expression": "vfio_pci_vmf_insert_pfn"
}Results: {
"Output": "drivers/vfio/pci/nvgrace-gpu/main.c=392=static vm_fault_t nvgrace_gpu_vfio_pci_huge_fault(struct vm_fault *vmf,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-440-\ndrivers/vfio/pci/nvgrace-gpu/main.c:441:\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf, pfn, order);\ndrivers/vfio/pci/nvgrace-gpu/main.c-442-\t\t}\n--\ndrivers/vfio/pci/vfio_pci_core.c=1772=void vfio_pci_memory_unlock_and_restore(struct vfio_pci_core_device *vdev, u16 cmd)\n--\ndrivers/vfio/pci/vfio_pci_core.c-1777-\ndrivers/vfio/pci/vfio_pci_core.c:1778:vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\ndrivers/vfio/pci/vfio_pci_core.c-1779-\t\t\t\t struct vm_fault *vmf,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1798-}\ndrivers/vfio/pci/vfio_pci_core.c:1799:EXPORT_SYMBOL_GPL(vfio_pci_vmf_insert_pfn);\ndrivers/vfio/pci/vfio_pci_core.c-1800-\ndrivers/vfio/pci/vfio_pci_core.c=1801=static vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1880-\ndrivers/vfio/pci/vfio_pci_core.c:1881:\t/* memory_lock for vfio_pci_vmf_insert_pfn() */\ndrivers/vfio/pci/vfio_pci_core.c-1882-\tdown_read(\u0026vdev-\u003ememory_lock);\n--\ndrivers/vfio/pci/vfio_pci_core.c-1890-\t\tif (pres == 0)\ndrivers/vfio/pci/vfio_pci_core.c:1891:\t\t\tret = vfio_pci_vmf_insert_pfn(vdev, vmf,\ndrivers/vfio/pci/vfio_pci_core.c-1892-\t\t\t\t\t\t pfn, order);\n--\ninclude/linux/vfio_pci_core.h=186=ssize_t vfio_pci_core_write(struct vfio_device *core_vdev, const char __user *buf,\ninclude/linux/vfio_pci_core.h-187-\t\tsize_t count, loff_t *ppos);\ninclude/linux/vfio_pci_core.h:188:vm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\ninclude/linux/vfio_pci_core.h-189-\t\t\t\t struct vm_fault *vmf, unsigned long pfn,\n"
} |
| 45/3 | 2026/10/06 20:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 34449 output: 46 thoughts: 46 Thoughts: Okay, let's dive into this. Here's my thought process on the matter, summarized as requested: **Analyzing vfio_pci_vmf_insert_pfn in vfio_pci_core.c** Alright, let's get down to brass tacks. I need to understand what's happening at the core of this function, specifically `vfio_pci_vmf_insert_pfn`. It's located in `drivers/vfio/pci/vfio_pci_core.c`, and I'm targeting roughly line 1778. I know this is where the action is likely happening related to page frame number (PFN) insertion, which is a critical piece of the VFIO puzzle. I need to understand how VFIO is managing physical memory regions, especially when it comes to attaching them to a virtual machine (VM). I'm probably looking at the details of how the driver handles requests from the user space for mapping physical memory into the VM's address space. Let's see how they're handling potential issues related to memory access permissions, potential conflicts, and the overall integrity of the system's memory management. I suspect it's going to involve some kind of locking mechanism to ensure coherency. Hopefully the code comments are good, that'll save me some time. |
| 46/3 | 2026/10/06 20:19 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_core.c",
"FirstLine": 1775,
"LineCount": 30
}Results: {
"Contents": "1775:\t\tup_write(\u0026vdev-\u003ememory_lock);\n1776:\t}\n1777:\t\n1778:\tvm_fault_t vfio_pci_vmf_insert_pfn(struct vfio_pci_core_device *vdev,\n1779:\t\t\t\t\t struct vm_fault *vmf,\n1780:\t\t\t\t\t unsigned long pfn,\n1781:\t\t\t\t\t unsigned int order)\n1782:\t{\n1783:\t\tlockdep_assert_held_read(\u0026vdev-\u003ememory_lock);\n1784:\t\n1785:\t\tif (vdev-\u003epm_runtime_engaged || !__vfio_pci_memory_enabled(vdev))\n1786:\t\t\treturn VM_FAULT_SIGBUS;\n1787:\t\n1788:\t\tif (!order)\n1789:\t\t\treturn vmf_insert_pfn(vmf-\u003evma, vmf-\u003eaddress, pfn);\n1790:\t\n1791:\t\tif (IS_ENABLED(CONFIG_ARCH_SUPPORTS_PMD_PFNMAP) \u0026\u0026 order == PMD_ORDER)\n1792:\t\t\treturn vmf_insert_pfn_pmd(vmf, pfn, false);\n1793:\t\n1794:\t\tif (IS_ENABLED(CONFIG_ARCH_SUPPORTS_PUD_PFNMAP) \u0026\u0026 order == PUD_ORDER)\n1795:\t\t\treturn vmf_insert_pfn_pud(vmf, pfn, false);\n1796:\t\n1797:\t\treturn VM_FAULT_FALLBACK;\n1798:\t}\n1799:\tEXPORT_SYMBOL_GPL(vfio_pci_vmf_insert_pfn);\n1800:\t\n1801:\tstatic vm_fault_t vfio_pci_mmap_huge_fault(struct vm_fault *vmf,\n1802:\t\t\t\t\t\t unsigned int order)\n1803:\t{\n1804:\t\tstruct vm_area_struct *vma = vmf-\u003evma;\n"
} |
| 47/3 | 2026/10/06 20:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 35368 output: 47 thoughts: 401 Thoughts: **Analyzing VFIO PCI DMA-BUF Feature** Okay, let's dive into this `vfio_pci_core_ioctl_feature` function. I see it handles different VFIO device features, and right now I'm interested in the `VFIO_DEVICE_FEATURE_DMA_BUF` case. That's the one I need to understand. Specifically, I'll examine how it's implemented. I'm going to look at the `vfio_pci_core_feature_dma_buf` function, that seems like the key. Now, let's go to `drivers/vfio/pci/vfio_pci_dmabuf.c` and take a look at the code. Interesting, looks like there's a refactoring here. The creation of the `dmabuf` is now delegated to `vfio_pci_dmabuf_export`. Let me check that. I want to see how that function works and specifically, how it affects the `dmabuf` lifecycle. I'll read around line 640 for context. It looks like `get_dma_buf.open_flags` is an important value, I'll need to see how that is being used, or where it's being set. |
| 48/3 | 2026/10/06 20:19 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_dmabuf.c",
"FirstLine": 640,
"LineCount": 40
}Results: {
"Contents": " 640:\t\t}\n 641:\t\n 642:\t\tpriv-\u003evdev = vdev;\n 643:\t\tpriv-\u003enr_ranges = get_dma_buf.nr_ranges;\n 644:\t\tpriv-\u003esize = length;\n 645:\t\tret = vdev-\u003epci_ops-\u003eget_dmabuf_phys(vdev, \u0026priv-\u003eprovider,\n 646:\t\t\t\t\t\t get_dma_buf.region_index,\n 647:\t\t\t\t\t\t priv-\u003ephys_vec, dma_ranges,\n 648:\t\t\t\t\t\t priv-\u003enr_ranges);\n 649:\t\tif (ret)\n 650:\t\t\tgoto err_free_phys;\n 651:\t\n 652:\t\tkfree(dma_ranges);\n 653:\t\tdma_ranges = NULL;\n 654:\t\n 655:\t\tret = vfio_pci_dmabuf_export(vdev, priv, get_dma_buf.open_flags);\n 656:\t\tif (ret)\n 657:\t\t\tgoto err_free_phys;\n 658:\t\t/*\n 659:\t\t * dma_buf_fd() consumes the reference, when the file closes the dmabuf\n 660:\t\t * will be released.\n 661:\t\t */\n 662:\t\tret = dma_buf_fd(priv-\u003edmabuf, get_dma_buf.open_flags);\n 663:\t\tif (ret \u003c 0)\n 664:\t\t\tdma_buf_put(priv-\u003edmabuf);\n 665:\t\n 666:\t\treturn ret;\n 667:\t\n 668:\terr_free_phys:\n 669:\t\tkfree(priv-\u003ephys_vec);\n 670:\terr_free_priv:\n 671:\t\tkfree(priv);\n 672:\terr_free_ranges:\n 673:\t\tkfree(dma_ranges);\n 674:\t\treturn ret;\n 675:\t}\n 676:\t#endif /* CONFIG_VFIO_PCI_DMABUF */\n 677:\t\n 678:\t/*\n 679:\t * Set the DMABUF's revocation status (OK, REVOKED, DEAD): DEAD gives\n"
} |
| 49/3 | 2026/10/06 20:19 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 36698 output: 47 thoughts: 25 Thoughts: Okay, here's my attempt at summarizing that thought process, keeping in mind the expert user: **Deep Dive into `vfio_pci_core_feature_dma_buf` Implementation** Alright, let's get down to brass tacks. I need to refresh my memory on the genesis of `vfio_pci_core_feature_dma_buf`. I'm going to start by tracing the original implementation. I need to understand the initial design choices, the rationale behind those decisions, and what assumptions were baked into the code at the time. I'll be particularly interested in any trade-offs made, performance considerations, and potential pitfalls that might have been anticipated – or overlooked – during the initial development. I want to build a solid mental map of its foundational architecture before I start contemplating any modifications or potential interactions with more recent developments in the VFIO ecosystem. No shortcuts, I need a good grasp of the basics. |
| 50/3 | 2026/10/06 20:19 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_dmabuf.c",
"FirstLine": 590,
"LineCount": 55
}Results: {
"Contents": " 590:\tint vfio_pci_core_feature_dma_buf(struct vfio_pci_core_device *vdev, u32 flags,\n 591:\t\t\t\t\t struct vfio_device_feature_dma_buf __user *arg,\n 592:\t\t\t\t\t size_t argsz)\n 593:\t{\n 594:\t\tstruct vfio_device_feature_dma_buf get_dma_buf = {};\n 595:\t\tstruct vfio_region_dma_range *dma_ranges;\n 596:\t\tstruct vfio_pci_dma_buf *priv;\n 597:\t\tsize_t length;\n 598:\t\tint ret;\n 599:\t\n 600:\t\tif (!vdev-\u003epci_ops || !vdev-\u003epci_ops-\u003eget_dmabuf_phys)\n 601:\t\t\treturn -EOPNOTSUPP;\n 602:\t\n 603:\t\tret = vfio_check_feature(flags, argsz, VFIO_DEVICE_FEATURE_GET,\n 604:\t\t\t\t\t sizeof(get_dma_buf));\n 605:\t\tif (ret != 1)\n 606:\t\t\treturn ret;\n 607:\t\n 608:\t\tif (copy_from_user(\u0026get_dma_buf, arg, sizeof(get_dma_buf)))\n 609:\t\t\treturn -EFAULT;\n 610:\t\n 611:\t\tif (!get_dma_buf.nr_ranges || get_dma_buf.flags)\n 612:\t\t\treturn -EINVAL;\n 613:\t\n 614:\t\t/*\n 615:\t\t * For PCI the region_index is the BAR number like everything\n 616:\t\t * else. Check that PCI resources have been claimed for it.\n 617:\t\t */\n 618:\t\tif (get_dma_buf.region_index \u003e= VFIO_PCI_ROM_REGION_INDEX ||\n 619:\t\t IS_ERR(vfio_pci_core_get_iomap(vdev, get_dma_buf.region_index)))\n 620:\t\t\treturn -ENODEV;\n 621:\t\n 622:\t\tdma_ranges = memdup_array_user(\u0026arg-\u003edma_ranges, get_dma_buf.nr_ranges,\n 623:\t\t\t\t\t sizeof(*dma_ranges));\n 624:\t\tif (IS_ERR(dma_ranges))\n 625:\t\t\treturn PTR_ERR(dma_ranges);\n 626:\t\n 627:\t\tret = validate_dmabuf_input(\u0026get_dma_buf, dma_ranges, \u0026length);\n 628:\t\tif (ret)\n 629:\t\t\tgoto err_free_ranges;\n 630:\t\n 631:\t\tpriv = kzalloc_obj(*priv);\n 632:\t\tif (!priv) {\n 633:\t\t\tret = -ENOMEM;\n 634:\t\t\tgoto err_free_ranges;\n 635:\t\t}\n 636:\t\tpriv-\u003ephys_vec = kzalloc_objs(*priv-\u003ephys_vec, get_dma_buf.nr_ranges);\n 637:\t\tif (!priv-\u003ephys_vec) {\n 638:\t\t\tret = -ENOMEM;\n 639:\t\t\tgoto err_free_priv;\n 640:\t\t}\n 641:\t\n 642:\t\tpriv-\u003evdev = vdev;\n 643:\t\tpriv-\u003enr_ranges = get_dma_buf.nr_ranges;\n 644:\t\tpriv-\u003esize = length;\n"
} |
| 51/3 | 2026/10/06 20:19 | llm | kmsan-evaluator |
3mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 37919 output: 29 thoughts: 139 Thoughts: Okay, let's break this down. **Diving into the VFIO DMA Buffer Feature** Alright, so `vfio_pci_core_feature_dma_buf` is the function handling that `VFIO_DEVICE_FEATURE_GET` request, and it looks like it's giving me this new DMA buffer file descriptor via `dma_buf_fd`. The crucial question: is anything *copied* back to userspace directly? I need to know if there's data being sent, or if it's just handing over a descriptor. Hold on, `dma_buf_fd` returns an integer file descriptor. That’s promising, but I need to understand the data flow completely. Now, I'll go check `vfio_ioctl_device_feature` in `drivers/vfio/vfio_main.c`. That's where the magic of how `device_feature` responses are sent back to the application lives! I'll see how that `int fd` is translated into something usable in user space. |
| 52/3 | 2026/10/06 20:23 | tool | grepper |
0mArgs:{
"Expression": "ioctl_feature",
"PathPrefixes": [
"drivers/vfio/"
]
}Results: {
"Output": "drivers/vfio/cdx/main.c=84=static int vfio_cdx_bm_ctrl(struct vfio_device *core_vdev, u32 flags,\n--\ndrivers/vfio/cdx/main.c-115-\ndrivers/vfio/cdx/main.c:116:static int vfio_cdx_ioctl_feature(struct vfio_device *device, u32 flags,\ndrivers/vfio/cdx/main.c-117-\t\t\t\t void __user *arg, size_t argsz)\n--\ndrivers/vfio/cdx/main.c=291=static const struct vfio_device_ops vfio_cdx_ops = {\n--\ndrivers/vfio/cdx/main.c-298-\t.get_region_info_caps = vfio_cdx_ioctl_get_region_info,\ndrivers/vfio/cdx/main.c:299:\t.device_feature = vfio_cdx_ioctl_feature,\ndrivers/vfio/cdx/main.c-300-\t.mmap\t\t= vfio_cdx_mmap,\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1593=static const struct vfio_device_ops hisi_acc_vfio_pci_migrn_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1600-\t.get_region_info_caps = hisi_acc_vfio_ioctl_get_region,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1601:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1602-\t.read = hisi_acc_vfio_pci_read,\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c=1614=static const struct vfio_device_ops hisi_acc_vfio_pci_ops = {\n--\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1621-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c:1622:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/hisilicon/hisi_acc_vfio_pci.c-1623-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/ism/main.c=335=static const struct vfio_device_ops ism_pci_ops = {\n--\ndrivers/vfio/pci/ism/main.c-342-\t.get_region_info_caps = ism_vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/ism/main.c:343:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/ism/main.c-344-\t.read = ism_vfio_pci_read,\n--\ndrivers/vfio/pci/mlx5/main.c=1386=static const struct vfio_device_ops mlx5vf_pci_ops = {\n--\ndrivers/vfio/pci/mlx5/main.c-1393-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/mlx5/main.c:1394:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/mlx5/main.c-1395-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1065=static const struct vfio_device_ops nvgrace_gpu_pci_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1072-\t.get_region_info_caps = nvgrace_gpu_ioctl_get_region_info,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1073:\t.device_feature\t= vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1074-\t.read\t\t= nvgrace_gpu_read,\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c=1090=static const struct vfio_device_ops nvgrace_gpu_pci_core_ops = {\n--\ndrivers/vfio/pci/nvgrace-gpu/main.c-1097-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/nvgrace-gpu/main.c:1098:\t.device_feature\t= vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/nvgrace-gpu/main.c-1099-\t.read\t\t= vfio_pci_core_read,\n--\ndrivers/vfio/pci/pds/vfio_dev.c=191=static const struct vfio_device_ops pds_vfio_ops = {\n--\ndrivers/vfio/pci/pds/vfio_dev.c-198-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/pds/vfio_dev.c:199:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/pds/vfio_dev.c-200-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/vfio_pci.c=149=static const struct vfio_device_ops vfio_pci_ops = {\n--\ndrivers/vfio/pci/vfio_pci.c-156-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/vfio_pci.c:157:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/vfio_pci.c-158-\t.read\t\t= vfio_pci_core_read,\n--\ndrivers/vfio/pci/vfio_pci_core.c=1598=static int vfio_pci_core_feature_token(struct vfio_pci_core_device *vdev,\n--\ndrivers/vfio/pci/vfio_pci_core.c-1624-\ndrivers/vfio/pci/vfio_pci_core.c:1625:int vfio_pci_core_ioctl_feature(struct vfio_device *device, u32 flags,\ndrivers/vfio/pci/vfio_pci_core.c-1626-\t\t\t\tvoid __user *arg, size_t argsz)\n--\ndrivers/vfio/pci/vfio_pci_core.c-1650-}\ndrivers/vfio/pci/vfio_pci_core.c:1651:EXPORT_SYMBOL_GPL(vfio_pci_core_ioctl_feature);\ndrivers/vfio/pci/vfio_pci_core.c-1652-\n--\ndrivers/vfio/pci/virtio/main.c=84=static const struct vfio_device_ops virtiovf_vfio_pci_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-91-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:92:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-93-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/virtio/main.c=106=static const struct vfio_device_ops virtiovf_vfio_pci_tran_lm_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-113-\t.get_region_info_caps = virtiovf_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:114:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-115-\t.read = virtiovf_pci_core_read,\n--\ndrivers/vfio/pci/virtio/main.c=128=static const struct vfio_device_ops virtiovf_vfio_pci_ops = {\n--\ndrivers/vfio/pci/virtio/main.c-135-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/virtio/main.c:136:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/virtio/main.c-137-\t.read = vfio_pci_core_read,\n--\ndrivers/vfio/pci/xe/main.c=522=static const struct vfio_device_ops xe_vfio_pci_ops = {\n--\ndrivers/vfio/pci/xe/main.c-529-\t.get_region_info_caps = vfio_pci_ioctl_get_region_info,\ndrivers/vfio/pci/xe/main.c:530:\t.device_feature = vfio_pci_core_ioctl_feature,\ndrivers/vfio/pci/xe/main.c-531-\t.read = vfio_pci_core_read,\n"
} |
| 53/3 | 2026/10/06 20:23 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 40934 output: 35 thoughts: 23 Thoughts: Okay, here's my take on those thoughts, tailored for an expert audience, and written as if they were my own immediate processing: **Diving into vfio_main.c and Feature Handling** Alright, my brain's honed in on `drivers/vfio/vfio_main.c`, specifically around the `device_feature` area. That suggests we're dealing with the core logic for exposing device capabilities to the user space. I immediately start mentally parsing that section, assuming that `device_feature` probably refers to the structures or functions responsible for managing the feature bits, extended capabilities, and other device metadata that VFIO presents. I need to quickly assess how VFIO abstracts and translates those hardware features into something the user space can understand and control. Is it a direct mapping? Are there layers of abstraction? Are there any security considerations or limitations in how features are exposed? I'm particularly interested in how they're handling things like MSI-X, interrupt remapping, DMA addressing capabilities, and any vendor-specific extensions that might be present. I also need to assess whether the code is secure, so I am interested in checking how the device's features are presented to the user. I'm thinking about the implications for device assignment and isolation. Any misconfiguration or vulnerability in this area could compromise the entire VFIO system. So I'll be looking for potential race conditions, incorrect permission checks, or any logic that could allow unauthorized access to hardware resources. I'll need to check how the code interacts with the PCI configuration space, and make sure that it respects all the relevant specifications and standards. Finally, I'll be particularly interested in understanding how features are discovered, advertised, and potentially filtered or modified during the assignment process. Are there any clever tricks, or potential gotchas that I should be aware of? |
| 54/3 | 2026/10/06 20:23 | tool | grepper |
0mArgs:{
"Expression": "device_feature",
"PathPrefixes": [
"drivers/vfio/vfio_main.c"
]
}Results: {
"Output": "drivers/vfio/vfio_main.c=903=static int vfio_ioct_mig_return_fd(struct file *filp, void __user *arg,\ndrivers/vfio/vfio_main.c:904:\t\t\t\t struct vfio_device_feature_mig_state *mig)\ndrivers/vfio/vfio_main.c-905-{\n--\ndrivers/vfio/vfio_main.c=930=static int\ndrivers/vfio/vfio_main.c:931:vfio_ioctl_device_feature_mig_device_state(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-932-\t\t\t\t\t u32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c-935-\tsize_t minsz =\ndrivers/vfio/vfio_main.c:936:\t\toffsetofend(struct vfio_device_feature_mig_state, data_fd);\ndrivers/vfio/vfio_main.c:937:\tstruct vfio_device_feature_mig_state mig;\ndrivers/vfio/vfio_main.c-938-\tstruct file *filp = NULL;\n--\ndrivers/vfio/vfio_main.c=980=static int\ndrivers/vfio/vfio_main.c:981:vfio_ioctl_device_feature_migration_data_size(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-982-\t\t\t\t\t u32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c-984-{\ndrivers/vfio/vfio_main.c:985:\tstruct vfio_device_feature_mig_data_size data_size = {};\ndrivers/vfio/vfio_main.c-986-\tunsigned long stop_copy_length;\n--\ndrivers/vfio/vfio_main.c=1008=static int\ndrivers/vfio/vfio_main.c:1009:vfio_ioctl_device_feature_migration_precopy_info_v2(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1010-\t\t\t\t\t\t u32 flags, size_t argsz)\n--\ndrivers/vfio/vfio_main.c-1024-\ndrivers/vfio/vfio_main.c:1025:static int vfio_ioctl_device_feature_migration(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1026-\t\t\t\t\t u32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c-1028-{\ndrivers/vfio/vfio_main.c:1029:\tstruct vfio_device_feature_migration mig = {\ndrivers/vfio/vfio_main.c-1030-\t\t.flags = device-\u003emigration_flags,\n--\ndrivers/vfio/vfio_main.c=1101=EXPORT_SYMBOL_GPL(vfio_combine_iova_ranges);\n--\ndrivers/vfio/vfio_main.c-1104-#define LOG_MAX_RANGES \\\ndrivers/vfio/vfio_main.c:1105:\t(PAGE_SIZE / sizeof(struct vfio_device_feature_dma_logging_range))\ndrivers/vfio/vfio_main.c-1106-\ndrivers/vfio/vfio_main.c=1107=static int\ndrivers/vfio/vfio_main.c:1108:vfio_ioctl_device_feature_logging_start(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1109-\t\t\t\t\tu32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c-1112-\tsize_t minsz =\ndrivers/vfio/vfio_main.c:1113:\t\toffsetofend(struct vfio_device_feature_dma_logging_control,\ndrivers/vfio/vfio_main.c-1114-\t\t\t ranges);\ndrivers/vfio/vfio_main.c:1115:\tstruct vfio_device_feature_dma_logging_range __user *ranges;\ndrivers/vfio/vfio_main.c:1116:\tstruct vfio_device_feature_dma_logging_control control;\ndrivers/vfio/vfio_main.c:1117:\tstruct vfio_device_feature_dma_logging_range range;\ndrivers/vfio/vfio_main.c-1118-\tstruct rb_root_cached root = RB_ROOT_CACHED;\n--\ndrivers/vfio/vfio_main.c=1191=static int\ndrivers/vfio/vfio_main.c:1192:vfio_ioctl_device_feature_logging_stop(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1193-\t\t\t\t u32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c=1218=static int\ndrivers/vfio/vfio_main.c:1219:vfio_ioctl_device_feature_logging_report(struct vfio_device *device,\ndrivers/vfio/vfio_main.c-1220-\t\t\t\t\t u32 flags, void __user *arg,\n--\ndrivers/vfio/vfio_main.c-1223-\tsize_t minsz =\ndrivers/vfio/vfio_main.c:1224:\t\toffsetofend(struct vfio_device_feature_dma_logging_report,\ndrivers/vfio/vfio_main.c-1225-\t\t\t bitmap);\ndrivers/vfio/vfio_main.c:1226:\tstruct vfio_device_feature_dma_logging_report report;\ndrivers/vfio/vfio_main.c-1227-\tstruct iova_bitmap *iter;\n--\ndrivers/vfio/vfio_main.c-1262-\ndrivers/vfio/vfio_main.c:1263:static int vfio_ioctl_device_feature(struct vfio_device *device,\ndrivers/vfio/vfio_main.c:1264:\t\t\t\t struct vfio_device_feature __user *arg)\ndrivers/vfio/vfio_main.c-1265-{\ndrivers/vfio/vfio_main.c:1266:\tsize_t minsz = offsetofend(struct vfio_device_feature, flags);\ndrivers/vfio/vfio_main.c:1267:\tstruct vfio_device_feature feature;\ndrivers/vfio/vfio_main.c-1268-\n--\ndrivers/vfio/vfio_main.c-1288-\tcase VFIO_DEVICE_FEATURE_MIGRATION:\ndrivers/vfio/vfio_main.c:1289:\t\treturn vfio_ioctl_device_feature_migration(\ndrivers/vfio/vfio_main.c-1290-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1292-\tcase VFIO_DEVICE_FEATURE_MIG_DEVICE_STATE:\ndrivers/vfio/vfio_main.c:1293:\t\treturn vfio_ioctl_device_feature_mig_device_state(\ndrivers/vfio/vfio_main.c-1294-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1296-\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_START:\ndrivers/vfio/vfio_main.c:1297:\t\treturn vfio_ioctl_device_feature_logging_start(\ndrivers/vfio/vfio_main.c-1298-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1300-\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_STOP:\ndrivers/vfio/vfio_main.c:1301:\t\treturn vfio_ioctl_device_feature_logging_stop(\ndrivers/vfio/vfio_main.c-1302-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1304-\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_REPORT:\ndrivers/vfio/vfio_main.c:1305:\t\treturn vfio_ioctl_device_feature_logging_report(\ndrivers/vfio/vfio_main.c-1306-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1308-\tcase VFIO_DEVICE_FEATURE_MIG_DATA_SIZE:\ndrivers/vfio/vfio_main.c:1309:\t\treturn vfio_ioctl_device_feature_migration_data_size(\ndrivers/vfio/vfio_main.c-1310-\t\t\tdevice, feature.flags, arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c-1312-\tcase VFIO_DEVICE_FEATURE_MIG_PRECOPY_INFOv2:\ndrivers/vfio/vfio_main.c:1313:\t\treturn vfio_ioctl_device_feature_migration_precopy_info_v2(\ndrivers/vfio/vfio_main.c-1314-\t\t\tdevice, feature.flags, feature.argsz - minsz);\ndrivers/vfio/vfio_main.c-1315-\tdefault:\ndrivers/vfio/vfio_main.c:1316:\t\tif (unlikely(!device-\u003eops-\u003edevice_feature))\ndrivers/vfio/vfio_main.c-1317-\t\t\treturn -ENOTTY;\ndrivers/vfio/vfio_main.c:1318:\t\treturn device-\u003eops-\u003edevice_feature(device, feature.flags,\ndrivers/vfio/vfio_main.c-1319-\t\t\t\t\t\t arg-\u003edata,\n--\ndrivers/vfio/vfio_main.c=1369=static long vfio_device_fops_unl_ioctl(struct file *filep,\n--\ndrivers/vfio/vfio_main.c-1402-\tcase VFIO_DEVICE_FEATURE:\ndrivers/vfio/vfio_main.c:1403:\t\tret = vfio_ioctl_device_feature(device, uptr);\ndrivers/vfio/vfio_main.c-1404-\t\tbreak;\n"
} |
| 55/3 | 2026/10/06 20:23 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 44055 output: 40 thoughts: 22 Thoughts: Here's a summary of my thought process, keeping in mind my expertise: **Diving into `vfio_ioctl_device_feature`** Okay, so we're looking at some VFIO interaction here, and I need to zero in on a specific area. My internal map of the VFIO code is pretty solid, but a quick refresher never hurts. Let's start with `vfio_ioctl_device_feature`. That function is key for device feature negotiation, and I know it's probably around line 1263 based on my previous analysis. That section is the place to start drilling down. Specifically, I'm expecting to find the handling of capabilities, vendor-specific features, or maybe even some initial sanity checks on feature requests in that area. Grabbing my mental debugger, I'm mentally checking what conditions lead to a call and what the general execution flow looks like. This initial focus will help guide my next steps: whether to look at the surrounding code for context or immediately follow a particular branch within the function itself. I suspect this is going to be the gateway into understanding some more complex device behavior. |
| 56/3 | 2026/10/06 20:23 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/vfio_main.c",
"FirstLine": 1263,
"LineCount": 40
}Results: {
"Contents": "1263:\tstatic int vfio_ioctl_device_feature(struct vfio_device *device,\n1264:\t\t\t\t\t struct vfio_device_feature __user *arg)\n1265:\t{\n1266:\t\tsize_t minsz = offsetofend(struct vfio_device_feature, flags);\n1267:\t\tstruct vfio_device_feature feature;\n1268:\t\n1269:\t\tif (copy_from_user(\u0026feature, arg, minsz))\n1270:\t\t\treturn -EFAULT;\n1271:\t\n1272:\t\tif (feature.argsz \u003c minsz)\n1273:\t\t\treturn -EINVAL;\n1274:\t\n1275:\t\t/* Check unknown flags */\n1276:\t\tif (feature.flags \u0026\n1277:\t\t ~(VFIO_DEVICE_FEATURE_MASK | VFIO_DEVICE_FEATURE_SET |\n1278:\t\t VFIO_DEVICE_FEATURE_GET | VFIO_DEVICE_FEATURE_PROBE))\n1279:\t\t\treturn -EINVAL;\n1280:\t\n1281:\t\t/* GET \u0026 SET are mutually exclusive except with PROBE */\n1282:\t\tif (!(feature.flags \u0026 VFIO_DEVICE_FEATURE_PROBE) \u0026\u0026\n1283:\t\t (feature.flags \u0026 VFIO_DEVICE_FEATURE_SET) \u0026\u0026\n1284:\t\t (feature.flags \u0026 VFIO_DEVICE_FEATURE_GET))\n1285:\t\t\treturn -EINVAL;\n1286:\t\n1287:\t\tswitch (feature.flags \u0026 VFIO_DEVICE_FEATURE_MASK) {\n1288:\t\tcase VFIO_DEVICE_FEATURE_MIGRATION:\n1289:\t\t\treturn vfio_ioctl_device_feature_migration(\n1290:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1291:\t\t\t\tfeature.argsz - minsz);\n1292:\t\tcase VFIO_DEVICE_FEATURE_MIG_DEVICE_STATE:\n1293:\t\t\treturn vfio_ioctl_device_feature_mig_device_state(\n1294:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1295:\t\t\t\tfeature.argsz - minsz);\n1296:\t\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_START:\n1297:\t\t\treturn vfio_ioctl_device_feature_logging_start(\n1298:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n1299:\t\t\t\tfeature.argsz - minsz);\n1300:\t\tcase VFIO_DEVICE_FEATURE_DMA_LOGGING_STOP:\n1301:\t\t\treturn vfio_ioctl_device_feature_logging_stop(\n1302:\t\t\t\tdevice, feature.flags, arg-\u003edata,\n"
} |
| 57/3 | 2026/10/06 20:23 | llm | kmsan-evaluator |
1mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 45058 output: 46 thoughts: 544 Thoughts: **Analyzing the DMA Buffer Revocation Flow in VFIO-PCI** Okay, so I'm diving into this `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` functionality. It's routed through `device->ops->device_feature`, which resolves to `vfio_pci_core_ioctl_feature`. That makes sense. Now, let's look at `vfio_pci_core_feature_dma_buf_revoke`. It seems like this function is the core of the revocation process. It takes a `vfio_pci_core_device` struct, some flags, a user-space argument (`arg`), and its size (`argsz`). First, there's a check to see if the device has a `pci_ops` with a `get_dmabuf_phys` method, which is reasonable. Then it uses `vfio_check_feature` to validate some things about the feature. If that passes, it copies data from user space into a local `vfio_device_feature_dma_buf_revoke` struct, and grabs a reference to the `dma_buf` associated with the file descriptor, `db_revoke.dmabuf_fd`. If the `dma_buf` is valid, it proceeds. The function returns a zero or an error code, and, importantly, doesn't copy any data back to user space. Now, I'm thinking about the bigger picture. I need to understand how this revocation fits into the DMA buffer mmap operations in the vfio-pci driver. I need to see if this is tied into the mmap implementation. Specifically, I'm thinking about `vfio_pci_core_mmap` and, in relation, `vfio_pci_core_mmap_prep_dmabuf`, and how DMA buffers are mapped using those functions. Prior to this patch, `vfio_pci_core_mmap` did the memory mapping using either `remap_pfn_range` or setting `vma->vm_ops = &vfio_pci_mmap_ops`. I know there was a fault handler called `vfio_pci_mmap_huge_fault` at one point. But hold on a second... how *was* `vfio_pci_core_mmap` implemented before this patch? I need to check the source code in `drivers/vfio/pci/vfio_pci_core.c` to see the historical context. That will give me the full picture. |
| 58/3 | 2026/10/06 20:24 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_core.c",
"FirstLine": 1920,
"LineCount": 70
}Results: {
"Contents": "1920:\tvoid vfio_pci_set_vma_ops(struct vm_area_struct *vma)\n1921:\t{\n1922:\t\tvma-\u003evm_ops = \u0026vfio_pci_mmap_ops;\n1923:\t}\n1924:\t\n1925:\tint vfio_pci_core_mmap(struct vfio_device *core_vdev, struct vm_area_struct *vma)\n1926:\t{\n1927:\t\tstruct vfio_pci_core_device *vdev =\n1928:\t\t\tcontainer_of(core_vdev, struct vfio_pci_core_device, vdev);\n1929:\t\tstruct pci_dev *pdev = vdev-\u003epdev;\n1930:\t\tunsigned int index;\n1931:\t\tu64 phys_len, req_len, pgoff, req_start;\n1932:\t\tvoid __iomem *bar_io;\n1933:\t\tint ret;\n1934:\t\n1935:\t\tindex = vma-\u003evm_pgoff \u003e\u003e (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT);\n1936:\t\n1937:\t\tif (index \u003e= VFIO_PCI_NUM_REGIONS + vdev-\u003enum_regions)\n1938:\t\t\treturn -EINVAL;\n1939:\t\tif (vma-\u003evm_end \u003c vma-\u003evm_start)\n1940:\t\t\treturn -EINVAL;\n1941:\t\tif ((vma-\u003evm_flags \u0026 VM_SHARED) == 0)\n1942:\t\t\treturn -EINVAL;\n1943:\t\tif (index \u003e= VFIO_PCI_NUM_REGIONS) {\n1944:\t\t\tint regnum = index - VFIO_PCI_NUM_REGIONS;\n1945:\t\t\tstruct vfio_pci_region *region = vdev-\u003eregion + regnum;\n1946:\t\n1947:\t\t\tif (region-\u003eops \u0026\u0026 region-\u003eops-\u003emmap \u0026\u0026\n1948:\t\t\t (region-\u003eflags \u0026 VFIO_REGION_INFO_FLAG_MMAP))\n1949:\t\t\t\treturn region-\u003eops-\u003emmap(vdev, region, vma);\n1950:\t\t\treturn -EINVAL;\n1951:\t\t}\n1952:\t\tif (index \u003e= VFIO_PCI_ROM_REGION_INDEX)\n1953:\t\t\treturn -EINVAL;\n1954:\t\tif (!vdev-\u003ebar_mmap_supported[index])\n1955:\t\t\treturn -EINVAL;\n1956:\t\n1957:\t\tphys_len = PAGE_ALIGN(pci_resource_len(pdev, index));\n1958:\t\treq_len = vma-\u003evm_end - vma-\u003evm_start;\n1959:\t\tpgoff = vma-\u003evm_pgoff \u0026\n1960:\t\t\t((1U \u003c\u003c (VFIO_PCI_OFFSET_SHIFT - PAGE_SHIFT)) - 1);\n1961:\t\treq_start = pgoff \u003c\u003c PAGE_SHIFT;\n1962:\t\n1963:\t\tif (req_start + req_len \u003e phys_len)\n1964:\t\t\treturn -EINVAL;\n1965:\t\n1966:\t\t/*\n1967:\t\t * Ensure the BAR resource region is reserved for use.\n1968:\t\t */\n1969:\t\tbar_io = vfio_pci_core_get_iomap(vdev, index);\n1970:\t\tif (IS_ERR(bar_io))\n1971:\t\t\treturn PTR_ERR(bar_io);\n1972:\t\n1973:\t\tret = vfio_pci_core_mmap_prep_dmabuf(vdev, vma,\n1974:\t\t\t\t\t\t pci_resource_start(pdev, index),\n1975:\t\t\t\t\t\t req_len, index);\n1976:\t\tif (ret)\n1977:\t\t\treturn ret;\n1978:\t\n1979:\t\tvma-\u003evm_page_prot = pgprot_noncached(vma-\u003evm_page_prot);\n1980:\t\tvma-\u003evm_page_prot = pgprot_decrypted(vma-\u003evm_page_prot);\n1981:\t\n1982:\t\t/*\n1983:\t\t * Set vm_flags now, they should not be changed in the fault handler.\n1984:\t\t * We want the same flags and page protection (decrypted above) as\n1985:\t\t * io_remap_pfn_range() would set.\n1986:\t\t *\n1987:\t\t * VM_ALLOW_ANY_UNCACHED: The VMA flag is implemented for ARM64,\n1988:\t\t * allowing KVM stage 2 device mapping attributes to use Normal-NC\n1989:\t\t * rather than DEVICE_nGnRE, which allows guest mappings\n"
} |
| 59/3 | 2026/10/06 20:24 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 47467 output: 46 thoughts: 206 Thoughts: Here's my understanding of the change, broken down in my own head: **VFIO PCI MMAP Optimization - Stepping Through the Logic** Okay, so I'm looking at this `vfio_pci_core_mmap` function. Before this patch was applied, it was straightforward: the `vma->vm_private_data` was directly set to the `vdev`. Now, things are different. Instead of that direct assignment, we're calling `vfio_pci_core_mmap_prep_dmabuf(vdev, vma, pci_resource_start(pdev, index), req_len, index)`. My understanding is that this is introducing some sort of DMA buffer management. This `vfio_pci_core_mmap_prep_dmabuf` function does a few critical things. It's setting `priv->vfile = vma->vm_file;` which looks like it's saving the original file associated with the VMA. Then, and this is the core of it, it's changing `vma->vm_file = priv->dmabuf->file;`. So the file associated with the VMA is being changed to something related to a DMA buffer, likely a file descriptor representing the DMA buffer itself. Finally, we now have `vma->vm_private_data = priv;`, which is now set to `priv` instead of `vdev`. `priv` is most likely a private structure related to the DMA buffer management. The last thing `vfio_pci_core_mmap` sets is `vma->vm_ops = &vfio_pci_mmap_ops;`. This is interesting, because it looks like the memory operations are also being replaced. I'll need to go look at `vfio_pci_mmap_ops` to understand what this new mmap operations struct does. I should check line 1989 in `drivers/vfio/pci/vfio_pci_core.c` to understand where this new code is and how these components play together. |
| 60/3 | 2026/10/06 20:25 | tool | read-file |
0mArgs:{
"File": "drivers/vfio/pci/vfio_pci_core.c",
"FirstLine": 1985,
"LineCount": 25
}Results: {
"Contents": "1985:\t\t * io_remap_pfn_range() would set.\n1986:\t\t *\n1987:\t\t * VM_ALLOW_ANY_UNCACHED: The VMA flag is implemented for ARM64,\n1988:\t\t * allowing KVM stage 2 device mapping attributes to use Normal-NC\n1989:\t\t * rather than DEVICE_nGnRE, which allows guest mappings\n1990:\t\t * supporting write-combining attributes (WC). ARM does not\n1991:\t\t * architecturally guarantee this is safe, and indeed some MMIO\n1992:\t\t * regions like the GICv2 VCPU interface can trigger uncontained\n1993:\t\t * faults if Normal-NC is used.\n1994:\t\t *\n1995:\t\t * To safely use VFIO in KVM the platform must guarantee full\n1996:\t\t * safety in the guest where no action taken against a MMIO\n1997:\t\t * mapping can trigger an uncontained failure. The assumption is\n1998:\t\t * that most VFIO PCI platforms support this for both mapping types,\n1999:\t\t * at least in common flows, based on some expectations of how\n2000:\t\t * PCI IP is integrated. Hence VM_ALLOW_ANY_UNCACHED is set in\n2001:\t\t * the VMA flags.\n2002:\t\t */\n2003:\t\tvm_flags_set(vma, VM_ALLOW_ANY_UNCACHED | VM_IO | VM_PFNMAP |\n2004:\t\t\t\tVM_DONTEXPAND | VM_DONTDUMP);\n2005:\t\tvma-\u003evm_ops = \u0026vfio_pci_mmap_ops;\n2006:\t\n2007:\t\treturn 0;\n2008:\t}\n2009:\tEXPORT_SYMBOL_GPL(vfio_pci_core_mmap);\n"
} |
| 61/3 | 2026/10/06 20:25 | llm | kmsan-evaluator |
0mModel:gemini-3.8-flash Args: null Results: null Tokens: input: 48658 output: 470 thoughts: 2883 Thoughts: **Analysis of DMA-BUF and VFIO-PCI Integration for KMSAN Needs** Okay, let's break down this code and assess whether KMSAN (Kernel Memory Sanitizer) is needed. My immediate focus is on whether this patch introduces any potential for uninitialized memory reads or information leakage to userspace, which is the primary trigger for `NeedsKMSAN=true`. I'm looking at the `vfio_pci_dma_buf_find_pfn` function and the surrounding code, especially how the DMA buffer functionality is integrated with VFIO-PCI. The core of this seems to be about mapping device BARs using dma-bufs. I'm scrutinizing data flow, looking for any instances where uninitialized memory might be accidentally exposed through user-space interactions. I need to confirm if kernel structures are being sent to userspace, paying attention to cases where fields or padding might not be fully initialized. Also, I need to check for conditional logic that depends on potentially uninitialized variables or struct fields, or if there's allocation and initialization of complex data structures that might leave uninitialized fields accessible. Finally, I'm checking for bounds checks, lengths, or logic that could lead to access of uninitialized bytes within existing buffers. Analyzing the code, I see no explicit information leakage to userspace. Functions like `dma_buf_set_name` and other user-facing ioctls seem to be carefully handling data coming from user space, and no output is shared to user space. The code appears to meticulously initialize all memory using `kzalloc_obj`, and `init_rwsem`, and all struct members are explicitly assigned values. The `vfio_pci_dma_buf_find_pfn` function itself seems to have robust checks, ensuring that address calculations are safe and that the PFNs are validated appropriately. The `check_add_overflow` usage caught my attention; these should have prevented potential vulnerabilities from integer overflows. Given the absence of any apparent uninitialized memory reads or potential data leaks, this patch seems to primarily introduce concurrency/lifetime concerns. I have checked for synchronization issues (locking), use-after-free conditions, potential address calculation/out-of-bounds errors. These are the kinds of issues that should be picked up by existing debugging tools like KASAN and LOCKDEP. These tools are far more appropriate for these issues. Therefore, because this patch doesn't introduce a risk of uninitialized memory reads or information leaks to userspace, it does not warrant a KMSAN fuzzing session. The `NeedsKMSAN=false` flag is appropriate here. |
| 62/3 | 2026/10/06 20:25 | tool | set-results |
0mArgs:{
"KMSANReasoning": "The patch refactors dma-buf naming by exporting `dma_buf_set_name()`, transitions VFIO PCI BAR mmaps to be backed by dma-buf allocations, and adds a `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl feature along with associated locking (`dmabuf_lock`) and fault-handling logic (`vfio_pci_dma_buf_find_pfn`).\n\nAnalysis of uninitialized memory vs other sanitizer applicability:\n1. No info leaks to userspace: The new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl only consumes input (`copy_from_user` for SET) and copies nothing to userspace. The `dma_buf_set_name` ioctl path reads from userspace via `strndup_user` and outputs no data to userspace.\n2. Memory allocations: All dynamically allocated structures (`priv`, `phys_vec`) are allocated using `kzalloc_obj`, zeroing all fields. String formatting via `kasprintf` ensures proper null-termination.\n3. Fault handling and address calculations: `vfio_pci_dma_buf_find_pfn()` uses overflow checks (`check_add_overflow`) and strictly writes `*out_pfn` on success, while the caller initializes `pfn = 0`. Mapped PFNs correspond to device MMIO BAR physical addresses, not kernel RAM.\n4. Bug profile: The changes primarily involve complex concurrency management, locking hierarchy (`memory_lock`, `dmabuf_lock`, and reservation locks), object lifetimes (`dma_buf`, `vfio_pci_core_device`, `vfile`), and refcounting (`kref`, `vfio_device_registration`). These issues (use-after-free, deadlocks, out-of-bounds accesses) are detected by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory risks or kernel information leaks that would exclusively be detected by KMSAN, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
}Results: {
"KMSANReasoning": "The patch refactors dma-buf naming by exporting `dma_buf_set_name()`, transitions VFIO PCI BAR mmaps to be backed by dma-buf allocations, and adds a `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl feature along with associated locking (`dmabuf_lock`) and fault-handling logic (`vfio_pci_dma_buf_find_pfn`).\n\nAnalysis of uninitialized memory vs other sanitizer applicability:\n1. No info leaks to userspace: The new `VFIO_DEVICE_FEATURE_DMA_BUF_REVOKE` ioctl only consumes input (`copy_from_user` for SET) and copies nothing to userspace. The `dma_buf_set_name` ioctl path reads from userspace via `strndup_user` and outputs no data to userspace.\n2. Memory allocations: All dynamically allocated structures (`priv`, `phys_vec`) are allocated using `kzalloc_obj`, zeroing all fields. String formatting via `kasprintf` ensures proper null-termination.\n3. Fault handling and address calculations: `vfio_pci_dma_buf_find_pfn()` uses overflow checks (`check_add_overflow`) and strictly writes `*out_pfn` on success, while the caller initializes `pfn = 0`. Mapped PFNs correspond to device MMIO BAR physical addresses, not kernel RAM.\n4. Bug profile: The changes primarily involve complex concurrency management, locking hierarchy (`memory_lock`, `dmabuf_lock`, and reservation locks), object lifetimes (`dma_buf`, `vfio_pci_core_device`, `vfile`), and refcounting (`kref`, `vfio_device_registration`). These issues (use-after-free, deadlocks, out-of-bounds accesses) are detected by standard KASAN and LOCKDEP configurations.\n\nBecause there are no uninitialized memory risks or kernel information leaks that would exclusively be detected by KMSAN, a dedicated KMSAN fuzzing session is not warranted.",
"NeedsKMSAN": false
} |
| Total Calls | Total Tokens | Avg Tokens | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|
| Total Calls | Total Duration (Seconds) | Avg Duration (Seconds) |
|---|