AI Trajectory


Trajectory: (Expand/Collapse All)
Seq Timestamp Type Name Duration
0/0 2026/09/25 22:21 flow patch-triage
1m Args:
null

Results:
{
  "EnableConfigs": [
    "KVM",
    "KVM_INTEL",
    "KVM_AMD",
    "VIRTUALIZATION"
  ],
  "FocusSymbols": [
    "kvm_vcpu_pre_fault_memory",
    "kvm_arch_pre_fault_allowed",
    "kvm_arch_vcpu_pre_fault_memory"
  ],
  "KMSANReasoning": "The patch series adds support for KVM_PRE_FAULT_MEMORY on arm64 and refactors the generic ioctl path (virt/kvm/kvm_main.c) to invoke a new architecture hook kvm_arch_pre_fault_allowed() before vcpu_load(). On the target architecture (amd64), the change simply moves the existing pre-fault allowance check (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed) from kvm_arch_vcpu_pre_fault_memory() to kvm_arch_pre_fault_allowed().\n\nNo uninitialized memory hazards are introduced:\n1. No uninitialized memory is returned to userspace: KVM_PRE_FAULT_MEMORY reads struct kvm_pre_fault_memory fully from userspace via copy_from_user() before mutating its fields and copying it back.\n2. In the arm64 implementation, all local structs (such as kvm_s2_fault_desc and kvm_s2_fault_result) are fully initialized before use.\n3. No conditional branches rely on uninitialized values, and no padding leaks or uninitialized page/heap reads are exposed.\n\nStandard bug detectors and KASAN are fully sufficient to test these changes. A dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false,
  "Reasoning": "While the patch primarily implements KVM_PRE_FAULT_MEMORY for arm64, it also refactors generic KVM (virt/kvm/kvm_main.c) and x86 KVM (arch/x86/kvm/mmu/mmu.c). It introduces the kvm_arch_pre_fault_allowed() hook, altering the control flow of the KVM_PRE_FAULT_MEMORY vcpu ioctl on amd64 to validate prefault permissions early prior to calling vcpu_load() and taking the SRCU lock. Because reachable executable code in the KVM subsystem is modified on amd64, it warrants fuzzing.",
  "WorthFuzzing": true
}

1/1 2026/09/25 22:21 action read-patch-diff
0m Args:
null

Results:
{
  "PatchDiff": "commit d79cc5272139a53b1da70940b59fa09bf4f16b16\nAuthor: syz-cluster \u003ctriage@syzkaller.com\u003e\nDate:   Fri Sep 25 22:21:57 2026 +0000\n\n    syz-cluster: applied patch under review\n\ndiff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst\nindex 1a90598901c57..71b642148b7f6 100644\n--- a/Documentation/virt/kvm/api.rst\n+++ b/Documentation/virt/kvm/api.rst\n@@ -6489,7 +6489,7 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.\n ---------------------------\n \n :Capability: KVM_CAP_PRE_FAULT_MEMORY\n-:Architectures: none\n+:Architectures: x86, s390, arm64\n :Type: vcpu ioctl\n :Parameters: struct kvm_pre_fault_memory (in/out)\n :Returns: 0 if at least one page is processed, \u003c 0 on error\n@@ -6497,12 +6497,15 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.\n Errors:\n \n   ========== ===============================================================\n+  EAGAIN     A race occurred before progress was made, but a retry may succeed.\n   EINVAL     The specified `gpa` and `size` were invalid (e.g. not\n              page aligned, causes an overflow, or size is zero), or the VM\n              is UCONTROL (s390).\n   ENOENT     The specified `gpa` is outside defined memslots.\n+  ENOEXEC    The vCPU has not been initialised (arm64).\n   EINTR      An unmasked signal is pending and no page was processed.\n   EFAULT     The parameter address was invalid.\n+  EHWPOISON  A poisoned host page was encountered.\n   EOPNOTSUPP Mapping memory for a GPA is unsupported by the\n              hypervisor, and/or for the current vCPU state/mode.\n   EIO        unexpected error conditions (also causes a WARN)\n@@ -6522,7 +6525,17 @@ Errors:\n KVM_PRE_FAULT_MEMORY populates KVM's stage-2 page tables used to map memory\n for the current vCPU state.  KVM maps memory as if the vCPU generated a\n stage-2 read page fault, e.g. faults in memory as needed, but doesn't break\n-CoW.  On x86, KVM does not mark any newly created stage-2 PTE as Accessed.\n+CoW.  On arm64, KVM marks newly created stage-2 PTEs as Accessed, as it\n+does for any stage-2 fault, but leaves the Accessed state of existing PTEs\n+unchanged.  On x86, KVM does not mark any newly created stage-2 PTE as\n+Accessed, and for s390 it is not applicable.\n+\n+On arm64, a GPA is interpreted as an IPA, and never interpreted as the IPA\n+of a nested guest. Pre-faulting only populates canonical stage-2 page\n+tables.\n+\n+The feature is not supported on arm64 if the protected KVM (pKVM) feature\n+is enabled.\n \n In the case of confidential VM types where there is an initial set up of\n private guest memory before the guest is 'finalized'/measured, this ioctl\n@@ -6537,7 +6550,7 @@ When the ioctl returns, the input values are updated to point to the\n remaining range.  If `size` \u003e 0 on return, the caller can just issue\n the ioctl again with the same `struct kvm_map_memory` argument.\n \n-Shadow page tables cannot support this ioctl because they\n+On x86, shadow page tables cannot support this ioctl because they\n are indexed by virtual address or nested guest physical address.\n Calling this ioctl when the guest is using shadow page tables (for\n example because it is running a nested guest with nested page tables)\ndiff --git a/arch/arm64/include/asm/esr.h b/arch/arm64/include/asm/esr.h\nindex f816f5d77f1a5..86756fc9bb5a2 100644\n--- a/arch/arm64/include/asm/esr.h\n+++ b/arch/arm64/include/asm/esr.h\n@@ -437,6 +437,35 @@\n #ifndef __ASSEMBLER__\n #include \u003casm/types.h\u003e\n \n+static __always_inline bool esr_trap_is_iabt(unsigned long esr)\n+{\n+\treturn ESR_ELx_EC(esr) == ESR_ELx_EC_IABT_LOW;\n+}\n+\n+/* Always check for S1PTW *before* using this. */\n+static __always_inline bool esr_dabt_is_write(unsigned long esr)\n+{\n+\treturn esr \u0026 ESR_ELx_WNR;\n+}\n+\n+static __always_inline bool esr_dabt_is_cm(unsigned long esr)\n+{\n+\treturn esr \u0026 ESR_ELx_CM;\n+}\n+\n+static __always_inline bool esr_abt_is_sea(unsigned long esr)\n+{\n+\tswitch (esr \u0026 ESR_ELx_FSC) {\n+\tcase ESR_ELx_FSC_EXTABT:\n+\tcase ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):\n+\tcase ESR_ELx_FSC_SECC:\n+\tcase ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):\n+\t\treturn true;\n+\tdefault:\n+\t\treturn false;\n+\t}\n+}\n+\n static inline unsigned long esr_brk_comment(unsigned long esr)\n {\n \treturn esr \u0026 ESR_ELx_BRK64_ISS_COMMENT_MASK;\ndiff --git a/arch/arm64/include/asm/kvm_emulate.h b/arch/arm64/include/asm/kvm_emulate.h\nindex a3c1928bdf743..708a4b6d88230 100644\n--- a/arch/arm64/include/asm/kvm_emulate.h\n+++ b/arch/arm64/include/asm/kvm_emulate.h\n@@ -336,6 +336,16 @@ static __always_inline u64 kvm_vcpu_get_esr(const struct kvm_vcpu *vcpu)\n \treturn vcpu-\u003earch.fault.esr_el2;\n }\n \n+static __always_inline bool esr_abt_is_s1ptw(unsigned long esr)\n+{\n+\treturn esr \u0026 ESR_ELx_S1PTW;\n+}\n+\n+static __always_inline bool esr_abt_is_exec_fault(unsigned long esr)\n+{\n+\treturn esr_trap_is_iabt(esr) \u0026\u0026 !esr_abt_is_s1ptw(esr);\n+}\n+\n static inline bool guest_hyp_wfx_traps_enabled(const struct kvm_vcpu *vcpu)\n {\n \tu64 esr = kvm_vcpu_get_esr(vcpu);\n@@ -411,18 +421,13 @@ static __always_inline int kvm_vcpu_dabt_get_rd(const struct kvm_vcpu *vcpu)\n \n static __always_inline bool kvm_vcpu_abt_iss1tw(const struct kvm_vcpu *vcpu)\n {\n-\treturn !!(kvm_vcpu_get_esr(vcpu) \u0026 ESR_ELx_S1PTW);\n+\treturn esr_abt_is_s1ptw(kvm_vcpu_get_esr(vcpu));\n }\n \n /* Always check for S1PTW *before* using this. */\n static __always_inline bool kvm_vcpu_dabt_iswrite(const struct kvm_vcpu *vcpu)\n {\n-\treturn kvm_vcpu_get_esr(vcpu) \u0026 ESR_ELx_WNR;\n-}\n-\n-static inline bool kvm_vcpu_dabt_is_cm(const struct kvm_vcpu *vcpu)\n-{\n-\treturn !!(kvm_vcpu_get_esr(vcpu) \u0026 ESR_ELx_CM);\n+\treturn esr_dabt_is_write(kvm_vcpu_get_esr(vcpu));\n }\n \n static __always_inline unsigned int kvm_vcpu_dabt_get_as(const struct kvm_vcpu *vcpu)\n@@ -443,12 +448,7 @@ static __always_inline u8 kvm_vcpu_trap_get_class(const struct kvm_vcpu *vcpu)\n \n static inline bool kvm_vcpu_trap_is_iabt(const struct kvm_vcpu *vcpu)\n {\n-\treturn kvm_vcpu_trap_get_class(vcpu) == ESR_ELx_EC_IABT_LOW;\n-}\n-\n-static inline bool kvm_vcpu_trap_is_exec_fault(const struct kvm_vcpu *vcpu)\n-{\n-\treturn kvm_vcpu_trap_is_iabt(vcpu) \u0026\u0026 !kvm_vcpu_abt_iss1tw(vcpu);\n+\treturn esr_trap_is_iabt(kvm_vcpu_get_esr(vcpu));\n }\n \n static __always_inline u8 kvm_vcpu_trap_get_fault(const struct kvm_vcpu *vcpu)\n@@ -468,26 +468,9 @@ bool kvm_vcpu_trap_is_translation_fault(const struct kvm_vcpu *vcpu)\n \treturn esr_fsc_is_translation_fault(kvm_vcpu_get_esr(vcpu));\n }\n \n-static inline\n-u64 kvm_vcpu_trap_get_perm_fault_granule(const struct kvm_vcpu *vcpu)\n-{\n-\tunsigned long esr = kvm_vcpu_get_esr(vcpu);\n-\n-\tBUG_ON(!esr_fsc_is_permission_fault(esr));\n-\treturn BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(esr \u0026 ESR_ELx_FSC_LEVEL));\n-}\n-\n static __always_inline bool kvm_vcpu_abt_issea(const struct kvm_vcpu *vcpu)\n {\n-\tswitch (kvm_vcpu_trap_get_fault(vcpu)) {\n-\tcase ESR_ELx_FSC_EXTABT:\n-\tcase ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):\n-\tcase ESR_ELx_FSC_SECC:\n-\tcase ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):\n-\t\treturn true;\n-\tdefault:\n-\t\treturn false;\n-\t}\n+\treturn esr_abt_is_sea(kvm_vcpu_get_esr(vcpu));\n }\n \n static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)\n@@ -496,9 +479,9 @@ static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)\n \treturn ESR_ELx_SYS64_ISS_RT(esr);\n }\n \n-static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)\n+static inline bool esr_abt_is_write_fault(unsigned long esr)\n {\n-\tif (kvm_vcpu_abt_iss1tw(vcpu)) {\n+\tif (esr_abt_is_s1ptw(esr)) {\n \t\t/*\n \t\t * Only a permission fault on a S1PTW should be\n \t\t * considered as a write. Otherwise, page tables baked\n@@ -511,13 +494,18 @@ static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)\n \t\t * first), then a permission fault to allow the flags\n \t\t * to be set.\n \t\t */\n-\t\treturn kvm_vcpu_trap_is_permission_fault(vcpu);\n+\t\treturn esr_fsc_is_permission_fault(esr);\n \t}\n \n-\tif (kvm_vcpu_trap_is_iabt(vcpu))\n+\tif (esr_trap_is_iabt(esr))\n \t\treturn false;\n \n-\treturn kvm_vcpu_dabt_iswrite(vcpu);\n+\treturn esr_dabt_is_write(esr);\n+}\n+\n+static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)\n+{\n+\treturn esr_abt_is_write_fault(kvm_vcpu_get_esr(vcpu));\n }\n \n static inline unsigned long kvm_vcpu_get_mpidr_aff(struct kvm_vcpu *vcpu)\ndiff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h\nindex 20d12da3d28ee..589db1d6405fa 100644\n--- a/arch/arm64/include/asm/kvm_pgtable.h\n+++ b/arch/arm64/include/asm/kvm_pgtable.h\n@@ -859,6 +859,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,\n  * @addr:\tInput address for the start of the walk.\n  * @ptep:\tPointer to storage for the retrieved PTE.\n  * @level:\tPointer to storage for the level of the retrieved PTE.\n+ * @flags:\tFlags to control the page-table walk\n+ *\t\t(see struct kvm_pgtable_visit_ctx).\n  *\n  * The offset of @addr within a page is ignored.\n  *\n@@ -869,7 +871,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,\n  * Return: 0 on success, negative error code on failure.\n  */\n int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,\n-\t\t\t kvm_pte_t *ptep, s8 *level);\n+\t\t\t kvm_pte_t *ptep, s8 *level,\n+\t\t\t enum kvm_pgtable_walk_flags flags);\n \n /**\n  * kvm_pgtable_stage2_pte_prot() - Retrieve the protection attributes of a\ndiff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h\nindex cad60569f0619..c1831c8e421d4 100644\n--- a/arch/arm64/include/asm/kvm_pkvm.h\n+++ b/arch/arm64/include/asm/kvm_pkvm.h\n@@ -44,9 +44,9 @@ static inline bool kvm_pkvm_ext_allowed(struct kvm *kvm, long ext)\n \tcase KVM_CAP_ARM_PTRAUTH_GENERIC:\n \t\treturn true;\n \tcase KVM_CAP_ARM_MTE:\n-\t\treturn false;\n \tcase KVM_CAP_ARM_EAGER_SPLIT_CHUNK_SIZE:\n \tcase KVM_CAP_ARM_SUPPORTED_BLOCK_SIZES:\n+\tcase KVM_CAP_PRE_FAULT_MEMORY:\n \t\treturn false;\n \tdefault:\n \t\treturn !kvm || !kvm_vm_is_protected(kvm);\ndiff --git a/arch/arm64/kvm/Kconfig b/arch/arm64/kvm/Kconfig\nindex 449154f9a4852..71233068b7cb6 100644\n--- a/arch/arm64/kvm/Kconfig\n+++ b/arch/arm64/kvm/Kconfig\n@@ -37,6 +37,7 @@ menuconfig KVM\n \tselect SCHED_INFO\n \tselect GUEST_PERF_EVENTS if PERF_EVENTS\n \tselect KVM_GUEST_MEMFD\n+\tselect KVM_GENERIC_PRE_FAULT_MEMORY\n \thelp\n \t  Support hosting virtualized guest machines.\n \ndiff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c\nindex 31f803010c57c..367638e1e0204 100644\n--- a/arch/arm64/kvm/arm.c\n+++ b/arch/arm64/kvm/arm.c\n@@ -410,6 +410,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)\n \tcase KVM_CAP_COUNTER_OFFSET:\n \tcase KVM_CAP_ARM_WRITABLE_IMP_ID_REGS:\n \tcase KVM_CAP_ARM_SEA_TO_USER:\n+\tcase KVM_CAP_PRE_FAULT_MEMORY:\n \t\tr = 1;\n \t\tbreak;\n \tcase KVM_CAP_SET_GUEST_DEBUG2:\ndiff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c\nindex a6a47c1e058b3..443ee433119ef 100644\n--- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c\n+++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c\n@@ -540,7 +540,7 @@ static int host_stage2_adjust_range(u64 addr, struct kvm_mem_range *range)\n \tint ret;\n \n \thyp_assert_lock_held(\u0026host_mmu.lock);\n-\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, addr, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, addr, \u0026pte, \u0026level, 0);\n \tif (ret)\n \t\treturn ret;\n \n@@ -930,7 +930,7 @@ static int get_valid_guest_pte(struct pkvm_hyp_vm *vm, u64 ipa, kvm_pte_t *ptep,\n \ts8 level;\n \tint ret;\n \n-\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level, 0);\n \tif (ret)\n \t\treturn ret;\n \tif (guest_pte_is_poisoned(pte))\n@@ -979,7 +979,7 @@ int __pkvm_vcpu_in_poison_fault(struct pkvm_hyp_vcpu *hyp_vcpu)\n \tipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(\u0026hyp_vcpu-\u003evcpu));\n \n \tguest_lock_component(vm);\n-\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level, 0);\n \tif (ret)\n \t\tgoto unlock;\n \n@@ -1335,7 +1335,7 @@ static int host_stage2_get_guest_info(phys_addr_t phys, struct pkvm_hyp_vm **vm,\n \t\treturn -EPERM;\n \t}\n \n-\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, phys, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, phys, \u0026pte, \u0026level, 0);\n \tif (ret)\n \t\treturn ret;\n \n@@ -1564,7 +1564,7 @@ static int __check_host_shared_guest(struct pkvm_hyp_vm *vm, u64 *__phys, u64 ip\n \ts8 level;\n \tint ret;\n \n-\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026vm-\u003epgt, ipa, \u0026pte, \u0026level, 0);\n \tif (ret)\n \t\treturn ret;\n \tif (!kvm_pte_valid(pte))\ndiff --git a/arch/arm64/kvm/hyp/nvhe/mm.c b/arch/arm64/kvm/hyp/nvhe/mm.c\nindex 29ab5ee9d57fc..bcf2fd1ddd75d 100644\n--- a/arch/arm64/kvm/hyp/nvhe/mm.c\n+++ b/arch/arm64/kvm/hyp/nvhe/mm.c\n@@ -486,7 +486,7 @@ static int check_page_ownership(phys_addr_t phys)\n \t\t\treturn -EPERM;\n \t}\n \n-\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, phys, \u0026pte, NULL);\n+\tret = kvm_pgtable_get_leaf(\u0026host_mmu.pgt, phys, \u0026pte, NULL, 0);\n \tif (ret)\n \t\treturn ret;\n \ndiff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c\nindex b74dd5ce1efd3..347eec3957d6a 100644\n--- a/arch/arm64/kvm/hyp/pgtable.c\n+++ b/arch/arm64/kvm/hyp/pgtable.c\n@@ -298,12 +298,13 @@ static int leaf_walker(const struct kvm_pgtable_visit_ctx *ctx,\n }\n \n int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,\n-\t\t\t kvm_pte_t *ptep, s8 *level)\n+\t\t\t kvm_pte_t *ptep, s8 *level,\n+\t\t\t enum kvm_pgtable_walk_flags flags)\n {\n \tstruct leaf_walk_data data;\n \tstruct kvm_pgtable_walker walker = {\n \t\t.cb\t= leaf_walker,\n-\t\t.flags\t= KVM_PGTABLE_WALK_LEAF,\n+\t\t.flags\t= flags | KVM_PGTABLE_WALK_LEAF,\n \t\t.arg\t= \u0026data,\n \t};\n \tint ret;\ndiff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c\nindex d6187295c3736..2a281488792fb 100644\n--- a/arch/arm64/kvm/mmu.c\n+++ b/arch/arm64/kvm/mmu.c\n@@ -5,6 +5,7 @@\n  */\n \n #include \u003clinux/acpi.h\u003e\n+#include \u003clinux/cleanup.h\u003e\n #include \u003clinux/mman.h\u003e\n #include \u003clinux/kvm_host.h\u003e\n #include \u003clinux/interval_tree.h\u003e\n@@ -895,7 +896,7 @@ static int get_user_mapping_size(struct kvm *kvm, u64 addr)\n \t * IPI-ing threads).\n \t */\n \tlocal_irq_save(flags);\n-\tret = kvm_pgtable_get_leaf(\u0026pgt, addr, \u0026pte, \u0026level);\n+\tret = kvm_pgtable_get_leaf(\u0026pgt, addr, \u0026pte, \u0026level, 0);\n \tlocal_irq_restore(flags);\n \n \tif (ret)\n@@ -1606,9 +1607,9 @@ static void *get_mmu_memcache(struct kvm_vcpu *vcpu)\n \t\treturn \u0026vcpu-\u003earch.pkvm_memcache;\n }\n \n-static int topup_mmu_memcache(struct kvm_vcpu *vcpu, void *memcache)\n+static int topup_mmu_memcache(struct kvm_s2_mmu *mmu, void *memcache)\n {\n-\tint min_pages = kvm_mmu_cache_min_pages(vcpu-\u003earch.hw_mmu);\n+\tint min_pages = kvm_mmu_cache_min_pages(mmu);\n \n \tif (!is_protected_kvm_enabled())\n \t\treturn kvm_mmu_topup_memory_cache(memcache, min_pages);\n@@ -1655,15 +1656,47 @@ struct kvm_s2_fault_desc {\n \tstruct kvm_s2_trans\t*nested;\n \tstruct kvm_memory_slot\t*memslot;\n \tunsigned long\t\thva;\n+\tunsigned long\t\tesr;\n+\tstruct kvm_s2_mmu\t*mmu;\n };\n \n-static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)\n+struct kvm_s2_fault_result {\n+\tunsigned long mapping_size;\n+};\n+\n+static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)\n+{\n+\treturn esr_fsc_is_permission_fault(s2fd-\u003eesr);\n+}\n+\n+static bool kvm_s2_fault_is_exec(const struct kvm_s2_fault_desc *s2fd)\n+{\n+\treturn esr_abt_is_exec_fault(s2fd-\u003eesr);\n+}\n+\n+static bool kvm_s2_fault_is_write(const struct kvm_s2_fault_desc *s2fd)\n+{\n+\treturn esr_abt_is_write_fault(s2fd-\u003eesr);\n+}\n+\n+static u64 kvm_s2_perm_fault_granule(const struct kvm_s2_fault_desc *s2fd)\n+{\n+\tu64 level;\n+\n+\tif (!kvm_s2_fault_is_perm(s2fd))\n+\t\treturn 0;\n+\tlevel = s2fd-\u003eesr \u0026 ESR_ELx_FSC_LEVEL;\n+\treturn BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level));\n+}\n+\n+static int gmem_abort(const struct kvm_s2_fault_desc *s2fd,\n+\t\t      struct kvm_s2_fault_result *result)\n {\n \tbool write_fault, exec_fault;\n-\tbool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd-\u003evcpu);\n+\tbool perm_fault = kvm_s2_fault_is_perm(s2fd);\n \tenum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;\n \tenum kvm_pgtable_prot prot = KVM_PGTABLE_PROT_R;\n-\tstruct kvm_pgtable *pgt = s2fd-\u003evcpu-\u003earch.hw_mmu-\u003epgt;\n+\tstruct kvm_pgtable *pgt = s2fd-\u003emmu-\u003epgt;\n \tstruct kvm_guest_s2_mapping *mapping = NULL;\n \tunsigned long mmu_seq;\n \tstruct page *page;\n@@ -1675,7 +1708,7 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)\n \n \tif (!perm_fault) {\n \t\tmemcache = get_mmu_memcache(s2fd-\u003evcpu);\n-\t\tret = topup_mmu_memcache(s2fd-\u003evcpu, memcache);\n+\t\tret = topup_mmu_memcache(s2fd-\u003emmu, memcache);\n \t\tif (ret)\n \t\t\treturn ret;\n \t\tif (kvm_is_nested_s2_mmu(kvm, pgt-\u003emmu)) {\n@@ -1690,8 +1723,8 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)\n \telse\n \t\tgfn = s2fd-\u003efault_ipa \u003e\u003e PAGE_SHIFT;\n \n-\twrite_fault = kvm_is_write_fault(s2fd-\u003evcpu);\n-\texec_fault = kvm_vcpu_trap_is_exec_fault(s2fd-\u003evcpu);\n+\twrite_fault = kvm_s2_fault_is_write(s2fd);\n+\texec_fault = kvm_s2_fault_is_exec(s2fd);\n \n \tVM_WARN_ON_ONCE(write_fault \u0026\u0026 exec_fault);\n \n@@ -1701,8 +1734,10 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)\n \n \tret = kvm_gmem_get_pfn(kvm, s2fd-\u003ememslot, gfn, \u0026pfn, \u0026page, NULL);\n \tif (ret) {\n-\t\tkvm_prepare_memory_fault_exit(s2fd-\u003evcpu, s2fd-\u003efault_ipa, PAGE_SIZE,\n-\t\t\t\t\t      write_fault, exec_fault, false);\n+\t\t/* If result is non-NULL this is a synthetic fault. */\n+\t\tif (!result)\n+\t\t\tkvm_prepare_memory_fault_exit(s2fd-\u003evcpu, s2fd-\u003efault_ipa, PAGE_SIZE,\n+\t\t\t\t\t\t      write_fault, exec_fault, false);\n \t\tkfree(mapping);\n \t\treturn ret;\n \t}\n@@ -1757,7 +1792,13 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)\n \tif ((prot \u0026 KVM_PGTABLE_PROT_W) \u0026\u0026 !ret)\n \t\tmark_page_dirty_in_slot(kvm, s2fd-\u003ememslot, gfn);\n \n-\treturn ret != -EAGAIN ? ret : 0;\n+\tif (ret == -EAGAIN)\n+\t\treturn result ? ret : 0;\n+\n+\tif (result \u0026\u0026 !ret)\n+\t\tresult-\u003emapping_size = PAGE_SIZE;\n+\n+\treturn ret;\n }\n \n struct kvm_s2_fault_vma_info {\n@@ -1779,7 +1820,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n {\n \tunsigned int flags = FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE;\n \tstruct kvm_vcpu *vcpu = s2fd-\u003evcpu;\n-\tstruct kvm_pgtable *pgt = vcpu-\u003earch.hw_mmu-\u003epgt;\n+\tstruct kvm_pgtable *pgt = s2fd-\u003emmu-\u003epgt;\n \tstruct mm_struct *mm = current-\u003emm;\n \tstruct kvm *kvm = vcpu-\u003ekvm;\n \tvoid *hyp_memcache;\n@@ -1787,7 +1828,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n \tint ret;\n \n \thyp_memcache = get_mmu_memcache(vcpu);\n-\tret = topup_mmu_memcache(vcpu, hyp_memcache);\n+\tret = topup_mmu_memcache(s2fd-\u003emmu, hyp_memcache);\n \tif (ret)\n \t\treturn -ENOMEM;\n \n@@ -1910,11 +1951,6 @@ static short kvm_s2_resolve_vma_size(const struct kvm_s2_fault_desc *s2fd,\n \treturn vma_shift;\n }\n \n-static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)\n-{\n-\treturn kvm_vcpu_trap_is_permission_fault(s2fd-\u003evcpu);\n-}\n-\n static int kvm_s2_fault_get_vma_info(const struct kvm_s2_fault_desc *s2fd,\n \t\t\t\t     struct kvm_s2_fault_vma_info *s2vi)\n {\n@@ -1980,13 +2016,11 @@ static int kvm_s2_fault_pin_pfn(const struct kvm_s2_fault_desc *s2fd,\n \t\treturn ret;\n \n \ts2vi-\u003epfn = __kvm_faultin_pfn(s2fd-\u003ememslot, get_canonical_gfn(s2fd, s2vi),\n-\t\t\t\t      kvm_is_write_fault(s2fd-\u003evcpu) ? FOLL_WRITE : 0,\n+\t\t\t\t      kvm_s2_fault_is_write(s2fd) ? FOLL_WRITE : 0,\n \t\t\t\t      \u0026s2vi-\u003emap_writable, \u0026s2vi-\u003epage);\n \tif (unlikely(is_error_noslot_pfn(s2vi-\u003epfn))) {\n-\t\tif (s2vi-\u003epfn == KVM_PFN_ERR_HWPOISON) {\n-\t\t\tkvm_send_hwpoison_signal(s2fd-\u003ehva, __ffs(s2vi-\u003evma_pagesize));\n-\t\t\treturn 0;\n-\t\t}\n+\t\tif (s2vi-\u003epfn == KVM_PFN_ERR_HWPOISON)\n+\t\t\treturn -EHWPOISON;\n \t\treturn -EFAULT;\n \t}\n \n@@ -2038,7 +2072,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,\n {\n \tstruct kvm *kvm = s2fd-\u003evcpu-\u003ekvm;\n \n-\tif (kvm_vcpu_trap_is_exec_fault(s2fd-\u003evcpu) \u0026\u0026 s2vi-\u003emap_non_cacheable)\n+\tif (kvm_s2_fault_is_exec(s2fd) \u0026\u0026 s2vi-\u003emap_non_cacheable)\n \t\treturn -ENOEXEC;\n \n \t/*\n@@ -2047,7 +2081,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,\n \t * and trigger the exception here. Since the memslot is valid, inject\n \t * the fault back to the guest.\n \t */\n-\tif (esr_fsc_is_excl_atomic_fault(kvm_vcpu_get_esr(s2fd-\u003evcpu))) {\n+\tif (esr_fsc_is_excl_atomic_fault(s2fd-\u003eesr)) {\n \t\tkvm_inject_dabt_excl_atomic(s2fd-\u003evcpu, kvm_vcpu_get_hfar(s2fd-\u003evcpu));\n \t\treturn 1;\n \t}\n@@ -2056,13 +2090,13 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,\n \n \tif (s2vi-\u003emap_writable \u0026\u0026 (s2vi-\u003edevice ||\n \t\t\t\t   !memslot_is_logging(s2fd-\u003ememslot) ||\n-\t\t\t\t   kvm_is_write_fault(s2fd-\u003evcpu)))\n+\t\t\t\t   kvm_s2_fault_is_write(s2fd)))\n \t\t*prot |= KVM_PGTABLE_PROT_W;\n \n \tif (s2fd-\u003enested)\n \t\t*prot = adjust_nested_fault_perms(s2fd-\u003enested, *prot);\n \n-\tif (kvm_vcpu_trap_is_exec_fault(s2fd-\u003evcpu))\n+\tif (kvm_s2_fault_is_exec(s2fd))\n \t\t*prot |= KVM_PGTABLE_PROT_X;\n \n \tif (s2vi-\u003emap_non_cacheable)\n@@ -2086,7 +2120,8 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,\n static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,\n \t\t\t    const struct kvm_s2_fault_vma_info *s2vi,\n \t\t\t    enum kvm_pgtable_prot prot,\n-\t\t\t    void *memcache)\n+\t\t\t    void *memcache,\n+\t\t\t    struct kvm_s2_fault_result *result)\n {\n \tenum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;\n \tstruct kvm_guest_s2_mapping *mapping = NULL;\n@@ -2100,7 +2135,7 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,\n \tgfn_t gfn;\n \tint ret;\n \n-\tif (kvm_is_nested_s2_mmu(kvm, s2fd-\u003evcpu-\u003earch.hw_mmu)) {\n+\tif (kvm_is_nested_s2_mmu(kvm, s2fd-\u003emmu)) {\n \t\tmapping = kmalloc_obj(struct kvm_guest_s2_mapping,\n \t\t\t\t      GFP_KERNEL_ACCOUNT);\n \t\tif (!mapping) {\n@@ -2110,13 +2145,12 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,\n \t}\n \n \tkvm_fault_lock(kvm);\n-\tpgt = s2fd-\u003evcpu-\u003earch.hw_mmu-\u003epgt;\n+\tpgt = s2fd-\u003emmu-\u003epgt;\n \tret = -EAGAIN;\n \tif (mmu_invalidate_retry(kvm, s2vi-\u003emmu_seq))\n \t\tgoto out_unlock;\n \n-\tperm_fault_granule = (kvm_s2_fault_is_perm(s2fd) ?\n-\t\t\t      kvm_vcpu_trap_get_perm_fault_granule(s2fd-\u003evcpu) : 0);\n+\tperm_fault_granule = kvm_s2_perm_fault_granule(s2fd);\n \tmapping_size = s2vi-\u003evma_pagesize;\n \tpfn = s2vi-\u003epfn;\n \tgfn = s2vi-\u003egfn;\n@@ -2188,14 +2222,19 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,\n \t\tmark_page_dirty_in_slot(kvm, s2fd-\u003ememslot,\n \t\t\t\t\tgpa_to_gfn(canonical_ipa));\n \n-\tif (ret != -EAGAIN)\n-\t\treturn ret;\n-\treturn 0;\n+\tif (ret == -EAGAIN)\n+\t\treturn result ? ret : 0;\n+\n+\tif (result \u0026\u0026 !ret)\n+\t\tresult-\u003emapping_size = mapping_size;\n+\n+\treturn ret;\n }\n \n-static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n+static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd,\n+\t\t\t  struct kvm_s2_fault_result *result)\n {\n-\tbool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd-\u003evcpu);\n+\tbool perm_fault = kvm_s2_fault_is_perm(s2fd);\n \tstruct kvm_s2_fault_vma_info s2vi = {};\n \tenum kvm_pgtable_prot prot;\n \tvoid *memcache;\n@@ -2213,7 +2252,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n \tmemcache = get_mmu_memcache(s2fd-\u003evcpu);\n \tif (!perm_fault || memslot_is_logging(s2fd-\u003ememslot) ||\n \t    is_protected_kvm_enabled()) {\n-\t\tret = topup_mmu_memcache(s2fd-\u003evcpu, memcache);\n+\t\tret = topup_mmu_memcache(s2fd-\u003emmu, memcache);\n \t\tif (ret)\n \t\t\treturn ret;\n \t}\n@@ -2223,6 +2262,13 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n \t * get block mapping for device MMIO region.\n \t */\n \tret = kvm_s2_fault_pin_pfn(s2fd, \u0026s2vi);\n+\tif (ret == -EHWPOISON) {\n+\t\t/* If result is specified, let the caller handle this. */\n+\t\tif (result)\n+\t\t\treturn -EHWPOISON;\n+\t\tkvm_send_hwpoison_signal(s2fd-\u003ehva, __ffs(s2vi.vma_pagesize));\n+\t\treturn 0;\n+\t}\n \tif (ret != 1)\n \t\treturn ret;\n \n@@ -2232,7 +2278,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)\n \t\treturn ret;\n \t}\n \n-\treturn kvm_s2_fault_map(s2fd, \u0026s2vi, prot, memcache);\n+\treturn kvm_s2_fault_map(s2fd, \u0026s2vi, prot, memcache, result);\n }\n \n /* Resolve the access fault by making the page young again. */\n@@ -2342,7 +2388,8 @@ int kvm_handle_guest_sea(struct kvm_vcpu *vcpu)\n int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n {\n \tstruct kvm_s2_trans nested_trans, *nested = NULL;\n-\tunsigned long esr;\n+\tunsigned long esr = kvm_vcpu_get_esr(vcpu);\n+\tstruct kvm_s2_mmu *mmu = vcpu-\u003earch.hw_mmu;\n \tphys_addr_t fault_ipa; /* The address we faulted on */\n \tphys_addr_t ipa; /* Always the IPA in the L1 guest phys space */\n \tstruct kvm_memory_slot *memslot;\n@@ -2351,11 +2398,9 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \tgfn_t gfn;\n \tint ret, idx;\n \n-\tif (kvm_vcpu_abt_issea(vcpu))\n+\tif (esr_abt_is_sea(esr))\n \t\treturn kvm_handle_guest_sea(vcpu);\n \n-\tesr = kvm_vcpu_get_esr(vcpu);\n-\n \t/*\n \t * The fault IPA should be reliable at this point as we're not dealing\n \t * with an SEA.\n@@ -2364,7 +2409,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \tif (KVM_BUG_ON(ipa == INVALID_GPA, vcpu-\u003ekvm))\n \t\treturn -EFAULT;\n \n-\tis_iabt = kvm_vcpu_trap_is_iabt(vcpu);\n+\tis_iabt = esr_trap_is_iabt(esr);\n \n \tif (esr_fsc_is_translation_fault(esr)) {\n \t\t/* Beyond sanitised PARange (which is the IPA limit) */\n@@ -2374,14 +2419,14 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t\t}\n \n \t\t/* Falls between the IPA range and the PARange? */\n-\t\tif (fault_ipa \u003e= BIT_ULL(VTCR_EL2_IPA(vcpu-\u003earch.hw_mmu-\u003evtcr))) {\n+\t\tif (fault_ipa \u003e= BIT_ULL(VTCR_EL2_IPA(mmu-\u003evtcr))) {\n \t\t\tfault_ipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(vcpu));\n \n \t\t\treturn kvm_inject_sea(vcpu, is_iabt, fault_ipa);\n \t\t}\n \t}\n \n-\ttrace_kvm_guest_fault(*vcpu_pc(vcpu), kvm_vcpu_get_esr(vcpu),\n+\ttrace_kvm_guest_fault(*vcpu_pc(vcpu), esr,\n \t\t\t      kvm_vcpu_get_hfar(vcpu), fault_ipa);\n \n \t/* Check the stage-2 fault is trans. fault or write fault */\n@@ -2389,10 +2434,10 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t    !esr_fsc_is_permission_fault(esr) \u0026\u0026\n \t    !esr_fsc_is_access_flag_fault(esr) \u0026\u0026\n \t    !esr_fsc_is_excl_atomic_fault(esr)) {\n-\t\tkvm_err(\"Unsupported FSC: EC=%#x xFSC=%#lx ESR_EL2=%#lx\\n\",\n-\t\t\tkvm_vcpu_trap_get_class(vcpu),\n-\t\t\t(unsigned long)kvm_vcpu_trap_get_fault(vcpu),\n-\t\t\t(unsigned long)kvm_vcpu_get_esr(vcpu));\n+\t\tkvm_err(\"Unsupported FSC: EC=%#lx xFSC=%#lx ESR_EL2=%#lx\\n\",\n+\t\t\tESR_ELx_EC(esr),\n+\t\t\t(unsigned long)(esr \u0026 ESR_ELx_FSC),\n+\t\t\t(unsigned long)esr);\n \t\treturn -EFAULT;\n \t}\n \n@@ -2411,8 +2456,8 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t * nothing to walk and we treat it as a 1:1 before going through the\n \t * canonical translation.\n \t */\n-\tif (kvm_is_nested_s2_mmu(vcpu-\u003ekvm,vcpu-\u003earch.hw_mmu) \u0026\u0026\n-\t    vcpu-\u003earch.hw_mmu-\u003enested_stage2_enabled) {\n+\tif (kvm_is_nested_s2_mmu(vcpu-\u003ekvm, mmu) \u0026\u0026\n+\t    mmu-\u003enested_stage2_enabled) {\n \t\tu32 esr;\n \n \t\tret = kvm_walk_nested_s2(vcpu, fault_ipa, \u0026nested_trans);\n@@ -2441,7 +2486,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \tgfn = ipa \u003e\u003e PAGE_SHIFT;\n \tmemslot = gfn_to_memslot(vcpu-\u003ekvm, gfn);\n \thva = gfn_to_hva_memslot_prot(memslot, gfn, \u0026writable);\n-\twrite_fault = kvm_is_write_fault(vcpu);\n+\twrite_fault = esr_abt_is_write_fault(esr);\n \tif (kvm_is_error_hva(hva) || (write_fault \u0026\u0026 !writable)) {\n \t\t/*\n \t\t * The guest has put either its instructions or its page-tables\n@@ -2454,7 +2499,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t\t\tgoto out;\n \t\t}\n \n-\t\tif (kvm_vcpu_abt_iss1tw(vcpu)) {\n+\t\tif (esr_abt_is_s1ptw(esr)) {\n \t\t\tret = kvm_inject_sea_dabt(vcpu, kvm_vcpu_get_hfar(vcpu));\n \t\t\tgoto out_unlock;\n \t\t}\n@@ -2469,7 +2514,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t\t * So let's assume that the guest is just being\n \t\t * cautious, and skip the instruction.\n \t\t */\n-\t\tif (kvm_is_error_hva(hva) \u0026\u0026 kvm_vcpu_dabt_is_cm(vcpu)) {\n+\t\tif (kvm_is_error_hva(hva) \u0026\u0026 esr_dabt_is_cm(esr)) {\n \t\t\tkvm_incr_pc(vcpu);\n \t\t\tret = 1;\n \t\t\tgoto out_unlock;\n@@ -2487,7 +2532,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t}\n \n \t/* Userspace should not be able to register out-of-bounds IPAs */\n-\tVM_BUG_ON(ipa \u003e= kvm_phys_size(vcpu-\u003earch.hw_mmu));\n+\tVM_BUG_ON(ipa \u003e= kvm_phys_size(mmu));\n \n \tif (esr_fsc_is_access_flag_fault(esr)) {\n \t\thandle_access_fault(vcpu, fault_ipa);\n@@ -2501,19 +2546,20 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)\n \t\t.nested\t\t= nested,\n \t\t.memslot\t= memslot,\n \t\t.hva\t\t= hva,\n+\t\t.esr\t\t= esr,\n+\t\t.mmu\t\t= mmu,\n \t};\n \n \tif (kvm_vm_is_protected(vcpu-\u003ekvm)) {\n \t\tret = pkvm_mem_abort(\u0026s2fd);\n \t} else {\n-\t\tVM_WARN_ON_ONCE(kvm_vcpu_trap_is_permission_fault(vcpu) \u0026\u0026\n-\t\t\t\t!write_fault \u0026\u0026\n-\t\t\t\t!kvm_vcpu_trap_is_exec_fault(vcpu));\n+\t\tVM_WARN_ON_ONCE(kvm_s2_fault_is_perm(\u0026s2fd) \u0026\u0026 !write_fault \u0026\u0026\n+\t\t\t\t!kvm_s2_fault_is_exec(\u0026s2fd));\n \n \t\tif (kvm_slot_has_gmem(memslot))\n-\t\t\tret = gmem_abort(\u0026s2fd);\n+\t\t\tret = gmem_abort(\u0026s2fd, NULL);\n \t\telse\n-\t\t\tret = user_mem_abort(\u0026s2fd);\n+\t\t\tret = user_mem_abort(\u0026s2fd, NULL);\n \t}\n \n \tif (ret == 0)\n@@ -2890,3 +2936,146 @@ void kvm_toggle_cache(struct kvm_vcpu *vcpu, bool was_enabled)\n \n \ttrace_kvm_toggle_cache(*vcpu_pc(vcpu), was_enabled, now_enabled);\n }\n+\n+/*\n+ * Try to walk to the specified GPA in canonical mmu - if unmapped returns 0, if\n+ * mapped returns the granule size, otherwise returns an error.\n+ */\n+static long kvm_walk_s2(struct kvm_pgtable *pgt,\n+\t\t\tgpa_t gpa, s8 *level)\n+{\n+\tstruct kvm *kvm = kvm_s2_mmu_to_kvm(pgt-\u003emmu);\n+\tkvm_pte_t pte;\n+\tlong ret;\n+\n+\tguard(read_lock)(\u0026kvm-\u003emmu_lock);\n+\n+\tret = kvm_pgtable_get_leaf(pgt, gpa, \u0026pte, level,\n+\t\t\t\t   KVM_PGTABLE_WALK_SHARED);\n+\tif (ret)\n+\t\treturn ret;\n+\t/* Unpopulated, must fault. */\n+\tif (!kvm_pte_valid(pte))\n+\t\treturn 0;\n+\treturn kvm_granule_size(*level);\n+}\n+\n+/* Synthesised data abort at specified page table level. */\n+#define PRE_FAULT_ESR(level)\t\t\t\t\\\n+\t ((ESR_ELx_EC_DABT_LOW \u003c\u003c ESR_ELx_EC_SHIFT) |\t\\\n+\t  ESR_ELx_IL | ESR_ELx_FSC_FAULT_L(level))\n+\n+/* Retrieve either a read-only or a read/write hva. */\n+static hva_t gfn_to_hva_memslot_read(struct kvm_memory_slot *slot, gfn_t gfn)\n+{\n+\treturn gfn_to_hva_memslot_prot(slot, gfn, /*writable=*/NULL);\n+}\n+\n+static long __pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,\n+\t\t\t   gpa_t gpa, struct kvm_memory_slot *memslot, s8 level)\n+{\n+\tconst bool is_gmem = kvm_slot_has_gmem(memslot);\n+\tconst gfn_t gfn = gpa_to_gfn(gpa);\n+\tconst hva_t hva = is_gmem ? 0 : gfn_to_hva_memslot_read(memslot, gfn);\n+\tconst struct kvm_s2_fault_desc s2fd = {\n+\t\t.vcpu\t\t= vcpu,\n+\t\t.fault_ipa\t= gpa,\n+\t\t.nested\t\t= NULL,\n+\t\t.memslot\t= memslot,\n+\t\t.hva\t\t= hva,\n+\t\t.esr\t\t= PRE_FAULT_ESR(level),\n+\t\t.mmu\t\t= mmu,\n+\t};\n+\tstruct kvm_s2_fault_result result = {};\n+\tlong ret;\n+\n+\tif (kvm_is_error_hva(hva))\n+\t\treturn -EFAULT;\n+\n+\tif (is_gmem)\n+\t\tret = gmem_abort(\u0026s2fd, \u0026result);\n+\telse\n+\t\tret = user_mem_abort(\u0026s2fd, \u0026result);\n+\tif (IS_ERR_VALUE(ret))\n+\t\treturn ret;\n+\treturn result.mapping_size;\n+}\n+\n+static long pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,\n+\t\t\t gpa_t gpa, struct kvm_memory_slot *memslot)\n+{\n+\ts8 level;\n+\tlong ret;\n+\n+\t/* Try a walk first. */\n+\tret = kvm_walk_s2(mmu-\u003epgt, gpa, \u0026level);\n+\tif (ret)\n+\t\treturn ret;\n+\t/* OK, have to fault page in. */\n+\treturn __pre_fault_s2(mmu, vcpu, gpa, memslot, level);\n+}\n+\n+static unsigned long\n+pre_fault_bytes_consumed(gpa_t gpa, unsigned long granule_size,\n+\t\t\t unsigned long bytes_remaining)\n+{\n+\t/* Granules are always a power-of-2. */\n+\tconst unsigned long granule_bytes_remaining =\n+\t\tgranule_size - (gpa % granule_size);\n+\n+\treturn min(granule_bytes_remaining, bytes_remaining);\n+}\n+\n+/* If you lose the race this many times, time to give up. */\n+#define MAX_PRE_FAULT_RETRIES 3\n+\n+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)\n+{\n+\tif (is_protected_kvm_enabled())\n+\t\treturn -EOPNOTSUPP;\n+\tif (!kvm_vcpu_initialized(vcpu))\n+\t\treturn -ENOEXEC;\n+\n+\treturn 0;\n+}\n+\n+/**\n+ * kvm_arch_vcpu_pre_fault_memory - pre-fault stage-2 page tables for the\n+ * specified GPA.\n+ * @vcpu:\tThe VCPU pointer\n+ * @range:\t{gpa, size, flags} tuple\n+ *\n+ * The mapping performed is always best-effort - faulting in is necessarily\n+ * racey. The ranges faulted in are canonical, nested page tables are ignored.\n+ *\n+ * @range-\u003egpa specifies the GPA to pre-fault, @range-\u003esize specifies how many\n+ * bytes remain to be pre-faulted and @range-\u003eflags is reserved and must be 0.\n+ *\n+ * Returns: the number of bytes the pre-fault consumed, or an error.\n+ */\n+long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n+\t\t\t\t    struct kvm_pre_fault_memory *range)\n+{\n+\tstruct kvm *kvm = vcpu-\u003ekvm;\n+\tconst u64 bytes_remaining = range-\u003esize;\n+\tstruct kvm_s2_mmu *mmu = \u0026kvm-\u003earch.mmu; /* Canonical. */\n+\tstruct kvm_memory_slot *memslot;\n+\tconst gpa_t gpa = range-\u003egpa;\n+\tint num_retries = 0;\n+\tlong ret;\n+\n+\tmemslot = gfn_to_memslot(kvm, gpa_to_gfn(gpa));\n+\tif (!memslot)\n+\t\treturn -ENOENT;\n+\t/* SRCU must be released for progress and only userland can do that. */\n+\tif (memslot-\u003eflags \u0026 KVM_MEMSLOT_INVALID)\n+\t\treturn -EAGAIN;\n+\n+\tdo {\n+\t\tret = pre_fault_s2(mmu, vcpu, gpa, memslot);\n+\t} while (ret == -EAGAIN \u0026\u0026 num_retries++ \u003c MAX_PRE_FAULT_RETRIES);\n+\n+\tif (IS_ERR_VALUE(ret))\n+\t\treturn ret;\n+\treturn pre_fault_bytes_consumed(gpa, ret, bytes_remaining);\n+}\ndiff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c\nindex ec754a865a006..cd868af63b1b8 100644\n--- a/arch/arm64/kvm/nested.c\n+++ b/arch/arm64/kvm/nested.c\n@@ -662,7 +662,7 @@ static u8 get_guest_mapping_ttl(struct kvm_s2_mmu *mmu, u64 addr)\n \t\treturn 0;\n \n \ttmp \u0026= ~(sz - 1);\n-\tif (kvm_pgtable_get_leaf(mmu-\u003epgt, tmp, \u0026pte, NULL))\n+\tif (kvm_pgtable_get_leaf(mmu-\u003epgt, tmp, \u0026pte, NULL, 0))\n \t\tgoto again;\n \tif (!(pte \u0026 PTE_VALID))\n \t\tgoto again;\ndiff --git a/arch/s390/kvm/s390/s390.c b/arch/s390/kvm/s390/s390.c\nindex 5c73f43782a74..47fe032444f45 100644\n--- a/arch/s390/kvm/s390/s390.c\n+++ b/arch/s390/kvm/s390/s390.c\n@@ -5784,6 +5784,14 @@ void kvm_arch_commit_memory_region(struct kvm *kvm, struct kvm_memory_slot *old,\n \ts390_kvm_mmu_commit_memory_region(kvm, old, new, change);\n }\n \n+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)\n+{\n+\tif (kvm_is_ucontrol(vcpu-\u003ekvm))\n+\t\treturn -EINVAL;\n+\n+\treturn 0;\n+}\n+\n /**\n  * kvm_arch_vcpu_pre_fault_memory() -- pre-fault and link gmap dat tables\n  * @vcpu: the vcpu that shall appear to have generated the fault-in.\n@@ -5810,9 +5818,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, struct kvm_pre_fault_\n \tgpa_t end;\n \tint rc;\n \n-\tif (kvm_is_ucontrol(vcpu-\u003ekvm))\n-\t\treturn -EINVAL;\n-\n \trc = kvm_s390_faultin_gfn(vcpu, NULL, \u0026f);\n \tif (rc == PGM_ADDRESSING)\n \t\treturn -ENOENT;\ndiff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c\nindex 064ecc33b9267..c35fd2868c20b 100644\n--- a/arch/x86/kvm/mmu/mmu.c\n+++ b/arch/x86/kvm/mmu/mmu.c\n@@ -5086,6 +5086,14 @@ static int kvm_tdp_page_prefault(struct kvm_vcpu *vcpu, gpa_t gpa,\n \t}\n }\n \n+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)\n+{\n+\tif (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed)\n+\t\treturn -EOPNOTSUPP;\n+\n+\treturn 0;\n+}\n+\n long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n \t\t\t\t    struct kvm_pre_fault_memory *range)\n {\n@@ -5095,9 +5103,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n \tu64 end;\n \tint r;\n \n-\tif (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed)\n-\t\treturn -EOPNOTSUPP;\n-\n \tif (kvm_is_gfn_alias(vcpu-\u003ekvm, gpa_to_gfn(range-\u003egpa)))\n \t\treturn -EINVAL;\n \ndiff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h\nindex 03bfc92864b6e..39a4d345aeca8 100644\n--- a/include/linux/kvm_host.h\n+++ b/include/linux/kvm_host.h\n@@ -2643,6 +2643,7 @@ void kvm_arch_gmem_invalidate_range(struct kvm *kvm, struct kvm_gfn_range *range\n #endif\n \n #ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\n+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu);\n long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n \t\t\t\t    struct kvm_pre_fault_memory *range);\n #endif\ndiff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm\nindex 6a1482e3a286b..908bdc7cf4f58 100644\n--- a/tools/testing/selftests/kvm/Makefile.kvm\n+++ b/tools/testing/selftests/kvm/Makefile.kvm\n@@ -178,6 +178,7 @@ TEST_GEN_PROGS_arm64 += arm64/debug-exceptions\n TEST_GEN_PROGS_arm64 += arm64/hello_el2\n TEST_GEN_PROGS_arm64 += arm64/host_sve\n TEST_GEN_PROGS_arm64 += arm64/hypercalls\n+TEST_GEN_PROGS_arm64 += arm64/nv_pre_fault_memory_test\n TEST_GEN_PROGS_arm64 += arm64/external_aborts\n TEST_GEN_PROGS_arm64 += arm64/mmio_sign_ext\n TEST_GEN_PROGS_arm64 += arm64/page_fault_test\n@@ -205,6 +206,7 @@ TEST_GEN_PROGS_arm64 += guest_memfd_test\n TEST_GEN_PROGS_arm64 += mmu_stress_test\n TEST_GEN_PROGS_arm64 += rseq_test\n TEST_GEN_PROGS_arm64 += steal_time\n+TEST_GEN_PROGS_arm64 += pre_fault_memory_test\n \n TEST_GEN_PROGS_s390 = $(TEST_GEN_PROGS_COMMON)\n TEST_GEN_PROGS_s390 += s390/memop\ndiff --git a/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c\nnew file mode 100644\nindex 0000000000000..09c1db3038566\n--- /dev/null\n+++ b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c\n@@ -0,0 +1,158 @@\n+// SPDX-License-Identifier: GPL-2.0-only\n+/*\n+ * nv_pre_fault_memory_test - Test KVM_PRE_FAULT_MEMORY on a vCPU whose\n+ * last-run context is nested.\n+ *\n+ * The guest enters vEL2, sets up its EL2 translation configuration into the\n+ * real EL1 registers then ERETs to vEL1 and exits to userspace so its vCPU\n+ * last-run context is nested backed by a shadow stage 2 MMU.\n+ *\n+ * Assert that pre-faulting ignores that and targets the canonical stage-2\n+ * page tables only.\n+ */\n+#include \"kvm_util.h\"\n+#include \"processor.h\"\n+#include \"test_util.h\"\n+#include \"ucall.h\"\n+\n+#include \u003casm/sysreg.h\u003e\n+#include \u003clinux/sizes.h\u003e\n+\n+#define TEST_MEM_SLOT\t\t10\n+#define TEST_MEM_SIZE\t\tSZ_2M\n+#define TEST_MEM_GPA\t\tSZ_1G\n+\n+static void guest_el1_code(void)\n+{\n+\tu64 offset;\n+\n+\tGUEST_ASSERT_EQ(get_current_el(), 1);\n+\n+\t/* Exit to userspace with the vEL1 (nested) context live. */\n+\tGUEST_SYNC(1);\n+\n+\t/*\n+\t * Touch the prefaulted range. vstage-2 is disabled, so the shadow\n+\t * stage-2 is a 1:1 view of the canonical IPA space.\n+\t */\n+\tfor (offset = 0; offset \u003c TEST_MEM_SIZE; offset += SZ_4K)\n+\t\tREAD_ONCE(*(u64 *)(TEST_MEM_GPA + offset));\n+\n+\tGUEST_DONE();\n+}\n+\n+static void guest_code(void)\n+{\n+\tu64 sp;\n+\n+\tGUEST_ASSERT_EQ(get_current_el(), 2);\n+\n+\t/*\n+\t * Mirror the EL2 translation regime into the real EL1 registers so\n+\t * that vEL1 runs on the test's stage-1 page tables. With E2H=1, the\n+\t * _EL1 accessors read the EL2 registers, and the _EL12 accessors\n+\t * write the real EL1 registers.\n+\t */\n+\twrite_sysreg_s(read_sysreg(sctlr_el1), SYS_SCTLR_EL12);\n+\twrite_sysreg_s(read_sysreg(tcr_el1), SYS_TCR_EL12);\n+\twrite_sysreg_s(read_sysreg(ttbr0_el1), SYS_TTBR0_EL12);\n+\twrite_sysreg_s(read_sysreg(mair_el1), SYS_MAIR_EL12);\n+\twrite_sysreg_s(read_sysreg(cpacr_el1), SYS_CPACR_EL12);\n+\n+\t/* Run vEL1 on the same stack. */\n+\tasm volatile(\"mov %0, sp\" : \"=r\"(sp));\n+\twrite_sysreg(sp, sp_el1);\n+\n+\t/*\n+\t * Drop TGE so that vEL1 is a nested context rather than host EL0.\n+\t * KVM backs it with a shadow stage-2 MMU even though vstage-2 is\n+\t * disabled (HCR_EL2.VM=0).\n+\t */\n+\twrite_sysreg(read_sysreg(hcr_el2) \u0026 ~HCR_EL2_TGE, hcr_el2);\n+\tisb();\n+\n+\twrite_sysreg(PSR_MODE_EL1h | PSR_F_BIT | PSR_I_BIT | PSR_A_BIT |\n+\t\t     PSR_D_BIT, spsr_el2);\n+\twrite_sysreg((u64)guest_el1_code, elr_el2);\n+\tasm volatile(\"eret\");\n+\n+\tGUEST_ASSERT(false);\n+}\n+\n+static void pre_fault(struct kvm_vcpu *vcpu, u64 gpa, u64 size)\n+{\n+\tstruct kvm_pre_fault_memory range = {\n+\t\t.gpa = gpa,\n+\t\t.size = size,\n+\t};\n+\tint ret;\n+\n+\tdo {\n+\t\tret = __vcpu_ioctl(vcpu, KVM_PRE_FAULT_MEMORY, \u0026range);\n+\t} while ((!ret \u0026\u0026 range.size) ||\n+\t\t (ret \u003c 0 \u0026\u0026 (errno == EINTR || errno == EAGAIN)));\n+\n+\tTEST_ASSERT(!ret, \"KVM_PRE_FAULT_MEMORY failed, ret: %d errno: %d\",\n+\t\t    ret, errno);\n+\tTEST_ASSERT_EQ(range.size, 0);\n+}\n+\n+int main(void)\n+{\n+\tstruct kvm_vcpu_init init;\n+\tstruct kvm_vcpu *vcpu;\n+\tstruct kvm_vm *vm;\n+\tstruct ucall uc;\n+\tu64 npages;\n+\n+\tTEST_REQUIRE(test_supports_el2());\n+\tTEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));\n+\n+\tvm = vm_create(1);\n+\n+\tkvm_get_default_vcpu_target(vm, \u0026init);\n+\tinit.features[0] |= BIT(KVM_ARM_VCPU_HAS_EL2);\n+\tvcpu = aarch64_vcpu_add(vm, 0, \u0026init, guest_code);\n+\tkvm_arch_vm_finalize_vcpus(vm);\n+\n+\tnpages = TEST_MEM_SIZE / vm-\u003epage_size;\n+\tvm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, TEST_MEM_GPA,\n+\t\t\t\t    TEST_MEM_SLOT, npages, 0);\n+\tvirt_map(vm, TEST_MEM_GPA, TEST_MEM_GPA, npages);\n+\n+\t/* Run the guest until it has ERET'd from vEL2 to vEL1. */\n+\tvcpu_run(vcpu);\n+\tswitch (get_ucall(vcpu, \u0026uc)) {\n+\tcase UCALL_SYNC:\n+\t\tTEST_ASSERT_EQ(uc.args[1], 1);\n+\t\tbreak;\n+\tcase UCALL_ABORT:\n+\t\tREPORT_GUEST_ASSERT(uc);\n+\t\tbreak;\n+\tdefault:\n+\t\tTEST_FAIL(\"Unhandled ucall: %ld\", uc.cmd);\n+\t}\n+\n+\t/*\n+\t * The vCPU's last-run context is vEL1, so its hw_mmu is a shadow\n+\t * stage-2 MMU.\n+\t *\n+\t * Pre-faulting must ignore that and populate the canonical stage-2.\n+\t */\n+\tpre_fault(vcpu, TEST_MEM_GPA, TEST_MEM_SIZE);\n+\n+\t/* Resume at vEL1 and touch the prefaulted range. */\n+\tvcpu_run(vcpu);\n+\tswitch (get_ucall(vcpu, \u0026uc)) {\n+\tcase UCALL_DONE:\n+\t\tbreak;\n+\tcase UCALL_ABORT:\n+\t\tREPORT_GUEST_ASSERT(uc);\n+\t\tbreak;\n+\tdefault:\n+\t\tTEST_FAIL(\"Unhandled ucall: %ld\", uc.cmd);\n+\t}\n+\n+\tkvm_vm_free(vm);\n+\treturn 0;\n+}\ndiff --git a/tools/testing/selftests/kvm/pre_fault_memory_test.c b/tools/testing/selftests/kvm/pre_fault_memory_test.c\nindex c57631aab3d38..3082fe09b95fa 100644\n--- a/tools/testing/selftests/kvm/pre_fault_memory_test.c\n+++ b/tools/testing/selftests/kvm/pre_fault_memory_test.c\n@@ -12,19 +12,29 @@\n #include \u003cprocessor.h\u003e\n #include \u003cpthread.h\u003e\n #include \u003cucall_common.h\u003e\n+#include \u003cguest_modes.h\u003e\n \n /* Arbitrarily chosen values */\n-#define TEST_SIZE\t\t(SZ_2M + PAGE_SIZE)\n-#define TEST_NPAGES\t\t(TEST_SIZE / PAGE_SIZE)\n+#define TEST_BASE_SIZE\t\tSZ_2M\n #define TEST_SLOT\t\t10\n \n+/* Storage of test info to share with guest code */\n+struct test_config {\n+\tu64 page_size;\n+\tu64 test_size;\n+\tu64 test_num_pages;\n+};\n+\n+static struct test_config test_config;\n+\n static void guest_code(u64 base_gva)\n {\n \tvolatile u64 val __used;\n+\tstruct test_config *config = \u0026test_config;\n \tint i;\n \n-\tfor (i = 0; i \u003c TEST_NPAGES; i++) {\n-\t\tu64 *src = (u64 *)(base_gva + i * PAGE_SIZE);\n+\tfor (i = 0; i \u003c config-\u003etest_num_pages; i++) {\n+\t\tu64 *src = (u64 *)(base_gva + i * config-\u003epage_size);\n \n \t\tval = *src;\n \t}\n@@ -36,6 +46,7 @@ struct slot_worker_data {\n \tstruct kvm_vm *vm;\n \tgpa_t gpa;\n \tu32 flags;\n+\tenum vm_mem_backing_src_type mem_backing_src;\n \tbool worker_ready;\n \tbool prefault_ready;\n \tbool recreate_slot;\n@@ -56,14 +67,16 @@ static void *delete_slot_worker(void *__data)\n \twhile (!READ_ONCE(data-\u003erecreate_slot))\n \t\tcpu_relax();\n \n-\tvm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, data-\u003egpa,\n-\t\t\t\t    TEST_SLOT, TEST_NPAGES, data-\u003eflags);\n+\tvm_userspace_mem_region_add(vm, data-\u003emem_backing_src, data-\u003egpa,\n+\t\t\t\t    TEST_SLOT, test_config.test_num_pages, data-\u003eflags);\n \n \treturn NULL;\n }\n \n static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,\n-\t\t\t     u64 size, u64 expected_left, bool private)\n+\t\t\t     u64 size, u64 expected_left,\n+\t\t\t     enum vm_mem_backing_src_type mem_backing_src,\n+\t\t\t     bool private)\n {\n \tstruct kvm_pre_fault_memory range = {\n \t\t.gpa = base_gpa + offset,\n@@ -74,6 +87,7 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,\n \t\t.vm = vcpu-\u003evm,\n \t\t.gpa = base_gpa,\n \t\t.flags = private ? KVM_MEM_GUEST_MEMFD : 0,\n+\t\t.mem_backing_src = mem_backing_src,\n \t};\n \tbool slot_recreated = false;\n \tpthread_t slot_worker;\n@@ -150,8 +164,8 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,\n \t/*\n \t * Assert success if prefaulting the entire range should succeed, i.e.\n \t * complete with no bytes remaining.  Otherwise prefaulting should have\n-\t * failed due to ENOENT (due to RET_PF_EMULATE for emulated MMIO when\n-\t * no memslot exists).\n+\t * failed due to ENOENT (no memslot exists for the GPA; on x86 this\n+\t * surfaces via RET_PF_EMULATE).\n \t */\n \tif (!expected_left)\n \t\tTEST_ASSERT_VM_VCPU_IOCTL(!ret, KVM_PRE_FAULT_MEMORY, ret, vcpu-\u003evm);\n@@ -160,39 +174,82 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,\n \t\t\t\t\t  KVM_PRE_FAULT_MEMORY, ret, vcpu-\u003evm);\n }\n \n-static void __test_pre_fault_memory(unsigned long vm_type, bool private)\n+struct test_params {\n+\tunsigned long vm_type;\n+\tbool private;\n+\tenum vm_mem_backing_src_type mem_backing_src;\n+};\n+\n+static void __test_pre_fault_memory(enum vm_guest_mode guest_mode, void *arg)\n {\n-\tgpa_t gpa, gva, alignment, guest_page_size;\n+\tgpa_t gpa, gva, alignment, guest_page_size, host_page_size;\n+\tgpa_t backing_src_pagesz, mem_page_size;\n+\tstruct test_params *p = arg;\n \tconst struct vm_shape shape = {\n-\t\t.mode = VM_MODE_DEFAULT,\n-\t\t.type = vm_type,\n+\t\t.mode = guest_mode,\n+\t\t.type = p-\u003evm_type,\n \t};\n \tstruct kvm_vcpu *vcpu;\n+\tstruct kvm_run *run;\n \tstruct kvm_vm *vm;\n \tstruct ucall uc;\n \n+\tpr_info(\"Testing guest mode: %s\\n\", vm_guest_mode_string(guest_mode));\n+\tpr_info(\"Testing memory backing src type: %s\\n\",\n+\t\tvm_mem_backing_src_alias(p-\u003emem_backing_src)-\u003ename);\n+\n \tvm = vm_create_shape_with_one_vcpu(shape, \u0026vcpu, guest_code);\n \n-\talignment = guest_page_size = vm_guest_mode_params[VM_MODE_DEFAULT].page_size;\n-\tgpa = (vm-\u003emax_gfn - TEST_NPAGES) * guest_page_size;\n+\tguest_page_size = vm_guest_mode_params[guest_mode].page_size;\n+\thost_page_size = getpagesize();\n+\tbacking_src_pagesz = get_backing_src_pagesz(p-\u003emem_backing_src);\n+\tmem_page_size = max(host_page_size, backing_src_pagesz);\n+\n+\ttest_config.page_size = guest_page_size;\n+\ttest_config.test_size = align_up(TEST_BASE_SIZE + test_config.page_size,\n+\t\t\t\t\t mem_page_size);\n+\ttest_config.test_num_pages = vm_calc_num_guest_pages(vm-\u003emode, test_config.test_size);\n+\n+\tgpa = (vm-\u003emax_gfn - test_config.test_num_pages) * test_config.page_size;\n \talignment = SZ_2M;\n+\talignment = max(alignment, mem_page_size);\n \tgpa = align_down(gpa, alignment);\n \tgva = gpa \u0026 ((1ULL \u003c\u003c (vm-\u003eva_bits - 1)) - 1);\n \n-\tvm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa, TEST_SLOT,\n-\t\t\t\t    TEST_NPAGES, private ? KVM_MEM_GUEST_MEMFD : 0);\n-\tvirt_map(vm, gva, gpa, TEST_NPAGES);\n+\tvm_userspace_mem_region_add(vm, p-\u003emem_backing_src,\n+\t\t\t\t    gpa, TEST_SLOT, test_config.test_num_pages,\n+\t\t\t\t    p-\u003eprivate ? KVM_MEM_GUEST_MEMFD : 0);\n+\tvirt_map(vm, gva, gpa, test_config.test_num_pages);\n+\n+\tif (p-\u003eprivate)\n+\t\tvm_mem_set_private(vm, gpa, test_config.test_size);\n+\n+\tpre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,\n+\t\t\t p-\u003emem_backing_src, p-\u003eprivate);\n+\t/* Retry the same range after the first prefault attempt. */\n+\tpre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,\n+\t\t\t p-\u003emem_backing_src, p-\u003eprivate);\n+\tpre_fault_memory(vcpu, gpa,\n+\t\t\t test_config.test_size - host_page_size,\n+\t\t\t host_page_size * 2, host_page_size,\n+\t\t\t p-\u003emem_backing_src, p-\u003eprivate);\n+\tpre_fault_memory(vcpu, gpa, test_config.test_size,\n+\t\t\t host_page_size, host_page_size,\n+\t\t\t p-\u003emem_backing_src, p-\u003eprivate);\n \n-\tif (private)\n-\t\tvm_mem_set_private(vm, gpa, TEST_SIZE);\n+\tvcpu_args_set(vcpu, 1, gva);\n \n-\tpre_fault_memory(vcpu, gpa, 0, SZ_2M, 0, private);\n-\tpre_fault_memory(vcpu, gpa, SZ_2M, PAGE_SIZE * 2, PAGE_SIZE, private);\n-\tpre_fault_memory(vcpu, gpa, TEST_SIZE, PAGE_SIZE, PAGE_SIZE, private);\n+\t/* Export the shared variables to the guest. */\n+\tsync_global_to_guest(vm, test_config);\n \n-\tvcpu_args_set(vcpu, 1, gva);\n \tvcpu_run(vcpu);\n \n+\trun = vcpu-\u003erun;\n+\tTEST_ASSERT(run-\u003eexit_reason == UCALL_EXIT_REASON,\n+\t\t    \"Wanted %s, got exit reason: %u (%s)\",\n+\t\t    exit_reason_str(UCALL_EXIT_REASON),\n+\t\t    run-\u003eexit_reason, exit_reason_str(run-\u003eexit_reason));\n+\n \tswitch (get_ucall(vcpu, \u0026uc)) {\n \tcase UCALL_ABORT:\n \t\tREPORT_GUEST_ASSERT(uc);\n@@ -207,24 +264,61 @@ static void __test_pre_fault_memory(unsigned long vm_type, bool private)\n \tkvm_vm_free(vm);\n }\n \n-static void test_pre_fault_memory(unsigned long vm_type, bool private)\n+static void test_pre_fault_memory(unsigned long vm_type, enum vm_mem_backing_src_type backing_src,\n+\t\t\t\t  bool private)\n {\n+\tstruct test_params p = {\n+\t\t.vm_type = vm_type,\n+\t\t.private = private,\n+\t\t.mem_backing_src = backing_src,\n+\t};\n+\n \tif (vm_type \u0026\u0026 !(kvm_check_cap(KVM_CAP_VM_TYPES) \u0026 BIT(vm_type))) {\n \t\tpr_info(\"Skipping tests for vm_type 0x%lx\\n\", vm_type);\n \t\treturn;\n \t}\n \n-\t__test_pre_fault_memory(vm_type, private);\n+\tfor_each_guest_mode(__test_pre_fault_memory, \u0026p);\n+}\n+\n+static void help(char *name)\n+{\n+\tputs(\"\");\n+\tprintf(\"usage: %s [-h] [-m mode] [-s mem-type]\\n\", name);\n+\tputs(\"\");\n+\tguest_modes_help();\n+\tbacking_src_help(\"-s\");\n+\tputs(\"\");\n }\n \n int main(int argc, char *argv[])\n {\n+\tenum vm_mem_backing_src_type backing = DEFAULT_VM_MEM_SRC;\n+\tint opt;\n+\n+\tguest_modes_append_default();\n+\n+\twhile ((opt = getopt(argc, argv, \"hm:s:\")) != -1) {\n+\t\tswitch (opt) {\n+\t\tcase 'm':\n+\t\t\tguest_modes_cmdline(optarg);\n+\t\t\tbreak;\n+\t\tcase 's':\n+\t\t\tbacking = parse_backing_src_type(optarg);\n+\t\t\tbreak;\n+\t\tcase 'h':\n+\t\tdefault:\n+\t\t\thelp(argv[0]);\n+\t\t\texit(0);\n+\t\t}\n+\t}\n+\n \tTEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));\n \n-\ttest_pre_fault_memory(0, false);\n+\ttest_pre_fault_memory(0, backing, false);\n #ifdef __x86_64__\n-\ttest_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, false);\n-\ttest_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, true);\n+\ttest_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, false);\n+\ttest_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, true);\n #endif\n \treturn 0;\n }\ndiff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c\nindex 65eb26a0520d8..41356577bee79 100644\n--- a/virt/kvm/kvm_main.c\n+++ b/virt/kvm/kvm_main.c\n@@ -4353,7 +4353,7 @@ static int kvm_vcpu_ioctl_get_stats_fd(struct kvm_vcpu *vcpu)\n static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n \t\t\t\t     struct kvm_pre_fault_memory *range)\n {\n-\tint idx;\n+\tint idx, err;\n \tlong r;\n \tu64 full_size;\n \n@@ -4365,6 +4365,14 @@ static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n \t    range-\u003egpa + range-\u003esize \u003c= range-\u003egpa)\n \t\treturn -EINVAL;\n \n+\t/*\n+\t * Certain architectures (e.g. arm64) need to reject the ioctl 'early'\n+\t * before vcpu_load().\n+\t */\n+\terr = kvm_arch_pre_fault_allowed(vcpu);\n+\tif (err)\n+\t\treturn err;\n+\n \tvcpu_load(vcpu);\n \tidx = srcu_read_lock(\u0026vcpu-\u003ekvm-\u003esrcu);\n \n"
}

2/1 2026/09/25 22:21 agent patch-evaluator
1m Model:
core

Args:
null

Results:
{
  "EnableConfigs": [
    "KVM",
    "KVM_INTEL",
    "KVM_AMD",
    "VIRTUALIZATION"
  ],
  "FocusSymbols": [
    "kvm_vcpu_pre_fault_memory",
    "kvm_arch_pre_fault_allowed",
    "kvm_arch_vcpu_pre_fault_memory"
  ],
  "Reasoning": "While the patch primarily implements KVM_PRE_FAULT_MEMORY for arm64, it also refactors generic KVM (virt/kvm/kvm_main.c) and x86 KVM (arch/x86/kvm/mmu/mmu.c). It introduces the kvm_arch_pre_fault_allowed() hook, altering the control flow of the KVM_PRE_FAULT_MEMORY vcpu ioctl on amd64 to validate prefault permissions early prior to calling vcpu_load() and taking the SRCU lock. Because reachable executable code in the KVM subsystem is modified on amd64, it warrants fuzzing.",
  "WorthFuzzing": true
}

Instruction:
You are an expert Linux kernel maintainer and security engineer.
Your job is to review a provided patch series and evaluate whether it warrants fuzzing with syzkaller.

IMPORTANT: The changes have ALREADY been applied and committed as the HEAD commit in
your workspace. Do NOT rely on internal assumptions. You must actively use your code access
tools to inspect the actual source code, callers, and surrounding context.

================================================================================
1. CORE TRIAGE PHILOSOPHY
================================================================================
The goal of patch fuzzing is to discover crashes, regressions, exposed latent bugs,
and newly triggered assertions introduced by the patch series.

- REACHABILITY IS THE PRIMARY GATE:
  Fuzzing can only discover bugs in code that can actually execute in standard virtualized
  environments (GCE or QEMU, utilizing software-emulated devices like USB gadgets, netdev, tun/tap).
  If the modified code is structurally unreachable (see Section 2), it MUST NOT be fuzzed,
  regardless of whether it adds assertions or complex logic.

- DO NOT BLINDLY TRUST "NO FUNCTIONAL CHANGE" (NFCI) OR "REFACTORING" CLAIMS:
  Patch authors routinely label changes as "cleanups", "refactorings", or state
  "No functional change intended". Do NOT take these claims at face value.
  Code refactorings that rearrange logic, introduce helper functions, or alter state management
  in core subsystems frequently introduce subtle semantic shifts or uncover latent kernel bugs.
  If reachable executable code is modified or refactored, it MUST be fuzzed.

- NEW OR MODIFIED ASSERTIONS IN REACHABLE CODE MUST BE FUZZED:
  When a patch introduces or modifies runtime checks or assertions (e.g., WARN_ON*, VM_WARN_ON*,
  BUG_ON*, lockdep_assert*) in reachable code paths, it enforces new or stricter invariants.
  Even if the author believes the invariant always holds, fuzzing is essential to verify whether
  an unusual sequence of operations can violate it.

================================================================================
2. WHEN TO RETURN WorthFuzzing=false (NEGATIVE CRITERIA)
================================================================================
Return WorthFuzzing=false ONLY IF all modified code falls strictly into one or more of these categories:

- Non-kernel and non-executable changes:
  * Modifications to Documentation/, comments, or spelling fixes.
  * User-space directories, self-tests, samples, or scripts (e.g., tools/, samples/, scripts/, usr/)
    that do not affect the compiled kernel image (vmlinux) or kernel modules.
  * Purely decorative logging (e.g., message strings in pr_err, printk, dev_info) or tracepoints
    that do not alter control flow or data structures.
  * Build system or Kconfig changes that do not alter compiled C logic.
- Structurally unreachable hardware:
  * Vendor-specific PCIe switches, SmartNICs, or GPU drivers (e.g., mlxsw, pds_core, qed,
    ionic, amdgpu) requiring physical ASIC/PCIe cards not emulated in standard QEMU.
- Unreachable execution paths:
  * Driver teardown callbacks (.remove, .shutdown, pci_unregister_driver) executed only during
    physical PCI hot-unplug or manual sysfs driver unbinding.
  * Code paths exclusive to architectures other than the target architecture.

================================================================================
3. WHEN TO RETURN WorthFuzzing=true (POSITIVE CRITERIA)
================================================================================
Return WorthFuzzing=true whenever the patch touches reachable executable code, including:
- Core Subsystems:
  * Any logic modifications in memory management (mm/), synchronization/locking (kernel/locking/),
    BPF, scheduler, core networking, VFS, or syscall handling.
- Refactorings and Code Cleanups:
  * Any restructuring of reachable data structures, helper abstractions, or algorithm flows.
- Runtime Assertions and Defensive Checks:
  * Any introduction or alteration of assertions (WARN_ON*, VM_WARN_ON*, BUG_ON*, etc.) in reachable paths.
- Reachable Drivers and Protocols:
  * Drivers accessible via virtual buses (virtio, USB gadget, loopback, netlink, binder, sockets, etc.).

================================================================================
4. EXTRACTING FocusSymbols (PREVENTING DILUTION)
================================================================================
When WorthFuzzing=true, you must extract specific kernel functions into FocusSymbols to guide the fuzzer:

- AVOID UBIQUITOUS LIFECYCLE HOT-PATHS:
  Do NOT list generic, ubiquitous functions called by almost every program in the corpus
  (including, but not limited to: general memory allocators and deallocators, page fault
  and trap handlers, or core synchronization primitives; this is not an exhaustive list).
  Listing ubiquitous functions causes the fuzzer to classify thousands of unrelated tests as "focused",
  which severely dilutes fuzzing effort away from the actual changes.

- TARGET SPECIFIC FEATURE LOGIC AND ENTRYPOINTS:
  List functions that specifically implement the logic being added or altered, or direct API entrypoints
  for the subsystem feature under review.

- HANDLING STATIC INLINE FUNCTIONS IN HEADERS (.h):
  Compiler-inlined static functions (such as static inlines in mm/*.h or include/linux/*.h) lack
  distinct symbol addresses in vmlinux and cannot be targeted directly by symbol coverage filters.
  If the changes are primarily in static inline helpers, identify non-static, feature-specific caller
  functions in .c files that exercise them (avoiding ubiquitous lifecycle wrappers).

================================================================================
5. IDENTIFYING EnableConfigs
================================================================================
Identify any specific CONFIG_ options required to properly compile and reach the modified code:
- Inspect Kconfig files and #ifdef guards; do not make assumptions.
- Check "depends on" lines in Kconfig to include any non-standard parent subsystem configs needed.
- Strip any 'CONFIG_' prefix (e.g., return "NET_IPV4" instead of "CONFIG_NET_IPV4").
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit d79cc5272139a53b1da70940b59fa09bf4f16b16
Author: syz-cluster <triage@syzkaller.com>
Date:   Fri Sep 25 22:21:57 2026 +0000

    syz-cluster: applied patch under review

diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index 1a90598901c57..71b642148b7f6 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6489,7 +6489,7 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.
 ---------------------------
 
 :Capability: KVM_CAP_PRE_FAULT_MEMORY
-:Architectures: none
+:Architectures: x86, s390, arm64
 :Type: vcpu ioctl
 :Parameters: struct kvm_pre_fault_memory (in/out)
 :Returns: 0 if at least one page is processed, < 0 on error
@@ -6497,12 +6497,15 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.
 Errors:
 
   ========== ===============================================================
+  EAGAIN     A race occurred before progress was made, but a retry may succeed.
   EINVAL     The specified `gpa` and `size` were invalid (e.g. not
              page aligned, causes an overflow, or size is zero), or the VM
              is UCONTROL (s390).
   ENOENT     The specified `gpa` is outside defined memslots.
+  ENOEXEC    The vCPU has not been initialised (arm64).
   EINTR      An unmasked signal is pending and no page was processed.
   EFAULT     The parameter address was invalid.
+  EHWPOISON  A poisoned host page was encountered.
   EOPNOTSUPP Mapping memory for a GPA is unsupported by the
              hypervisor, and/or for the current vCPU state/mode.
   EIO        unexpected error conditions (also causes a WARN)
@@ -6522,7 +6525,17 @@ Errors:
 KVM_PRE_FAULT_MEMORY populates KVM's stage-2 page tables used to map memory
 for the current vCPU state.  KVM maps memory as if the vCPU generated a
 stage-2 read page fault, e.g. faults in memory as needed, but doesn't break
-CoW.  On x86, KVM does not mark any newly created stage-2 PTE as Accessed.
+CoW.  On arm64, KVM marks newly created stage-2 PTEs as Accessed, as it
+does for any stage-2 fault, but leaves the Accessed state of existing PTEs
+unchanged.  On x86, KVM does not mark any newly created stage-2 PTE as
+Accessed, and for s390 it is not applicable.
+
+On arm64, a GPA is interpreted as an IPA, and never interpreted as the IPA
+of a nested guest. Pre-faulting only populates canonical stage-2 page
+tables.
+
+The feature is not supported on arm64 if the protected KVM (pKVM) feature
+is enabled.
 
 In the case of confidential VM types where there is an initial set up of
 private guest memory before the guest is 'finalized'/measured, this ioctl
@@ -6537,7 +6550,7 @@ When the ioctl returns, the input values are updated to point to the
 remaining range.  If `size` > 0 on return, the caller can just issue
 the ioctl again with the same `struct kvm_map_memory` argument.
 
-Shadow page tables cannot support this ioctl because they
+On x86, shadow page tables cannot support this ioctl because they
 are indexed by virtual address or nested guest physical address.
 Calling this ioctl when the guest is using shadow page tables (for
 example because it is running a nested guest with nested page tables)
diff --git a/arch/arm64/include/asm/esr.h b/arch/arm64/include/asm/esr.h
index f816f5d77f1a5..86756fc9bb5a2 100644
--- a/arch/arm64/include/asm/esr.h
+++ b/arch/arm64/include/asm/esr.h
@@ -437,6 +437,35 @@
 #ifndef __ASSEMBLER__
 #include <asm/types.h>
 
+static __always_inline bool esr_trap_is_iabt(unsigned long esr)
+{
+	return ESR_ELx_EC(esr) == ESR_ELx_EC_IABT_LOW;
+}
+
+/* Always check for S1PTW *before* using this. */
+static __always_inline bool esr_dabt_is_write(unsigned long esr)
+{
+	return esr & ESR_ELx_WNR;
+}
+
+static __always_inline bool esr_dabt_is_cm(unsigned long esr)
+{
+	return esr & ESR_ELx_CM;
+}
+
+static __always_inline bool esr_abt_is_sea(unsigned long esr)
+{
+	switch (esr & ESR_ELx_FSC) {
+	case ESR_ELx_FSC_EXTABT:
+	case ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):
+	case ESR_ELx_FSC_SECC:
+	case ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):
+		return true;
+	default:
+		return false;
+	}
+}
+
 static inline unsigned long esr_brk_comment(unsigned long esr)
 {
 	return esr & ESR_ELx_BRK64_ISS_COMMENT_MASK;
diff --git a/arch/arm64/include/asm/kvm_emulate.h b/arch/arm64/include/asm/kvm_emulate.h
index a3c1928bdf743..708a4b6d88230 100644
--- a/arch/arm64/include/asm/kvm_emulate.h
+++ b/arch/arm64/include/asm/kvm_emulate.h
@@ -336,6 +336,16 @@ static __always_inline u64 kvm_vcpu_get_esr(const struct kvm_vcpu *vcpu)
 	return vcpu->arch.fault.esr_el2;
 }
 
+static __always_inline bool esr_abt_is_s1ptw(unsigned long esr)
+{
+	return esr & ESR_ELx_S1PTW;
+}
+
+static __always_inline bool esr_abt_is_exec_fault(unsigned long esr)
+{
+	return esr_trap_is_iabt(esr) && !esr_abt_is_s1ptw(esr);
+}
+
 static inline bool guest_hyp_wfx_traps_enabled(const struct kvm_vcpu *vcpu)
 {
 	u64 esr = kvm_vcpu_get_esr(vcpu);
@@ -411,18 +421,13 @@ static __always_inline int kvm_vcpu_dabt_get_rd(const struct kvm_vcpu *vcpu)
 
 static __always_inline bool kvm_vcpu_abt_iss1tw(const struct kvm_vcpu *vcpu)
 {
-	return !!(kvm_vcpu_get_esr(vcpu) & ESR_ELx_S1PTW);
+	return esr_abt_is_s1ptw(kvm_vcpu_get_esr(vcpu));
 }
 
 /* Always check for S1PTW *before* using this. */
 static __always_inline bool kvm_vcpu_dabt_iswrite(const struct kvm_vcpu *vcpu)
 {
-	return kvm_vcpu_get_esr(vcpu) & ESR_ELx_WNR;
-}
-
-static inline bool kvm_vcpu_dabt_is_cm(const struct kvm_vcpu *vcpu)
-{
-	return !!(kvm_vcpu_get_esr(vcpu) & ESR_ELx_CM);
+	return esr_dabt_is_write(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline unsigned int kvm_vcpu_dabt_get_as(const struct kvm_vcpu *vcpu)
@@ -443,12 +448,7 @@ static __always_inline u8 kvm_vcpu_trap_get_class(const struct kvm_vcpu *vcpu)
 
 static inline bool kvm_vcpu_trap_is_iabt(const struct kvm_vcpu *vcpu)
 {
-	return kvm_vcpu_trap_get_class(vcpu) == ESR_ELx_EC_IABT_LOW;
-}
-
-static inline bool kvm_vcpu_trap_is_exec_fault(const struct kvm_vcpu *vcpu)
-{
-	return kvm_vcpu_trap_is_iabt(vcpu) && !kvm_vcpu_abt_iss1tw(vcpu);
+	return esr_trap_is_iabt(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline u8 kvm_vcpu_trap_get_fault(const struct kvm_vcpu *vcpu)
@@ -468,26 +468,9 @@ bool kvm_vcpu_trap_is_translation_fault(const struct kvm_vcpu *vcpu)
 	return esr_fsc_is_translation_fault(kvm_vcpu_get_esr(vcpu));
 }
 
-static inline
-u64 kvm_vcpu_trap_get_perm_fault_granule(const struct kvm_vcpu *vcpu)
-{
-	unsigned long esr = kvm_vcpu_get_esr(vcpu);
-
-	BUG_ON(!esr_fsc_is_permission_fault(esr));
-	return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(esr & ESR_ELx_FSC_LEVEL));
-}
-
 static __always_inline bool kvm_vcpu_abt_issea(const struct kvm_vcpu *vcpu)
 {
-	switch (kvm_vcpu_trap_get_fault(vcpu)) {
-	case ESR_ELx_FSC_EXTABT:
-	case ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):
-	case ESR_ELx_FSC_SECC:
-	case ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):
-		return true;
-	default:
-		return false;
-	}
+	return esr_abt_is_sea(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)
@@ -496,9 +479,9 @@ static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)
 	return ESR_ELx_SYS64_ISS_RT(esr);
 }
 
-static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
+static inline bool esr_abt_is_write_fault(unsigned long esr)
 {
-	if (kvm_vcpu_abt_iss1tw(vcpu)) {
+	if (esr_abt_is_s1ptw(esr)) {
 		/*
 		 * Only a permission fault on a S1PTW should be
 		 * considered as a write. Otherwise, page tables baked
@@ -511,13 +494,18 @@ static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
 		 * first), then a permission fault to allow the flags
 		 * to be set.
 		 */
-		return kvm_vcpu_trap_is_permission_fault(vcpu);
+		return esr_fsc_is_permission_fault(esr);
 	}
 
-	if (kvm_vcpu_trap_is_iabt(vcpu))
+	if (esr_trap_is_iabt(esr))
 		return false;
 
-	return kvm_vcpu_dabt_iswrite(vcpu);
+	return esr_dabt_is_write(esr);
+}
+
+static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
+{
+	return esr_abt_is_write_fault(kvm_vcpu_get_esr(vcpu));
 }
 
 static inline unsigned long kvm_vcpu_get_mpidr_aff(struct kvm_vcpu *vcpu)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 20d12da3d28ee..589db1d6405fa 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -859,6 +859,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,
  * @addr:	Input address for the start of the walk.
  * @ptep:	Pointer to storage for the retrieved PTE.
  * @level:	Pointer to storage for the level of the retrieved PTE.
+ * @flags:	Flags to control the page-table walk
+ *		(see struct kvm_pgtable_visit_ctx).
  *
  * The offset of @addr within a page is ignored.
  *
@@ -869,7 +871,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,
  * Return: 0 on success, negative error code on failure.
  */
 int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
-			 kvm_pte_t *ptep, s8 *level);
+			 kvm_pte_t *ptep, s8 *level,
+			 enum kvm_pgtable_walk_flags flags);
 
 /**
  * kvm_pgtable_stage2_pte_prot() - Retrieve the protection attributes of a
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index cad60569f0619..c1831c8e421d4 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -44,9 +44,9 @@ static inline bool kvm_pkvm_ext_allowed(struct kvm *kvm, long ext)
 	case KVM_CAP_ARM_PTRAUTH_GENERIC:
 		return true;
 	case KVM_CAP_ARM_MTE:
-		return false;
 	case KVM_CAP_ARM_EAGER_SPLIT_CHUNK_SIZE:
 	case KVM_CAP_ARM_SUPPORTED_BLOCK_SIZES:
+	case KVM_CAP_PRE_FAULT_MEMORY:
 		return false;
 	default:
 		return !kvm || !kvm_vm_is_protected(kvm);
diff --git a/arch/arm64/kvm/Kconfig b/arch/arm64/kvm/Kconfig
index 449154f9a4852..71233068b7cb6 100644
--- a/arch/arm64/kvm/Kconfig
+++ b/arch/arm64/kvm/Kconfig
@@ -37,6 +37,7 @@ menuconfig KVM
 	select SCHED_INFO
 	select GUEST_PERF_EVENTS if PERF_EVENTS
 	select KVM_GUEST_MEMFD
+	select KVM_GENERIC_PRE_FAULT_MEMORY
 	help
 	  Support hosting virtualized guest machines.
 
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 31f803010c57c..367638e1e0204 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -410,6 +410,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
 	case KVM_CAP_COUNTER_OFFSET:
 	case KVM_CAP_ARM_WRITABLE_IMP_ID_REGS:
 	case KVM_CAP_ARM_SEA_TO_USER:
+	case KVM_CAP_PRE_FAULT_MEMORY:
 		r = 1;
 		break;
 	case KVM_CAP_SET_GUEST_DEBUG2:
diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
index a6a47c1e058b3..443ee433119ef 100644
--- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c
+++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
@@ -540,7 +540,7 @@ static int host_stage2_adjust_range(u64 addr, struct kvm_mem_range *range)
 	int ret;
 
 	hyp_assert_lock_held(&host_mmu.lock);
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level, 0);
 	if (ret)
 		return ret;
 
@@ -930,7 +930,7 @@ static int get_valid_guest_pte(struct pkvm_hyp_vm *vm, u64 ipa, kvm_pte_t *ptep,
 	s8 level;
 	int ret;
 
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		return ret;
 	if (guest_pte_is_poisoned(pte))
@@ -979,7 +979,7 @@ int __pkvm_vcpu_in_poison_fault(struct pkvm_hyp_vcpu *hyp_vcpu)
 	ipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(&hyp_vcpu->vcpu));
 
 	guest_lock_component(vm);
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		goto unlock;
 
@@ -1335,7 +1335,7 @@ static int host_stage2_get_guest_info(phys_addr_t phys, struct pkvm_hyp_vm **vm,
 		return -EPERM;
 	}
 
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level, 0);
 	if (ret)
 		return ret;
 
@@ -1564,7 +1564,7 @@ static int __check_host_shared_guest(struct pkvm_hyp_vm *vm, u64 *__phys, u64 ip
 	s8 level;
 	int ret;
 
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		return ret;
 	if (!kvm_pte_valid(pte))
diff --git a/arch/arm64/kvm/hyp/nvhe/mm.c b/arch/arm64/kvm/hyp/nvhe/mm.c
index 29ab5ee9d57fc..bcf2fd1ddd75d 100644
--- a/arch/arm64/kvm/hyp/nvhe/mm.c
+++ b/arch/arm64/kvm/hyp/nvhe/mm.c
@@ -486,7 +486,7 @@ static int check_page_ownership(phys_addr_t phys)
 			return -EPERM;
 	}
 
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, NULL);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, NULL, 0);
 	if (ret)
 		return ret;
 
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index b74dd5ce1efd3..347eec3957d6a 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -298,12 +298,13 @@ static int leaf_walker(const struct kvm_pgtable_visit_ctx *ctx,
 }
 
 int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
-			 kvm_pte_t *ptep, s8 *level)
+			 kvm_pte_t *ptep, s8 *level,
+			 enum kvm_pgtable_walk_flags flags)
 {
 	struct leaf_walk_data data;
 	struct kvm_pgtable_walker walker = {
 		.cb	= leaf_walker,
-		.flags	= KVM_PGTABLE_WALK_LEAF,
+		.flags	= flags | KVM_PGTABLE_WALK_LEAF,
 		.arg	= &data,
 	};
 	int ret;
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index d6187295c3736..2a281488792fb 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -5,6 +5,7 @@
  */
 
 #include <linux/acpi.h>
+#include <linux/cleanup.h>
 #include <linux/mman.h>
 #include <linux/kvm_host.h>
 #include <linux/interval_tree.h>
@@ -895,7 +896,7 @@ static int get_user_mapping_size(struct kvm *kvm, u64 addr)
 	 * IPI-ing threads).
 	 */
 	local_irq_save(flags);
-	ret = kvm_pgtable_get_leaf(&pgt, addr, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&pgt, addr, &pte, &level, 0);
 	local_irq_restore(flags);
 
 	if (ret)
@@ -1606,9 +1607,9 @@ static void *get_mmu_memcache(struct kvm_vcpu *vcpu)
 		return &vcpu->arch.pkvm_memcache;
 }
 
-static int topup_mmu_memcache(struct kvm_vcpu *vcpu, void *memcache)
+static int topup_mmu_memcache(struct kvm_s2_mmu *mmu, void *memcache)
 {
-	int min_pages = kvm_mmu_cache_min_pages(vcpu->arch.hw_mmu);
+	int min_pages = kvm_mmu_cache_min_pages(mmu);
 
 	if (!is_protected_kvm_enabled())
 		return kvm_mmu_topup_memory_cache(memcache, min_pages);
@@ -1655,15 +1656,47 @@ struct kvm_s2_fault_desc {
 	struct kvm_s2_trans	*nested;
 	struct kvm_memory_slot	*memslot;
 	unsigned long		hva;
+	unsigned long		esr;
+	struct kvm_s2_mmu	*mmu;
 };
 
-static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
+struct kvm_s2_fault_result {
+	unsigned long mapping_size;
+};
+
+static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_fsc_is_permission_fault(s2fd->esr);
+}
+
+static bool kvm_s2_fault_is_exec(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_abt_is_exec_fault(s2fd->esr);
+}
+
+static bool kvm_s2_fault_is_write(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_abt_is_write_fault(s2fd->esr);
+}
+
+static u64 kvm_s2_perm_fault_granule(const struct kvm_s2_fault_desc *s2fd)
+{
+	u64 level;
+
+	if (!kvm_s2_fault_is_perm(s2fd))
+		return 0;
+	level = s2fd->esr & ESR_ELx_FSC_LEVEL;
+	return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level));
+}
+
+static int gmem_abort(const struct kvm_s2_fault_desc *s2fd,
+		      struct kvm_s2_fault_result *result)
 {
 	bool write_fault, exec_fault;
-	bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
+	bool perm_fault = kvm_s2_fault_is_perm(s2fd);
 	enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;
 	enum kvm_pgtable_prot prot = KVM_PGTABLE_PROT_R;
-	struct kvm_pgtable *pgt = s2fd->vcpu->arch.hw_mmu->pgt;
+	struct kvm_pgtable *pgt = s2fd->mmu->pgt;
 	struct kvm_guest_s2_mapping *mapping = NULL;
 	unsigned long mmu_seq;
 	struct page *page;
@@ -1675,7 +1708,7 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 
 	if (!perm_fault) {
 		memcache = get_mmu_memcache(s2fd->vcpu);
-		ret = topup_mmu_memcache(s2fd->vcpu, memcache);
+		ret = topup_mmu_memcache(s2fd->mmu, memcache);
 		if (ret)
 			return ret;
 		if (kvm_is_nested_s2_mmu(kvm, pgt->mmu)) {
@@ -1690,8 +1723,8 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 	else
 		gfn = s2fd->fault_ipa >> PAGE_SHIFT;
 
-	write_fault = kvm_is_write_fault(s2fd->vcpu);
-	exec_fault = kvm_vcpu_trap_is_exec_fault(s2fd->vcpu);
+	write_fault = kvm_s2_fault_is_write(s2fd);
+	exec_fault = kvm_s2_fault_is_exec(s2fd);
 
 	VM_WARN_ON_ONCE(write_fault && exec_fault);
 
@@ -1701,8 +1734,10 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 
 	ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL);
 	if (ret) {
-		kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE,
-					      write_fault, exec_fault, false);
+		/* If result is non-NULL this is a synthetic fault. */
+		if (!result)
+			kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE,
+						      write_fault, exec_fault, false);
 		kfree(mapping);
 		return ret;
 	}
@@ -1757,7 +1792,13 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 	if ((prot & KVM_PGTABLE_PROT_W) && !ret)
 		mark_page_dirty_in_slot(kvm, s2fd->memslot, gfn);
 
-	return ret != -EAGAIN ? ret : 0;
+	if (ret == -EAGAIN)
+		return result ? ret : 0;
+
+	if (result && !ret)
+		result->mapping_size = PAGE_SIZE;
+
+	return ret;
 }
 
 struct kvm_s2_fault_vma_info {
@@ -1779,7 +1820,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 {
 	unsigned int flags = FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE;
 	struct kvm_vcpu *vcpu = s2fd->vcpu;
-	struct kvm_pgtable *pgt = vcpu->arch.hw_mmu->pgt;
+	struct kvm_pgtable *pgt = s2fd->mmu->pgt;
 	struct mm_struct *mm = current->mm;
 	struct kvm *kvm = vcpu->kvm;
 	void *hyp_memcache;
@@ -1787,7 +1828,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	int ret;
 
 	hyp_memcache = get_mmu_memcache(vcpu);
-	ret = topup_mmu_memcache(vcpu, hyp_memcache);
+	ret = topup_mmu_memcache(s2fd->mmu, hyp_memcache);
 	if (ret)
 		return -ENOMEM;
 
@@ -1910,11 +1951,6 @@ static short kvm_s2_resolve_vma_size(const struct kvm_s2_fault_desc *s2fd,
 	return vma_shift;
 }
 
-static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)
-{
-	return kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
-}
-
 static int kvm_s2_fault_get_vma_info(const struct kvm_s2_fault_desc *s2fd,
 				     struct kvm_s2_fault_vma_info *s2vi)
 {
@@ -1980,13 +2016,11 @@ static int kvm_s2_fault_pin_pfn(const struct kvm_s2_fault_desc *s2fd,
 		return ret;
 
 	s2vi->pfn = __kvm_faultin_pfn(s2fd->memslot, get_canonical_gfn(s2fd, s2vi),
-				      kvm_is_write_fault(s2fd->vcpu) ? FOLL_WRITE : 0,
+				      kvm_s2_fault_is_write(s2fd) ? FOLL_WRITE : 0,
 				      &s2vi->map_writable, &s2vi->page);
 	if (unlikely(is_error_noslot_pfn(s2vi->pfn))) {
-		if (s2vi->pfn == KVM_PFN_ERR_HWPOISON) {
-			kvm_send_hwpoison_signal(s2fd->hva, __ffs(s2vi->vma_pagesize));
-			return 0;
-		}
+		if (s2vi->pfn == KVM_PFN_ERR_HWPOISON)
+			return -EHWPOISON;
 		return -EFAULT;
 	}
 
@@ -2038,7 +2072,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 {
 	struct kvm *kvm = s2fd->vcpu->kvm;
 
-	if (kvm_vcpu_trap_is_exec_fault(s2fd->vcpu) && s2vi->map_non_cacheable)
+	if (kvm_s2_fault_is_exec(s2fd) && s2vi->map_non_cacheable)
 		return -ENOEXEC;
 
 	/*
@@ -2047,7 +2081,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 	 * and trigger the exception here. Since the memslot is valid, inject
 	 * the fault back to the guest.
 	 */
-	if (esr_fsc_is_excl_atomic_fault(kvm_vcpu_get_esr(s2fd->vcpu))) {
+	if (esr_fsc_is_excl_atomic_fault(s2fd->esr)) {
 		kvm_inject_dabt_excl_atomic(s2fd->vcpu, kvm_vcpu_get_hfar(s2fd->vcpu));
 		return 1;
 	}
@@ -2056,13 +2090,13 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 
 	if (s2vi->map_writable && (s2vi->device ||
 				   !memslot_is_logging(s2fd->memslot) ||
-				   kvm_is_write_fault(s2fd->vcpu)))
+				   kvm_s2_fault_is_write(s2fd)))
 		*prot |= KVM_PGTABLE_PROT_W;
 
 	if (s2fd->nested)
 		*prot = adjust_nested_fault_perms(s2fd->nested, *prot);
 
-	if (kvm_vcpu_trap_is_exec_fault(s2fd->vcpu))
+	if (kvm_s2_fault_is_exec(s2fd))
 		*prot |= KVM_PGTABLE_PROT_X;
 
 	if (s2vi->map_non_cacheable)
@@ -2086,7 +2120,8 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 			    const struct kvm_s2_fault_vma_info *s2vi,
 			    enum kvm_pgtable_prot prot,
-			    void *memcache)
+			    void *memcache,
+			    struct kvm_s2_fault_result *result)
 {
 	enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;
 	struct kvm_guest_s2_mapping *mapping = NULL;
@@ -2100,7 +2135,7 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 	gfn_t gfn;
 	int ret;
 
-	if (kvm_is_nested_s2_mmu(kvm, s2fd->vcpu->arch.hw_mmu)) {
+	if (kvm_is_nested_s2_mmu(kvm, s2fd->mmu)) {
 		mapping = kmalloc_obj(struct kvm_guest_s2_mapping,
 				      GFP_KERNEL_ACCOUNT);
 		if (!mapping) {
@@ -2110,13 +2145,12 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 	}
 
 	kvm_fault_lock(kvm);
-	pgt = s2fd->vcpu->arch.hw_mmu->pgt;
+	pgt = s2fd->mmu->pgt;
 	ret = -EAGAIN;
 	if (mmu_invalidate_retry(kvm, s2vi->mmu_seq))
 		goto out_unlock;
 
-	perm_fault_granule = (kvm_s2_fault_is_perm(s2fd) ?
-			      kvm_vcpu_trap_get_perm_fault_granule(s2fd->vcpu) : 0);
+	perm_fault_granule = kvm_s2_perm_fault_granule(s2fd);
 	mapping_size = s2vi->vma_pagesize;
 	pfn = s2vi->pfn;
 	gfn = s2vi->gfn;
@@ -2188,14 +2222,19 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 		mark_page_dirty_in_slot(kvm, s2fd->memslot,
 					gpa_to_gfn(canonical_ipa));
 
-	if (ret != -EAGAIN)
-		return ret;
-	return 0;
+	if (ret == -EAGAIN)
+		return result ? ret : 0;
+
+	if (result && !ret)
+		result->mapping_size = mapping_size;
+
+	return ret;
 }
 
-static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
+static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd,
+			  struct kvm_s2_fault_result *result)
 {
-	bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
+	bool perm_fault = kvm_s2_fault_is_perm(s2fd);
 	struct kvm_s2_fault_vma_info s2vi = {};
 	enum kvm_pgtable_prot prot;
 	void *memcache;
@@ -2213,7 +2252,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	memcache = get_mmu_memcache(s2fd->vcpu);
 	if (!perm_fault || memslot_is_logging(s2fd->memslot) ||
 	    is_protected_kvm_enabled()) {
-		ret = topup_mmu_memcache(s2fd->vcpu, memcache);
+		ret = topup_mmu_memcache(s2fd->mmu, memcache);
 		if (ret)
 			return ret;
 	}
@@ -2223,6 +2262,13 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	 * get block mapping for device MMIO region.
 	 */
 	ret = kvm_s2_fault_pin_pfn(s2fd, &s2vi);
+	if (ret == -EHWPOISON) {
+		/* If result is specified, let the caller handle this. */
+		if (result)
+			return -EHWPOISON;
+		kvm_send_hwpoison_signal(s2fd->hva, __ffs(s2vi.vma_pagesize));
+		return 0;
+	}
 	if (ret != 1)
 		return ret;
 
@@ -2232,7 +2278,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 		return ret;
 	}
 
-	return kvm_s2_fault_map(s2fd, &s2vi, prot, memcache);
+	return kvm_s2_fault_map(s2fd, &s2vi, prot, memcache, result);
 }
 
 /* Resolve the access fault by making the page young again. */
@@ -2342,7 +2388,8 @@ int kvm_handle_guest_sea(struct kvm_vcpu *vcpu)
 int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 {
 	struct kvm_s2_trans nested_trans, *nested = NULL;
-	unsigned long esr;
+	unsigned long esr = kvm_vcpu_get_esr(vcpu);
+	struct kvm_s2_mmu *mmu = vcpu->arch.hw_mmu;
 	phys_addr_t fault_ipa; /* The address we faulted on */
 	phys_addr_t ipa; /* Always the IPA in the L1 guest phys space */
 	struct kvm_memory_slot *memslot;
@@ -2351,11 +2398,9 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	gfn_t gfn;
 	int ret, idx;
 
-	if (kvm_vcpu_abt_issea(vcpu))
+	if (esr_abt_is_sea(esr))
 		return kvm_handle_guest_sea(vcpu);
 
-	esr = kvm_vcpu_get_esr(vcpu);
-
 	/*
 	 * The fault IPA should be reliable at this point as we're not dealing
 	 * with an SEA.
@@ -2364,7 +2409,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	if (KVM_BUG_ON(ipa == INVALID_GPA, vcpu->kvm))
 		return -EFAULT;
 
-	is_iabt = kvm_vcpu_trap_is_iabt(vcpu);
+	is_iabt = esr_trap_is_iabt(esr);
 
 	if (esr_fsc_is_translation_fault(esr)) {
 		/* Beyond sanitised PARange (which is the IPA limit) */
@@ -2374,14 +2419,14 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		}
 
 		/* Falls between the IPA range and the PARange? */
-		if (fault_ipa >= BIT_ULL(VTCR_EL2_IPA(vcpu->arch.hw_mmu->vtcr))) {
+		if (fault_ipa >= BIT_ULL(VTCR_EL2_IPA(mmu->vtcr))) {
 			fault_ipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(vcpu));
 
 			return kvm_inject_sea(vcpu, is_iabt, fault_ipa);
 		}
 	}
 
-	trace_kvm_guest_fault(*vcpu_pc(vcpu), kvm_vcpu_get_esr(vcpu),
+	trace_kvm_guest_fault(*vcpu_pc(vcpu), esr,
 			      kvm_vcpu_get_hfar(vcpu), fault_ipa);
 
 	/* Check the stage-2 fault is trans. fault or write fault */
@@ -2389,10 +2434,10 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	    !esr_fsc_is_permission_fault(esr) &&
 	    !esr_fsc_is_access_flag_fault(esr) &&
 	    !esr_fsc_is_excl_atomic_fault(esr)) {
-		kvm_err("Unsupported FSC: EC=%#x xFSC=%#lx ESR_EL2=%#lx\n",
-			kvm_vcpu_trap_get_class(vcpu),
-			(unsigned long)kvm_vcpu_trap_get_fault(vcpu),
-			(unsigned long)kvm_vcpu_get_esr(vcpu));
+		kvm_err("Unsupported FSC: EC=%#lx xFSC=%#lx ESR_EL2=%#lx\n",
+			ESR_ELx_EC(esr),
+			(unsigned long)(esr & ESR_ELx_FSC),
+			(unsigned long)esr);
 		return -EFAULT;
 	}
 
@@ -2411,8 +2456,8 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	 * nothing to walk and we treat it as a 1:1 before going through the
 	 * canonical translation.
 	 */
-	if (kvm_is_nested_s2_mmu(vcpu->kvm,vcpu->arch.hw_mmu) &&
-	    vcpu->arch.hw_mmu->nested_stage2_enabled) {
+	if (kvm_is_nested_s2_mmu(vcpu->kvm, mmu) &&
+	    mmu->nested_stage2_enabled) {
 		u32 esr;
 
 		ret = kvm_walk_nested_s2(vcpu, fault_ipa, &nested_trans);
@@ -2441,7 +2486,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	gfn = ipa >> PAGE_SHIFT;
 	memslot = gfn_to_memslot(vcpu->kvm, gfn);
 	hva = gfn_to_hva_memslot_prot(memslot, gfn, &writable);
-	write_fault = kvm_is_write_fault(vcpu);
+	write_fault = esr_abt_is_write_fault(esr);
 	if (kvm_is_error_hva(hva) || (write_fault && !writable)) {
 		/*
 		 * The guest has put either its instructions or its page-tables
@@ -2454,7 +2499,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 			goto out;
 		}
 
-		if (kvm_vcpu_abt_iss1tw(vcpu)) {
+		if (esr_abt_is_s1ptw(esr)) {
 			ret = kvm_inject_sea_dabt(vcpu, kvm_vcpu_get_hfar(vcpu));
 			goto out_unlock;
 		}
@@ -2469,7 +2514,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		 * So let's assume that the guest is just being
 		 * cautious, and skip the instruction.
 		 */
-		if (kvm_is_error_hva(hva) && kvm_vcpu_dabt_is_cm(vcpu)) {
+		if (kvm_is_error_hva(hva) && esr_dabt_is_cm(esr)) {
 			kvm_incr_pc(vcpu);
 			ret = 1;
 			goto out_unlock;
@@ -2487,7 +2532,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	}
 
 	/* Userspace should not be able to register out-of-bounds IPAs */
-	VM_BUG_ON(ipa >= kvm_phys_size(vcpu->arch.hw_mmu));
+	VM_BUG_ON(ipa >= kvm_phys_size(mmu));
 
 	if (esr_fsc_is_access_flag_fault(esr)) {
 		handle_access_fault(vcpu, fault_ipa);
@@ -2501,19 +2546,20 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		.nested		= nested,
 		.memslot	= memslot,
 		.hva		= hva,
+		.esr		= esr,
+		.mmu		= mmu,
 	};
 
 	if (kvm_vm_is_protected(vcpu->kvm)) {
 		ret = pkvm_mem_abort(&s2fd);
 	} else {
-		VM_WARN_ON_ONCE(kvm_vcpu_trap_is_permission_fault(vcpu) &&
-				!write_fault &&
-				!kvm_vcpu_trap_is_exec_fault(vcpu));
+		VM_WARN_ON_ONCE(kvm_s2_fault_is_perm(&s2fd) && !write_fault &&
+				!kvm_s2_fault_is_exec(&s2fd));
 
 		if (kvm_slot_has_gmem(memslot))
-			ret = gmem_abort(&s2fd);
+			ret = gmem_abort(&s2fd, NULL);
 		else
-			ret = user_mem_abort(&s2fd);
+			ret = user_mem_abort(&s2fd, NULL);
 	}
 
 	if (ret == 0)
@@ -2890,3 +2936,146 @@ void kvm_toggle_cache(struct kvm_vcpu *vcpu, bool was_enabled)
 
 	trace_kvm_toggle_cache(*vcpu_pc(vcpu), was_enabled, now_enabled);
 }
+
+/*
+ * Try to walk to the specified GPA in canonical mmu - if unmapped returns 0, if
+ * mapped returns the granule size, otherwise returns an error.
+ */
+static long kvm_walk_s2(struct kvm_pgtable *pgt,
+			gpa_t gpa, s8 *level)
+{
+	struct kvm *kvm = kvm_s2_mmu_to_kvm(pgt->mmu);
+	kvm_pte_t pte;
+	long ret;
+
+	guard(read_lock)(&kvm->mmu_lock);
+
+	ret = kvm_pgtable_get_leaf(pgt, gpa, &pte, level,
+				   KVM_PGTABLE_WALK_SHARED);
+	if (ret)
+		return ret;
+	/* Unpopulated, must fault. */
+	if (!kvm_pte_valid(pte))
+		return 0;
+	return kvm_granule_size(*level);
+}
+
+/* Synthesised data abort at specified page table level. */
+#define PRE_FAULT_ESR(level)				\
+	 ((ESR_ELx_EC_DABT_LOW << ESR_ELx_EC_SHIFT) |	\
+	  ESR_ELx_IL | ESR_ELx_FSC_FAULT_L(level))
+
+/* Retrieve either a read-only or a read/write hva. */
+static hva_t gfn_to_hva_memslot_read(struct kvm_memory_slot *slot, gfn_t gfn)
+{
+	return gfn_to_hva_memslot_prot(slot, gfn, /*writable=*/NULL);
+}
+
+static long __pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,
+			   gpa_t gpa, struct kvm_memory_slot *memslot, s8 level)
+{
+	const bool is_gmem = kvm_slot_has_gmem(memslot);
+	const gfn_t gfn = gpa_to_gfn(gpa);
+	const hva_t hva = is_gmem ? 0 : gfn_to_hva_memslot_read(memslot, gfn);
+	const struct kvm_s2_fault_desc s2fd = {
+		.vcpu		= vcpu,
+		.fault_ipa	= gpa,
+		.nested		= NULL,
+		.memslot	= memslot,
+		.hva		= hva,
+		.esr		= PRE_FAULT_ESR(level),
+		.mmu		= mmu,
+	};
+	struct kvm_s2_fault_result result = {};
+	long ret;
+
+	if (kvm_is_error_hva(hva))
+		return -EFAULT;
+
+	if (is_gmem)
+		ret = gmem_abort(&s2fd, &result);
+	else
+		ret = user_mem_abort(&s2fd, &result);
+	if (IS_ERR_VALUE(ret))
+		return ret;
+	return result.mapping_size;
+}
+
+static long pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,
+			 gpa_t gpa, struct kvm_memory_slot *memslot)
+{
+	s8 level;
+	long ret;
+
+	/* Try a walk first. */
+	ret = kvm_walk_s2(mmu->pgt, gpa, &level);
+	if (ret)
+		return ret;
+	/* OK, have to fault page in. */
+	return __pre_fault_s2(mmu, vcpu, gpa, memslot, level);
+}
+
+static unsigned long
+pre_fault_bytes_consumed(gpa_t gpa, unsigned long granule_size,
+			 unsigned long bytes_remaining)
+{
+	/* Granules are always a power-of-2. */
+	const unsigned long granule_bytes_remaining =
+		granule_size - (gpa % granule_size);
+
+	return min(granule_bytes_remaining, bytes_remaining);
+}
+
+/* If you lose the race this many times, time to give up. */
+#define MAX_PRE_FAULT_RETRIES 3
+
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (is_protected_kvm_enabled())
+		return -EOPNOTSUPP;
+	if (!kvm_vcpu_initialized(vcpu))
+		return -ENOEXEC;
+
+	return 0;
+}
+
+/**
+ * kvm_arch_vcpu_pre_fault_memory - pre-fault stage-2 page tables for the
+ * specified GPA.
+ * @vcpu:	The VCPU pointer
+ * @range:	{gpa, size, flags} tuple
+ *
+ * The mapping performed is always best-effort - faulting in is necessarily
+ * racey. The ranges faulted in are canonical, nested page tables are ignored.
+ *
+ * @range->gpa specifies the GPA to pre-fault, @range->size specifies how many
+ * bytes remain to be pre-faulted and @range->flags is reserved and must be 0.
+ *
+ * Returns: the number of bytes the pre-fault consumed, or an error.
+ */
+long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
+				    struct kvm_pre_fault_memory *range)
+{
+	struct kvm *kvm = vcpu->kvm;
+	const u64 bytes_remaining = range->size;
+	struct kvm_s2_mmu *mmu = &kvm->arch.mmu; /* Canonical. */
+	struct kvm_memory_slot *memslot;
+	const gpa_t gpa = range->gpa;
+	int num_retries = 0;
+	long ret;
+
+	memslot = gfn_to_memslot(kvm, gpa_to_gfn(gpa));
+	if (!memslot)
+		return -ENOENT;
+	/* SRCU must be released for progress and only userland can do that. */
+	if (memslot->flags & KVM_MEMSLOT_INVALID)
+		return -EAGAIN;
+
+	do {
+		ret = pre_fault_s2(mmu, vcpu, gpa, memslot);
+	} while (ret == -EAGAIN && num_retries++ < MAX_PRE_FAULT_RETRIES);
+
+	if (IS_ERR_VALUE(ret))
+		return ret;
+	return pre_fault_bytes_consumed(gpa, ret, bytes_remaining);
+}
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index ec754a865a006..cd868af63b1b8 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -662,7 +662,7 @@ static u8 get_guest_mapping_ttl(struct kvm_s2_mmu *mmu, u64 addr)
 		return 0;
 
 	tmp &= ~(sz - 1);
-	if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL))
+	if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL, 0))
 		goto again;
 	if (!(pte & PTE_VALID))
 		goto again;
diff --git a/arch/s390/kvm/s390/s390.c b/arch/s390/kvm/s390/s390.c
index 5c73f43782a74..47fe032444f45 100644
--- a/arch/s390/kvm/s390/s390.c
+++ b/arch/s390/kvm/s390/s390.c
@@ -5784,6 +5784,14 @@ void kvm_arch_commit_memory_region(struct kvm *kvm, struct kvm_memory_slot *old,
 	s390_kvm_mmu_commit_memory_region(kvm, old, new, change);
 }
 
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (kvm_is_ucontrol(vcpu->kvm))
+		return -EINVAL;
+
+	return 0;
+}
+
 /**
  * kvm_arch_vcpu_pre_fault_memory() -- pre-fault and link gmap dat tables
  * @vcpu: the vcpu that shall appear to have generated the fault-in.
@@ -5810,9 +5818,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, struct kvm_pre_fault_
 	gpa_t end;
 	int rc;
 
-	if (kvm_is_ucontrol(vcpu->kvm))
-		return -EINVAL;
-
 	rc = kvm_s390_faultin_gfn(vcpu, NULL, &f);
 	if (rc == PGM_ADDRESSING)
 		return -ENOENT;
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 064ecc33b9267..c35fd2868c20b 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -5086,6 +5086,14 @@ static int kvm_tdp_page_prefault(struct kvm_vcpu *vcpu, gpa_t gpa,
 	}
 }
 
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (!vcpu->kvm->arch.pre_fault_allowed)
+		return -EOPNOTSUPP;
+
+	return 0;
+}
+
 long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				    struct kvm_pre_fault_memory *range)
 {
@@ -5095,9 +5103,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 	u64 end;
 	int r;
 
-	if (!vcpu->kvm->arch.pre_fault_allowed)
-		return -EOPNOTSUPP;
-
 	if (kvm_is_gfn_alias(vcpu->kvm, gpa_to_gfn(range->gpa)))
 		return -EINVAL;
 
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 03bfc92864b6e..39a4d345aeca8 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -2643,6 +2643,7 @@ void kvm_arch_gmem_invalidate_range(struct kvm *kvm, struct kvm_gfn_range *range
 #endif
 
 #ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu);
 long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				    struct kvm_pre_fault_memory *range);
 #endif
diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 6a1482e3a286b..908bdc7cf4f58 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -178,6 +178,7 @@ TEST_GEN_PROGS_arm64 += arm64/debug-exceptions
 TEST_GEN_PROGS_arm64 += arm64/hello_el2
 TEST_GEN_PROGS_arm64 += arm64/host_sve
 TEST_GEN_PROGS_arm64 += arm64/hypercalls
+TEST_GEN_PROGS_arm64 += arm64/nv_pre_fault_memory_test
 TEST_GEN_PROGS_arm64 += arm64/external_aborts
 TEST_GEN_PROGS_arm64 += arm64/mmio_sign_ext
 TEST_GEN_PROGS_arm64 += arm64/page_fault_test
@@ -205,6 +206,7 @@ TEST_GEN_PROGS_arm64 += guest_memfd_test
 TEST_GEN_PROGS_arm64 += mmu_stress_test
 TEST_GEN_PROGS_arm64 += rseq_test
 TEST_GEN_PROGS_arm64 += steal_time
+TEST_GEN_PROGS_arm64 += pre_fault_memory_test
 
 TEST_GEN_PROGS_s390 = $(TEST_GEN_PROGS_COMMON)
 TEST_GEN_PROGS_s390 += s390/memop
diff --git a/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c
new file mode 100644
index 0000000000000..09c1db3038566
--- /dev/null
+++ b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c
@@ -0,0 +1,158 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * nv_pre_fault_memory_test - Test KVM_PRE_FAULT_MEMORY on a vCPU whose
+ * last-run context is nested.
+ *
+ * The guest enters vEL2, sets up its EL2 translation configuration into the
+ * real EL1 registers then ERETs to vEL1 and exits to userspace so its vCPU
+ * last-run context is nested backed by a shadow stage 2 MMU.
+ *
+ * Assert that pre-faulting ignores that and targets the canonical stage-2
+ * page tables only.
+ */
+#include "kvm_util.h"
+#include "processor.h"
+#include "test_util.h"
+#include "ucall.h"
+
+#include <asm/sysreg.h>
+#include <linux/sizes.h>
+
+#define TEST_MEM_SLOT		10
+#define TEST_MEM_SIZE		SZ_2M
+#define TEST_MEM_GPA		SZ_1G
+
+static void guest_el1_code(void)
+{
+	u64 offset;
+
+	GUEST_ASSERT_EQ(get_current_el(), 1);
+
+	/* Exit to userspace with the vEL1 (nested) context live. */
+	GUEST_SYNC(1);
+
+	/*
+	 * Touch the prefaulted range. vstage-2 is disabled, so the shadow
+	 * stage-2 is a 1:1 view of the canonical IPA space.
+	 */
+	for (offset = 0; offset < TEST_MEM_SIZE; offset += SZ_4K)
+		READ_ONCE(*(u64 *)(TEST_MEM_GPA + offset));
+
+	GUEST_DONE();
+}
+
+static void guest_code(void)
+{
+	u64 sp;
+
+	GUEST_ASSERT_EQ(get_current_el(), 2);
+
+	/*
+	 * Mirror the EL2 translation regime into the real EL1 registers so
+	 * that vEL1 runs on the test's stage-1 page tables. With E2H=1, the
+	 * _EL1 accessors read the EL2 registers, and the _EL12 accessors
+	 * write the real EL1 registers.
+	 */
+	write_sysreg_s(read_sysreg(sctlr_el1), SYS_SCTLR_EL12);
+	write_sysreg_s(read_sysreg(tcr_el1), SYS_TCR_EL12);
+	write_sysreg_s(read_sysreg(ttbr0_el1), SYS_TTBR0_EL12);
+	write_sysreg_s(read_sysreg(mair_el1), SYS_MAIR_EL12);
+	write_sysreg_s(read_sysreg(cpacr_el1), SYS_CPACR_EL12);
+
+	/* Run vEL1 on the same stack. */
+	asm volatile("mov %0, sp" : "=r"(sp));
+	write_sysreg(sp, sp_el1);
+
+	/*
+	 * Drop TGE so that vEL1 is a nested context rather than host EL0.
+	 * KVM backs it with a shadow stage-2 MMU even though vstage-2 is
+	 * disabled (HCR_EL2.VM=0).
+	 */
+	write_sysreg(read_sysreg(hcr_el2) & ~HCR_EL2_TGE, hcr_el2);
+	isb();
+
+	write_sysreg(PSR_MODE_EL1h | PSR_F_BIT | PSR_I_BIT | PSR_A_BIT |
+		     PSR_D_BIT, spsr_el2);
+	write_sysreg((u64)guest_el1_code, elr_el2);
+	asm volatile("eret");
+
+	GUEST_ASSERT(false);
+}
+
+static void pre_fault(struct kvm_vcpu *vcpu, u64 gpa, u64 size)
+{
+	struct kvm_pre_fault_memory range = {
+		.gpa = gpa,
+		.size = size,
+	};
+	int ret;
+
+	do {
+		ret = __vcpu_ioctl(vcpu, KVM_PRE_FAULT_MEMORY, &range);
+	} while ((!ret && range.size) ||
+		 (ret < 0 && (errno == EINTR || errno == EAGAIN)));
+
+	TEST_ASSERT(!ret, "KVM_PRE_FAULT_MEMORY failed, ret: %d errno: %d",
+		    ret, errno);
+	TEST_ASSERT_EQ(range.size, 0);
+}
+
+int main(void)
+{
+	struct kvm_vcpu_init init;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	struct ucall uc;
+	u64 npages;
+
+	TEST_REQUIRE(test_supports_el2());
+	TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));
+
+	vm = vm_create(1);
+
+	kvm_get_default_vcpu_target(vm, &init);
+	init.features[0] |= BIT(KVM_ARM_VCPU_HAS_EL2);
+	vcpu = aarch64_vcpu_add(vm, 0, &init, guest_code);
+	kvm_arch_vm_finalize_vcpus(vm);
+
+	npages = TEST_MEM_SIZE / vm->page_size;
+	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, TEST_MEM_GPA,
+				    TEST_MEM_SLOT, npages, 0);
+	virt_map(vm, TEST_MEM_GPA, TEST_MEM_GPA, npages);
+
+	/* Run the guest until it has ERET'd from vEL2 to vEL1. */
+	vcpu_run(vcpu);
+	switch (get_ucall(vcpu, &uc)) {
+	case UCALL_SYNC:
+		TEST_ASSERT_EQ(uc.args[1], 1);
+		break;
+	case UCALL_ABORT:
+		REPORT_GUEST_ASSERT(uc);
+		break;
+	default:
+		TEST_FAIL("Unhandled ucall: %ld", uc.cmd);
+	}
+
+	/*
+	 * The vCPU's last-run context is vEL1, so its hw_mmu is a shadow
+	 * stage-2 MMU.
+	 *
+	 * Pre-faulting must ignore that and populate the canonical stage-2.
+	 */
+	pre_fault(vcpu, TEST_MEM_GPA, TEST_MEM_SIZE);
+
+	/* Resume at vEL1 and touch the prefaulted range. */
+	vcpu_run(vcpu);
+	switch (get_ucall(vcpu, &uc)) {
+	case UCALL_DONE:
+		break;
+	case UCALL_ABORT:
+		REPORT_GUEST_ASSERT(uc);
+		break;
+	default:
+		TEST_FAIL("Unhandled ucall: %ld", uc.cmd);
+	}
+
+	kvm_vm_free(vm);
+	return 0;
+}
diff --git a/tools/testing/selftests/kvm/pre_fault_memory_test.c b/tools/testing/selftests/kvm/pre_fault_memory_test.c
index c57631aab3d38..3082fe09b95fa 100644
--- a/tools/testing/selftests/kvm/pre_fault_memory_test.c
+++ b/tools/testing/selftests/kvm/pre_fault_memory_test.c
@@ -12,19 +12,29 @@
 #include <processor.h>
 #include <pthread.h>
 #include <ucall_common.h>
+#include <guest_modes.h>
 
 /* Arbitrarily chosen values */
-#define TEST_SIZE		(SZ_2M + PAGE_SIZE)
-#define TEST_NPAGES		(TEST_SIZE / PAGE_SIZE)
+#define TEST_BASE_SIZE		SZ_2M
 #define TEST_SLOT		10
 
+/* Storage of test info to share with guest code */
+struct test_config {
+	u64 page_size;
+	u64 test_size;
+	u64 test_num_pages;
+};
+
+static struct test_config test_config;
+
 static void guest_code(u64 base_gva)
 {
 	volatile u64 val __used;
+	struct test_config *config = &test_config;
 	int i;
 
-	for (i = 0; i < TEST_NPAGES; i++) {
-		u64 *src = (u64 *)(base_gva + i * PAGE_SIZE);
+	for (i = 0; i < config->test_num_pages; i++) {
+		u64 *src = (u64 *)(base_gva + i * config->page_size);
 
 		val = *src;
 	}
@@ -36,6 +46,7 @@ struct slot_worker_data {
 	struct kvm_vm *vm;
 	gpa_t gpa;
 	u32 flags;
+	enum vm_mem_backing_src_type mem_backing_src;
 	bool worker_ready;
 	bool prefault_ready;
 	bool recreate_slot;
@@ -56,14 +67,16 @@ static void *delete_slot_worker(void *__data)
 	while (!READ_ONCE(data->recreate_slot))
 		cpu_relax();
 
-	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, data->gpa,
-				    TEST_SLOT, TEST_NPAGES, data->flags);
+	vm_userspace_mem_region_add(vm, data->mem_backing_src, data->gpa,
+				    TEST_SLOT, test_config.test_num_pages, data->flags);
 
 	return NULL;
 }
 
 static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
-			     u64 size, u64 expected_left, bool private)
+			     u64 size, u64 expected_left,
+			     enum vm_mem_backing_src_type mem_backing_src,
+			     bool private)
 {
 	struct kvm_pre_fault_memory range = {
 		.gpa = base_gpa + offset,
@@ -74,6 +87,7 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 		.vm = vcpu->vm,
 		.gpa = base_gpa,
 		.flags = private ? KVM_MEM_GUEST_MEMFD : 0,
+		.mem_backing_src = mem_backing_src,
 	};
 	bool slot_recreated = false;
 	pthread_t slot_worker;
@@ -150,8 +164,8 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 	/*
 	 * Assert success if prefaulting the entire range should succeed, i.e.
 	 * complete with no bytes remaining.  Otherwise prefaulting should have
-	 * failed due to ENOENT (due to RET_PF_EMULATE for emulated MMIO when
-	 * no memslot exists).
+	 * failed due to ENOENT (no memslot exists for the GPA; on x86 this
+	 * surfaces via RET_PF_EMULATE).
 	 */
 	if (!expected_left)
 		TEST_ASSERT_VM_VCPU_IOCTL(!ret, KVM_PRE_FAULT_MEMORY, ret, vcpu->vm);
@@ -160,39 +174,82 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 					  KVM_PRE_FAULT_MEMORY, ret, vcpu->vm);
 }
 
-static void __test_pre_fault_memory(unsigned long vm_type, bool private)
+struct test_params {
+	unsigned long vm_type;
+	bool private;
+	enum vm_mem_backing_src_type mem_backing_src;
+};
+
+static void __test_pre_fault_memory(enum vm_guest_mode guest_mode, void *arg)
 {
-	gpa_t gpa, gva, alignment, guest_page_size;
+	gpa_t gpa, gva, alignment, guest_page_size, host_page_size;
+	gpa_t backing_src_pagesz, mem_page_size;
+	struct test_params *p = arg;
 	const struct vm_shape shape = {
-		.mode = VM_MODE_DEFAULT,
-		.type = vm_type,
+		.mode = guest_mode,
+		.type = p->vm_type,
 	};
 	struct kvm_vcpu *vcpu;
+	struct kvm_run *run;
 	struct kvm_vm *vm;
 	struct ucall uc;
 
+	pr_info("Testing guest mode: %s\n", vm_guest_mode_string(guest_mode));
+	pr_info("Testing memory backing src type: %s\n",
+		vm_mem_backing_src_alias(p->mem_backing_src)->name);
+
 	vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code);
 
-	alignment = guest_page_size = vm_guest_mode_params[VM_MODE_DEFAULT].page_size;
-	gpa = (vm->max_gfn - TEST_NPAGES) * guest_page_size;
+	guest_page_size = vm_guest_mode_params[guest_mode].page_size;
+	host_page_size = getpagesize();
+	backing_src_pagesz = get_backing_src_pagesz(p->mem_backing_src);
+	mem_page_size = max(host_page_size, backing_src_pagesz);
+
+	test_config.page_size = guest_page_size;
+	test_config.test_size = align_up(TEST_BASE_SIZE + test_config.page_size,
+					 mem_page_size);
+	test_config.test_num_pages = vm_calc_num_guest_pages(vm->mode, test_config.test_size);
+
+	gpa = (vm->max_gfn - test_config.test_num_pages) * test_config.page_size;
 	alignment = SZ_2M;
+	alignment = max(alignment, mem_page_size);
 	gpa = align_down(gpa, alignment);
 	gva = gpa & ((1ULL << (vm->va_bits - 1)) - 1);
 
-	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa, TEST_SLOT,
-				    TEST_NPAGES, private ? KVM_MEM_GUEST_MEMFD : 0);
-	virt_map(vm, gva, gpa, TEST_NPAGES);
+	vm_userspace_mem_region_add(vm, p->mem_backing_src,
+				    gpa, TEST_SLOT, test_config.test_num_pages,
+				    p->private ? KVM_MEM_GUEST_MEMFD : 0);
+	virt_map(vm, gva, gpa, test_config.test_num_pages);
+
+	if (p->private)
+		vm_mem_set_private(vm, gpa, test_config.test_size);
+
+	pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,
+			 p->mem_backing_src, p->private);
+	/* Retry the same range after the first prefault attempt. */
+	pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,
+			 p->mem_backing_src, p->private);
+	pre_fault_memory(vcpu, gpa,
+			 test_config.test_size - host_page_size,
+			 host_page_size * 2, host_page_size,
+			 p->mem_backing_src, p->private);
+	pre_fault_memory(vcpu, gpa, test_config.test_size,
+			 host_page_size, host_page_size,
+			 p->mem_backing_src, p->private);
 
-	if (private)
-		vm_mem_set_private(vm, gpa, TEST_SIZE);
+	vcpu_args_set(vcpu, 1, gva);
 
-	pre_fault_memory(vcpu, gpa, 0, SZ_2M, 0, private);
-	pre_fault_memory(vcpu, gpa, SZ_2M, PAGE_SIZE * 2, PAGE_SIZE, private);
-	pre_fault_memory(vcpu, gpa, TEST_SIZE, PAGE_SIZE, PAGE_SIZE, private);
+	/* Export the shared variables to the guest. */
+	sync_global_to_guest(vm, test_config);
 
-	vcpu_args_set(vcpu, 1, gva);
 	vcpu_run(vcpu);
 
+	run = vcpu->run;
+	TEST_ASSERT(run->exit_reason == UCALL_EXIT_REASON,
+		    "Wanted %s, got exit reason: %u (%s)",
+		    exit_reason_str(UCALL_EXIT_REASON),
+		    run->exit_reason, exit_reason_str(run->exit_reason));
+
 	switch (get_ucall(vcpu, &uc)) {
 	case UCALL_ABORT:
 		REPORT_GUEST_ASSERT(uc);
@@ -207,24 +264,61 @@ static void __test_pre_fault_memory(unsigned long vm_type, bool private)
 	kvm_vm_free(vm);
 }
 
-static void test_pre_fault_memory(unsigned long vm_type, bool private)
+static void test_pre_fault_memory(unsigned long vm_type, enum vm_mem_backing_src_type backing_src,
+				  bool private)
 {
+	struct test_params p = {
+		.vm_type = vm_type,
+		.private = private,
+		.mem_backing_src = backing_src,
+	};
+
 	if (vm_type && !(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(vm_type))) {
 		pr_info("Skipping tests for vm_type 0x%lx\n", vm_type);
 		return;
 	}
 
-	__test_pre_fault_memory(vm_type, private);
+	for_each_guest_mode(__test_pre_fault_memory, &p);
+}
+
+static void help(char *name)
+{
+	puts("");
+	printf("usage: %s [-h] [-m mode] [-s mem-type]\n", name);
+	puts("");
+	guest_modes_help();
+	backing_src_help("-s");
+	puts("");
 }
 
 int main(int argc, char *argv[])
 {
+	enum vm_mem_backing_src_type backing = DEFAULT_VM_MEM_SRC;
+	int opt;
+
+	guest_modes_append_default();
+
+	while ((opt = getopt(argc, argv, "hm:s:")) != -1) {
+		switch (opt) {
+		case 'm':
+			guest_modes_cmdline(optarg);
+			break;
+		case 's':
+			backing = parse_backing_src_type(optarg);
+			break;
+		case 'h':
+		default:
+			help(argv[0]);
+			exit(0);
+		}
+	}
+
 	TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));
 
-	test_pre_fault_memory(0, false);
+	test_pre_fault_memory(0, backing, false);
 #ifdef __x86_64__
-	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, false);
-	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, true);
+	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, false);
+	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, true);
 #endif
 	return 0;
 }
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 65eb26a0520d8..41356577bee79 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -4353,7 +4353,7 @@ static int kvm_vcpu_ioctl_get_stats_fd(struct kvm_vcpu *vcpu)
 static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				     struct kvm_pre_fault_memory *range)
 {
-	int idx;
+	int idx, err;
 	long r;
 	u64 full_size;
 
@@ -4365,6 +4365,14 @@ static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 	    range->gpa + range->size <= range->gpa)
 		return -EINVAL;
 
+	/*
+	 * Certain architectures (e.g. arm64) need to reject the ioctl 'early'
+	 * before vcpu_load().
+	 */
+	err = kvm_arch_pre_fault_allowed(vcpu);
+	if (err)
+		return err;
+
 	vcpu_load(vcpu);
 	idx = srcu_read_lock(&vcpu->kvm->srcu);
 

3/2 2026/09/25 22:21 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 25241
							output: 40
							thoughts: 1708
						

Thoughts:
**Analyzing the Patch for Fuzzing Potential on amd64**

Okay, let's break this down. First, the patch is specifically targeting the amd64 architecture, which immediately piques my interest from a fuzzing perspective. The key here is to determine *where* it touches amd64 code and what functionality it introduces or modifies.

I see a flurry of files modified, a lot of ARM64 specific files but... wait, there are a few that are relevant to x86/amd64: `virt/kvm/kvm_main.c` and `arch/x86/kvm/mmu/mmu.c`. My spidey-sense is tingling. These are KVM core files, so they are critical.

Looking at `virt/kvm/kvm_main.c`, the focus seems to be on a function called `kvm_vcpu_pre_fault_memory`.  It appears the patch inserts a call to `kvm_arch_pre_fault_allowed(vcpu)` *before* the `vcpu_load(vcpu)` call. And also before the call to `srcu_read_lock(&vcpu->kvm->srcu)`. This is a crucial change in execution order.

Then, inspecting `arch/x86/kvm/mmu/mmu.c`, `kvm_arch_pre_fault_allowed` is defined. This function simply checks `!vcpu->kvm->arch.pre_fault_allowed` and returns an error if `pre_fault_allowed` is not set.

The original intent of the patch seems to be to implement `KVM_PRE_FAULT_MEMORY` for ARM64, as I can see from the file structure, and the code changes. Now, it looks like, x86 architecture is being touched. The x86 change seems to have just added the call to the function and implementation of `kvm_arch_pre_fault_allowed` function.

This is a good spot for fuzzing! The `kvm_vcpu_pre_fault_memory` function, which now calls the newly introduced `kvm_arch_pre_fault_allowed`, is exposed via a KVM ioctl: `KVM_PRE_FAULT_MEMORY`. Now, a new pre-condition check has been introduced before any locking or arch-specific code gets called.  This could be a valuable point of entry.  By crafting specific inputs to this ioctl, we could potentially trigger interesting code paths and edge cases, especially given the change in execution order. It looks like it is now being called earlier in the ioctl handling.  This is a good candidate for focused fuzzing.



4/2 2026/09/25 22:22 tool read-file
0m Args:
{
  "File": "virt/kvm/kvm_main.c",
  "FirstLine": 4350,
  "LineCount": 50
}

Results:
{
  "Contents": "4350:\t}\n4351:\t\n4352:\t#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\n4353:\tstatic int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n4354:\t\t\t\t\t     struct kvm_pre_fault_memory *range)\n4355:\t{\n4356:\t\tint idx, err;\n4357:\t\tlong r;\n4358:\t\tu64 full_size;\n4359:\t\n4360:\t\tif (range-\u003eflags)\n4361:\t\t\treturn -EINVAL;\n4362:\t\n4363:\t\tif (!PAGE_ALIGNED(range-\u003egpa) ||\n4364:\t\t    !PAGE_ALIGNED(range-\u003esize) ||\n4365:\t\t    range-\u003egpa + range-\u003esize \u003c= range-\u003egpa)\n4366:\t\t\treturn -EINVAL;\n4367:\t\n4368:\t\t/*\n4369:\t\t * Certain architectures (e.g. arm64) need to reject the ioctl 'early'\n4370:\t\t * before vcpu_load().\n4371:\t\t */\n4372:\t\terr = kvm_arch_pre_fault_allowed(vcpu);\n4373:\t\tif (err)\n4374:\t\t\treturn err;\n4375:\t\n4376:\t\tvcpu_load(vcpu);\n4377:\t\tidx = srcu_read_lock(\u0026vcpu-\u003ekvm-\u003esrcu);\n4378:\t\n4379:\t\tfull_size = range-\u003esize;\n4380:\t\tdo {\n4381:\t\t\tif (signal_pending(current)) {\n4382:\t\t\t\tr = -EINTR;\n4383:\t\t\t\tbreak;\n4384:\t\t\t}\n4385:\t\n4386:\t\t\tr = kvm_arch_vcpu_pre_fault_memory(vcpu, range);\n4387:\t\t\tif (WARN_ON_ONCE(r == 0 || r == -EIO))\n4388:\t\t\t\tbreak;\n4389:\t\n4390:\t\t\tif (r \u003c 0)\n4391:\t\t\t\tbreak;\n4392:\t\n4393:\t\t\trange-\u003esize -= r;\n4394:\t\t\trange-\u003egpa += r;\n4395:\t\t\tcond_resched();\n4396:\t\t} while (range-\u003esize);\n4397:\t\n4398:\t\tsrcu_read_unlock(\u0026vcpu-\u003ekvm-\u003esrcu, idx);\n4399:\t\tvcpu_put(vcpu);\n"
}

5/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 28269
							output: 40
							thoughts: 22
						

Thoughts:
Here's my thought process, summarized as you requested:

**Investigating the Caller of `kvm_vcpu_pre_fault_memory`**

Okay, so I need to understand the execution path that leads to `kvm_vcpu_pre_fault_memory`. It's crucial for performance analysis and potentially debugging memory access issues within the KVM virtualization layer. I need to trace back from that function to see which other kernel components and, ultimately, user-space operations are involved. Let's start by identifying the call sites. I'll need to use my preferred debugging tools (likely `grep` or `objdump` with careful filtering, possibly augmented by source code navigation tools like cscope or ctags) to locate all instances where `kvm_vcpu_pre_fault_memory` is directly invoked. I need to be precise, as there could be similar function names that are not the same function, or even similar variables, that could be misleading. I'll take notes to create a list of the call sites, noting the module and the context (e.g., function name, line number) for each. This initial list will be the foundation for constructing the higher-level call chain.



6/2 2026/09/25 22:22 tool read-file
0m Args:
{
  "File": "virt/kvm/kvm_main.c",
  "FirstLine": 4400,
  "LineCount": 40
}

Results:
{
  "Contents": "4400:\t\n4401:\t\t/* Return success if at least one page was mapped successfully.  */\n4402:\t\treturn full_size == range-\u003esize ? r : 0;\n4403:\t}\n4404:\t#endif\n4405:\t\n4406:\tstatic int kvm_wait_for_vcpu_online(struct kvm_vcpu *vcpu)\n4407:\t{\n4408:\t\tstruct kvm *kvm = vcpu-\u003ekvm;\n4409:\t\n4410:\t\t/*\n4411:\t\t * In practice, this happy path will always be taken, as a well-behaved\n4412:\t\t * VMM will never invoke a vCPU ioctl() before KVM_CREATE_VCPU returns.\n4413:\t\t */\n4414:\t\tif (likely(vcpu-\u003evcpu_idx \u003c atomic_read(\u0026kvm-\u003eonline_vcpus)))\n4415:\t\t\treturn 0;\n4416:\t\n4417:\t\t/*\n4418:\t\t * Acquire and release the vCPU's mutex to wait for vCPU creation to\n4419:\t\t * complete (kvm_vm_ioctl_create_vcpu() holds the mutex until the vCPU\n4420:\t\t * is fully online).\n4421:\t\t */\n4422:\t\tif (mutex_lock_killable(\u0026vcpu-\u003emutex))\n4423:\t\t\treturn -EINTR;\n4424:\t\n4425:\t\tmutex_unlock(\u0026vcpu-\u003emutex);\n4426:\t\n4427:\t\tif (WARN_ON_ONCE(!kvm_get_vcpu(kvm, vcpu-\u003evcpu_idx)))\n4428:\t\t\treturn -EIO;\n4429:\t\n4430:\t\treturn 0;\n4431:\t}\n4432:\t\n4433:\tstatic long kvm_vcpu_ioctl(struct file *filp,\n4434:\t\t\t\t   unsigned int ioctl, unsigned long arg)\n4435:\t{\n4436:\t\tstruct kvm_vcpu *vcpu = filp-\u003eprivate_data;\n4437:\t\tvoid __user *argp = (void __user *)arg;\n4438:\t\tint r;\n4439:\t\tstruct kvm_fpu *fpu = NULL;\n"
}

7/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 29204
							output: 40
							thoughts: 9
						

Thoughts:
Okay, here's my interpretation of that thought process, tailored for an expert audience, and written in the first person:

**Diving into the ioctl Handler**

Alright, so the task at hand is to understand how a particular `ioctl` handler is functioning. My immediate thought is, "Let's get right to it." I need to dissect this code; there's no time for preamble. Where's the handler? I need to quickly identify the entry point, the critical functions and data structures involved, and the execution flow. I'm looking for the meat of the operation. Let's see how this thing actually *works*. I'm expecting to find the usual suspects: argument parsing, validation, device access, error handling, and the core device-specific logic. Let's see it in action!



8/2 2026/09/25 22:22 tool read-file
0m Args:
{
  "File": "virt/kvm/kvm_main.c",
  "FirstLine": 4440,
  "LineCount": 50
}

Results:
{
  "Contents": "4440:\t\tstruct kvm_sregs *kvm_sregs = NULL;\n4441:\t\n4442:\t\tif (vcpu-\u003ekvm-\u003emm != current-\u003emm || vcpu-\u003ekvm-\u003evm_dead)\n4443:\t\t\treturn -EIO;\n4444:\t\n4445:\t\tif (unlikely(_IOC_TYPE(ioctl) != KVMIO))\n4446:\t\t\treturn -EINVAL;\n4447:\t\n4448:\t\t/*\n4449:\t\t * Wait for the vCPU to be online before handling the ioctl(), as KVM\n4450:\t\t * assumes the vCPU is reachable via vcpu_array, i.e. may dereference\n4451:\t\t * a NULL pointer if userspace invokes an ioctl() before KVM is ready.\n4452:\t\t */\n4453:\t\tr = kvm_wait_for_vcpu_online(vcpu);\n4454:\t\tif (r)\n4455:\t\t\treturn r;\n4456:\t\n4457:\t\t/*\n4458:\t\t * Let arch code handle select vCPU ioctls without holding vcpu-\u003emutex,\n4459:\t\t * e.g. to support ioctls that can run asynchronous to vCPU execution.\n4460:\t\t */\n4461:\t\tr = kvm_arch_vcpu_unlocked_ioctl(filp, ioctl, arg);\n4462:\t\tif (r != -ENOIOCTLCMD)\n4463:\t\t\treturn r;\n4464:\t\n4465:\t\tif (mutex_lock_killable(\u0026vcpu-\u003emutex))\n4466:\t\t\treturn -EINTR;\n4467:\t\tswitch (ioctl) {\n4468:\t\tcase KVM_RUN: {\n4469:\t\t\tstruct pid *oldpid;\n4470:\t\t\tr = -EINVAL;\n4471:\t\t\tif (arg)\n4472:\t\t\t\tgoto out;\n4473:\t\n4474:\t\t\t/*\n4475:\t\t\t * Note, vcpu-\u003epid is primarily protected by vcpu-\u003emutex. The\n4476:\t\t\t * dedicated r/w lock allows other tasks, e.g. other vCPUs, to\n4477:\t\t\t * read vcpu-\u003epid while this vCPU is in KVM_RUN, e.g. to yield\n4478:\t\t\t * directly to this vCPU\n4479:\t\t\t */\n4480:\t\t\toldpid = vcpu-\u003epid;\n4481:\t\t\tif (unlikely(oldpid != task_pid(current))) {\n4482:\t\t\t\t/* The thread running this VCPU changed. */\n4483:\t\t\t\tstruct pid *newpid;\n4484:\t\n4485:\t\t\t\tr = kvm_arch_vcpu_run_pid_change(vcpu);\n4486:\t\t\t\tif (r)\n4487:\t\t\t\t\tbreak;\n4488:\t\n4489:\t\t\t\tnewpid = get_task_pid(current, PIDTYPE_PID);\n"
}

9/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 30209
							output: 39
							thoughts: 20
						

Thoughts:
Okay, here's my take on those thoughts, tailored for an expert audience:

**Hunting Down KVM_PRE_FAULT_MEMORY**

Alright, let's get down to brass tacks. I need to track down `KVM_PRE_FAULT_MEMORY`. Given the context, my initial hunch is that it's going to be defined and, more importantly, likely *used*, within the core KVM codebase. Specifically, `kvm_main.c` seems like the most logical starting point. It's the central hub for a lot of core KVM functionality. Searching there will be the first and most efficient step. Hopefully I can find the definition, find where it's called and the context in which it operates. Let's dig in.



10/2 2026/09/25 22:22 tool grepper
0m Args:
{
  "Expression": "KVM_PRE_FAULT_MEMORY",
  "PathPrefix": "virt/kvm/kvm_main.c"
}

Results:
{
  "Output": "virt/kvm/kvm_main.c=4433=static long kvm_vcpu_ioctl(struct file *filp,\n--\nvirt/kvm/kvm_main.c-4659-#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\nvirt/kvm/kvm_main.c:4660:\tcase KVM_PRE_FAULT_MEMORY: {\nvirt/kvm/kvm_main.c-4661-\t\tstruct kvm_pre_fault_memory range;\n"
}

11/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 30570
							output: 25
							thoughts: 41
						

Thoughts:
Okay, let's see. My current train of thought leads me to this question: Does the amd64 architecture, specifically, utilize the `CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY` configuration option?

**Investigating KVM Pre-Fault Memory with AMD64**

My first instinct is to dive into the Kconfig files, specifically focusing on the x86 and KVM-related sections. This should be where the relevant configuration options and dependencies are defined. I'll need to meticulously examine these files to ascertain whether `CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY` is employed within the amd64 context, or if there are specific architectural overrides or alternatives in place. This level of detail is crucial for understanding the memory pre-faulting behavior within KVM on x86-64 systems.



12/2 2026/09/25 22:22 tool grepper
0m Args:
{
  "Expression": "KVM_GENERIC_PRE_FAULT_MEMORY"
}

Results:
{
  "Output": "arch/arm64/kvm/Kconfig=20=menuconfig KVM\n--\narch/arm64/kvm/Kconfig-39-\tselect KVM_GUEST_MEMFD\narch/arm64/kvm/Kconfig:40:\tselect KVM_GENERIC_PRE_FAULT_MEMORY\narch/arm64/kvm/Kconfig-41-\thelp\n--\narch/s390/kvm/Kconfig=19=config KVM\n--\narch/s390/kvm/Kconfig-32-\tselect KVM_MMU_LOCKLESS_AGING\narch/s390/kvm/Kconfig:33:\tselect KVM_GENERIC_PRE_FAULT_MEMORY\narch/s390/kvm/Kconfig-34-\thelp\n--\narch/x86/kvm/Kconfig=20=config KVM_X86\n--\narch/x86/kvm/Kconfig-47-\tselect KVM_GENERIC_HARDWARE_ENABLING\narch/x86/kvm/Kconfig:48:\tselect KVM_GENERIC_PRE_FAULT_MEMORY\narch/x86/kvm/Kconfig-49-\tselect KVM_WERROR if WERROR\n--\ninclude/linux/kvm_host.h=2642=void kvm_arch_gmem_invalidate_range(struct kvm *kvm, struct kvm_gfn_range *range);\n--\ninclude/linux/kvm_host.h-2644-\ninclude/linux/kvm_host.h:2645:#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\ninclude/linux/kvm_host.h-2646-int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu);\n--\nvirt/kvm/Kconfig=68=config KVM_GENERIC_DIRTYLOG_READ_PROTECT\n--\nvirt/kvm/Kconfig-70-\nvirt/kvm/Kconfig:71:config KVM_GENERIC_PRE_FAULT_MEMORY\nvirt/kvm/Kconfig-72-       bool\n--\nvirt/kvm/kvm_main.c=4327=static int kvm_vcpu_ioctl_get_stats_fd(struct kvm_vcpu *vcpu)\n--\nvirt/kvm/kvm_main.c-4351-\nvirt/kvm/kvm_main.c:4352:#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\nvirt/kvm/kvm_main.c-4353-static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n--\nvirt/kvm/kvm_main.c=4433=static long kvm_vcpu_ioctl(struct file *filp,\n--\nvirt/kvm/kvm_main.c-4658-\t}\nvirt/kvm/kvm_main.c:4659:#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\nvirt/kvm/kvm_main.c-4660-\tcase KVM_PRE_FAULT_MEMORY: {\n"
}

13/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 31615
							output: 34
							thoughts: 2950
						

Thoughts:
**KVM Code Change Analysis for Fuzzing Eligibility**

Okay, so I'm looking at this patch that seems to be shuffling some code around within KVM, specifically the `KVM_GENERIC_PRE_FAULT_MEMORY` functionality on x86, which, of course, includes amd64. My initial thought is to figure out whether this change warrants fuzzing. I'm seeing that the check `!vcpu->kvm->arch.pre_fault_allowed` has been moved from inside `kvm_arch_vcpu_pre_fault_memory()` to before `vcpu_load(vcpu)` in `kvm_vcpu_pre_fault_memory()`.  This seems to be a refactoring, but I know better than to blindly trust those claims.  Refactorings *can* introduce subtle shifts, and I need to be thorough.

I'm carefully reviewing the provided core triage philosophy. The primary gate is reachability. Is the modified code reachable in standard virtualized environments?  I know that `KVM_PRE_FAULT_MEMORY` is an ioctl on `/dev/kvm`, and that KVM is certainly reachable in QEMU and GCE, especially with syzkaller. The nested virtualization scenario is relevant.  So, yes, it's reachable.

Now, I'm going through the "negative criteria" to see if I can definitively *exclude* the patch from fuzzing. The patch touches the following files: `Documentation/virt/kvm/api.rst` (documentation change), arch-specific code for ARM64 and s390, user-space tests, and core KVM code in `arch/x86/kvm/mmu/mmu.c`, `include/linux/kvm_host.h`, and `virt/kvm/kvm_main.c`. The Documentation and user-space tests are out.  The ARM64 and s390 code is out. But what about the other files?

Is the code in `arch/x86/kvm/mmu/mmu.c` and `virt/kvm/kvm_main.c` exclusive to an architecture other than the target architecture, amd64? No. `arch/x86` is x86_64/amd64. These files are definitely *not* architecture-specific in the sense that they're compiled for a different architecture than the target. Thus, this doesn't apply.

The philosophy clearly states that changes to core KVM subsystems, like these, must be fuzzed if they are reachable. This is what's happening.

This patch refactors the control flow by moving a critical check (`kvm_arch_pre_fault_allowed`) earlier in the function, *before* vCPU loading and locking. That's a significant change! This could lead to a difference in concurrency or locking behavior. The comment referencing ARM64's need to "reject the ioctl 'early'" further underlines the fact that this change isn't trivial.

Given all of this, the answer is clear: the changes to core KVM code are reachable and involve a refactoring that alters control flow. This means that fuzzing is essential. The new check can short circuit the process, and potentially have new races or semantic issues related to state management.



14/2 2026/09/25 22:22 tool grepper
0m Args:
{
  "Expression": "pre_fault_allowed",
  "PathPrefix": "arch/x86/kvm/"
}

Results:
{
  "Output": "arch/x86/kvm/mmu/mmu.c=5045=static int kvm_tdp_page_prefault(struct kvm_vcpu *vcpu, gpa_t gpa,\n--\narch/x86/kvm/mmu/mmu.c-5088-\narch/x86/kvm/mmu/mmu.c:5089:int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)\narch/x86/kvm/mmu/mmu.c-5090-{\narch/x86/kvm/mmu/mmu.c:5091:\tif (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed)\narch/x86/kvm/mmu/mmu.c-5092-\t\treturn -EOPNOTSUPP;\n--\narch/x86/kvm/svm/sev.c=2562=static int snp_launch_finish(struct kvm *kvm, struct kvm_sev_cmd *argp)\n--\narch/x86/kvm/svm/sev.c-2624-\tif (!ret)\narch/x86/kvm/svm/sev.c:2625:\t\tkvm-\u003earch.pre_fault_allowed = true;\narch/x86/kvm/svm/sev.c-2626-\n--\narch/x86/kvm/svm/sev.c=2955=void sev_vm_init(struct kvm *kvm)\n--\narch/x86/kvm/svm/sev.c-2967-\tcase KVM_X86_SEV_VM:\narch/x86/kvm/svm/sev.c:2968:\t\tkvm-\u003earch.pre_fault_allowed = !kvm-\u003earch.has_private_mem;\narch/x86/kvm/svm/sev.c-2969-\t\tto_kvm_sev_info(kvm)-\u003eneed_init = true;\n--\narch/x86/kvm/vmx/tdx.c=1621=static int tdx_mem_page_add(struct kvm *kvm, gfn_t gfn, enum pg_level level,\n--\narch/x86/kvm/vmx/tdx.c-1629-\narch/x86/kvm/vmx/tdx.c:1630:\tif (KVM_BUG_ON(kvm-\u003earch.pre_fault_allowed, kvm) ||\narch/x86/kvm/vmx/tdx.c-1631-\t    KVM_BUG_ON(!kvm_tdx-\u003epage_add_src, kvm))\n--\narch/x86/kvm/vmx/tdx.c=1698=static int tdx_sept_map_leaf_spte(struct kvm *kvm, gfn_t gfn, enum pg_level level,\n--\narch/x86/kvm/vmx/tdx.c-1710-\t/*\narch/x86/kvm/vmx/tdx.c:1711:\t * Ensure pre_fault_allowed is read by kvm_arch_vcpu_pre_fault_memory()\narch/x86/kvm/vmx/tdx.c-1712-\t * before kvm_tdx-\u003estate.  Userspace must not be allowed to pre-fault\n--\narch/x86/kvm/vmx/tdx.c=2874=static int tdx_td_finalize(struct kvm *kvm, struct kvm_tdx_cmd *cmd)\n--\narch/x86/kvm/vmx/tdx.c-2887-\tkvm_tdx-\u003estate = TD_STATE_RUNNABLE;\narch/x86/kvm/vmx/tdx.c:2888:\t/* TD_STATE_RUNNABLE must be set before 'pre_fault_allowed' */\narch/x86/kvm/vmx/tdx.c-2889-\tsmp_wmb();\narch/x86/kvm/vmx/tdx.c:2890:\tkvm-\u003earch.pre_fault_allowed = true;\narch/x86/kvm/vmx/tdx.c-2891-\treturn 0;\n--\narch/x86/kvm/x86.c=9843=int kvm_arch_init_vm(struct kvm *kvm, unsigned long type)\n--\narch/x86/kvm/x86.c-9854-\t/* Decided by the vendor code for other VM types.  */\narch/x86/kvm/x86.c:9855:\tkvm-\u003earch.pre_fault_allowed =\narch/x86/kvm/x86.c-9856-\t\ttype == KVM_X86_DEFAULT_VM || type == KVM_X86_SW_PROTECTED_VM;\n"
}

15/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 36594
							output: 38
							thoughts: 865
						

Thoughts:
**WorthFuzzing Analysis: KVM Pre-Fault Memory Timing Change**

Okay, so I'm looking at this patch and my initial reaction is that we need to investigate if it's fuzz-worthy. Specifically, I'm focusing on the change in `arch/x86/kvm/vmx/tdx.c` where the comment flags a potential timing issue with `pre_fault_allowed`. I see that `pre_fault_allowed` is now read before `vcpu_load()`, which is a significant shift in the execution context on x86.

My next thought is to confirm KVM's reachability via syzkaller on amd64, since that's our target. Yep, no surprise, syzkaller heavily targets KVM on x86, using `/dev/kvm` and related syscalls like `ioctl$KVM_PRE_FAULT_MEMORY`. So, we're definitely in a fuzzable area.

Now, let's go through the `WorthFuzzing` criteria. I'm checking the positive criteria first. Does this patch touch reachable executable code, especially related to core subsystems? It's altering the flow of the `KVM_PRE_FAULT_MEMORY` ioctl to call `kvm_arch_pre_fault_allowed(vcpu)` earlier and introducing that function. I'm already seeing changes in memory management, so this is immediately raising a flag. The refactoring aspect is also important because it's restructuring a reachable data structure or altering algorithm flows in core KVM functionality. That's a strong positive indication.

Next, I need to check the negative criteria. Could this change be categorized as strictly refactoring or a minor change? No way. The modified code in `virt/kvm/kvm_main.c` and `arch/x86/kvm/mmu/mmu.c` doesn't fall into the negative criteria, because it's *not* unreachable. It's perfectly reachable on amd64. 

And, remember, never trust "No Functional Change" or "Refactoring" labels blindly. Refactorings can *absolutely* introduce semantic shifts, or expose bugs. This is a core subsystem change, and it's altering timing.

Given all of this, the negative criteria are not met, so the answer is clear. This *is* WorthFuzzing=true. I need to make sure to verify the configs needed for KVM on amd64 and check `arch/x86/kvm/Kconfig`.



16/2 2026/09/25 22:22 tool read-file
0m Args:
{
  "File": "arch/x86/kvm/Kconfig",
  "FirstLine": 15,
  "LineCount": 60
}

Results:
{
  "Contents": "  15:\t\n  16:\t\t  If you say N, all options in this submenu will be skipped and disabled.\n  17:\t\n  18:\tif VIRTUALIZATION\n  19:\t\n  20:\tconfig KVM_X86\n  21:\t\tdef_tristate KVM if (KVM_INTEL != n || KVM_AMD != n)\n  22:\t\tselect KVM_COMMON\n  23:\t\tselect KVM_ELIDE_TLB_FLUSH_IF_YOUNG\n  24:\t\tselect KVM_MMU_LOCKLESS_AGING\n  25:\t\tselect HAVE_KVM_IRQCHIP\n  26:\t\tselect HAVE_KVM_PFNCACHE\n  27:\t\tselect HAVE_KVM_DIRTY_RING_TSO\n  28:\t\tselect HAVE_KVM_DIRTY_RING_ACQ_REL\n  29:\t\tselect HAVE_KVM_IRQ_BYPASS\n  30:\t\tselect HAVE_KVM_IRQ_ROUTING\n  31:\t\tselect HAVE_KVM_READONLY_MEM\n  32:\t\tselect VHOST_TASK\n  33:\t\tselect KVM_ASYNC_PF\n  34:\t\tselect USER_RETURN_NOTIFIER\n  35:\t\tselect KVM_MMIO\n  36:\t\tselect SCHED_INFO\n  37:\t\tselect PERF_EVENTS\n  38:\t\tselect GUEST_PERF_EVENTS\n  39:\t\tselect PERF_GUEST_MEDIATED_PMU\n  40:\t\tselect HAVE_KVM_MSI\n  41:\t\tselect HAVE_KVM_CPU_RELAX_INTERCEPT\n  42:\t\tselect HAVE_KVM_NO_POLL\n  43:\t\tselect VIRT_XFER_TO_GUEST_WORK\n  44:\t\tselect KVM_GENERIC_DIRTYLOG_READ_PROTECT\n  45:\t\tselect KVM_VFIO\n  46:\t\tselect HAVE_KVM_PM_NOTIFIER if PM\n  47:\t\tselect KVM_GENERIC_HARDWARE_ENABLING\n  48:\t\tselect KVM_GENERIC_PRE_FAULT_MEMORY\n  49:\t\tselect KVM_WERROR if WERROR\n  50:\t\tselect KVM_GUEST_MEMFD if X86_64\n  51:\t\n  52:\tconfig KVM\n  53:\t\ttristate \"Kernel-based Virtual Machine (KVM) support\"\n  54:\t\tdepends on X86_LOCAL_APIC\n  55:\t\thelp\n  56:\t\t  Support hosting fully virtualized guest machines using hardware\n  57:\t\t  virtualization extensions.  You will need a fairly recent\n  58:\t\t  processor equipped with virtualization extensions. You will also\n  59:\t\t  need to select one or more of the processor modules below.\n  60:\t\n  61:\t\t  This module provides access to the hardware capabilities through\n  62:\t\t  a character device node named /dev/kvm.\n  63:\t\n  64:\t\t  To compile this as a module, choose M here: the module\n  65:\t\t  will be called kvm.\n  66:\t\n  67:\t\t  If unsure, say N.\n  68:\t\n  69:\tconfig KVM_WERROR\n  70:\t\tbool \"Compile KVM with -Werror\"\n  71:\t\t# Disallow KVM's -Werror if KASAN is enabled, e.g. to guard against\n  72:\t\t# randomized configs from selecting KVM_WERROR=y, which doesn't play\n  73:\t\t# nice with KASAN.  KASAN builds generates warnings for the default\n  74:\t\t# FRAME_WARN, i.e. KVM_WERROR=y with KASAN=y requires special tuning.\n"
}

17/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 38988
							output: 38
							thoughts: 27
						

Thoughts:
Okay, here's my take on those thoughts, tailored for an expert audience:

**Examining KVM_INTEL and KVM_AMD in Kconfig**

Alright, let's dive into the `Kconfig` file within the `arch/x86/kvm/` directory. My immediate interest is to pinpoint the configuration options specifically for Intel and AMD KVM support. I need to understand how these features are conditionally compiled, what dependencies they have, and how they relate to the broader KVM framework on x86. This helps me get the full picture of the build process and the architectural differences between KVM's Intel and AMD implementations. I'm looking for the logical structure: What dependencies does `KVM_INTEL` require? Is it linked to a more general `KVM` option, or does it stand alone? Are there any specific features within Intel's virtualization that are toggled via different configuration options? Similarly, I'll be looking into `KVM_AMD` to get a sense of AMD's implementation and how it compares. Is there any obvious overlap in features? Are there platform-specific considerations? Understanding the relationships between these options, and any associated dependencies, is critical for understanding KVM's build behavior and, ultimately, for debugging any KVM-related issues I might encounter.



18/2 2026/09/25 22:22 tool read-file
0m Args:
{
  "File": "arch/x86/kvm/Kconfig",
  "FirstLine": 75,
  "LineCount": 40
}

Results:
{
  "Contents": "  75:\t\t# Building KVM with -Werror and KASAN is still doable via enabling\n  76:\t\t# the kernel-wide WERROR=y.\n  77:\t\tdepends on KVM_X86 \u0026\u0026 ((EXPERT \u0026\u0026 !KASAN) || WERROR)\n  78:\t\thelp\n  79:\t\t  Add -Werror to the build flags for KVM.\n  80:\t\n  81:\t\t  If in doubt, say \"N\".\n  82:\t\n  83:\tconfig KVM_SW_PROTECTED_VM\n  84:\t\tbool \"Enable support for KVM software-protected VMs\"\n  85:\t\tdepends on EXPERT\n  86:\t\tdepends on KVM_X86 \u0026\u0026 X86_64\n  87:\t\tselect KVM_GENERIC_MEMORY_ATTRIBUTES\n  88:\t\thelp\n  89:\t\t  Enable support for KVM software-protected VMs.  Currently, software-\n  90:\t\t  protected VMs are purely a development and testing vehicle for\n  91:\t\t  KVM_CREATE_GUEST_MEMFD.  Attempting to run a \"real\" VM workload as a\n  92:\t\t  software-protected VM will fail miserably.\n  93:\t\n  94:\t\t  If unsure, say \"N\".\n  95:\t\n  96:\tconfig KVM_INTEL\n  97:\t\ttristate \"KVM for Intel (and compatible) processors support\"\n  98:\t\tdepends on KVM \u0026\u0026 IA32_FEAT_CTL\n  99:\t\tselect X86_FRED if X86_64\n 100:\t\thelp\n 101:\t\t  Provides support for KVM on processors equipped with Intel's VT\n 102:\t\t  extensions, a.k.a. Virtual Machine Extensions (VMX).\n 103:\t\n 104:\t\t  To compile this as a module, choose M here: the module\n 105:\t\t  will be called kvm-intel.\n 106:\t\n 107:\tconfig KVM_INTEL_PROVE_VE\n 108:\t        bool \"Check that guests do not receive #VE exceptions\"\n 109:\t        depends on KVM_INTEL \u0026\u0026 EXPERT\n 110:\t        help\n 111:\t          Checks that KVM's page table management code will not incorrectly\n 112:\t          let guests receive a virtualization exception.  Virtualization\n 113:\t          exceptions will be trapped by the hypervisor rather than injected\n 114:\t          in the guest.\n"
}

19/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 39974
							output: 20
							thoughts: 127
						

Thoughts:
Okay, let's break this down.

**KVM Configuration Deep Dive**

Alright, my initial thought is: we're dealing with KVM. Specifically, I need to pinpoint the kernel configurations that are essential. I'm immediately thinking of the core options: `KVM`, and the architecture-specific ones like `KVM_INTEL` and `KVM_AMD`. But hang on a second, before I get too far, let's not forget about software-protected VMs, I should also check into `KVM_SW_PROTECTED_VM`. Ah yes, and that will leverage `KVM_PRE_FAULT_MEMORY`, right? Now, to make sure I have the full picture, I need to confirm: what are the *baseline* configurations necessary for any KVM setup? Okay, looks like it's: `KVM`, `KVM_INTEL`, and `KVM_AMD` are the core requirements. Just to double-check my assumptions, it's always a good idea to cross-reference the `virt/kvm/Kconfig` file in the kernel source. I'll take a quick look to confirm if the overarching `VIRTUALIZATION` option is also a prerequisite or if it's handled differently within the KVM subsystem.



20/2 2026/09/25 22:22 tool grepper
0m Args:
{
  "Expression": "menuconfig VIRTUALIZATION"
}

Results:
{
  "Output": "arch/arm64/kvm/Kconfig=6=source \"virt/kvm/Kconfig\"\narch/arm64/kvm/Kconfig-7-\narch/arm64/kvm/Kconfig:8:menuconfig VIRTUALIZATION\narch/arm64/kvm/Kconfig-9-\tbool \"Virtualization\"\n--\narch/loongarch/kvm/Kconfig=6=source \"virt/kvm/Kconfig\"\narch/loongarch/kvm/Kconfig-7-\narch/loongarch/kvm/Kconfig:8:menuconfig VIRTUALIZATION\narch/loongarch/kvm/Kconfig-9-\tbool \"Virtualization\"\n--\narch/mips/kvm/Kconfig=5=source \"virt/kvm/Kconfig\"\narch/mips/kvm/Kconfig-6-\narch/mips/kvm/Kconfig:7:menuconfig VIRTUALIZATION\narch/mips/kvm/Kconfig-8-\tbool \"Virtualization\"\n--\narch/powerpc/kvm/Kconfig=6=source \"virt/kvm/Kconfig\"\narch/powerpc/kvm/Kconfig-7-\narch/powerpc/kvm/Kconfig:8:menuconfig VIRTUALIZATION\narch/powerpc/kvm/Kconfig-9-\tbool \"Virtualization\"\n--\narch/riscv/kvm/Kconfig=6=source \"virt/kvm/Kconfig\"\narch/riscv/kvm/Kconfig-7-\narch/riscv/kvm/Kconfig:8:menuconfig VIRTUALIZATION\narch/riscv/kvm/Kconfig-9-\tbool \"Virtualization\"\n--\narch/s390/kvm/Kconfig=5=source \"virt/kvm/Kconfig\"\narch/s390/kvm/Kconfig-6-\narch/s390/kvm/Kconfig:7:menuconfig VIRTUALIZATION\narch/s390/kvm/Kconfig-8-\tdef_bool y\n--\narch/x86/kvm/Kconfig=6=source \"virt/kvm/Kconfig\"\narch/x86/kvm/Kconfig-7-\narch/x86/kvm/Kconfig:8:menuconfig VIRTUALIZATION\narch/x86/kvm/Kconfig-9-\tbool \"Virtualization\"\n"
}

21/2 2026/09/25 22:22 llm patch-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 40982
							output: 241
							thoughts: 2362
						

Thoughts:
**KVM Patch Analysis: WorthFuzzing = TRUE**

Okay, so I've got a patch here that's primarily targeting arm64 for `KVM_PRE_FAULT_MEMORY` support. But, and this is crucial, it *also* touches code on amd64. Let's see... I've identified the key symbols that are modified: `kvm_vcpu_pre_fault_memory` in `virt/kvm/kvm_main.c`, `kvm_arch_pre_fault_allowed` in `arch/x86/kvm/mmu/mmu.c`, and `kvm_arch_vcpu_pre_fault_memory` (also in `arch/x86/kvm/mmu/mmu.c`).

The real question is, is this code *reachable* and is it changing logic in a way that needs fuzzing? I have to go through my mental checklist: "Non-kernel and non-executable changes"? Nope. These are core KVM files. "Structurally unreachable hardware"? No, we're dealing with KVM, so it's emulated and should be testable. "Unreachable execution paths"? Nope, the code is under active development.

Now, looking at the criteria for *WorthFuzzing=true*, this falls squarely into the "Refactorings and Code Cleanups" category. They are restructuring data structures and helper abstractions, and "Runtime Assertions and Defensive Checks" since the algorithm flow of `kvm_vcpu_pre_fault_memory()` was modified so that `kvm_arch_pre_fault_allowed(vcpu)` is called before `vcpu_load(vcpu)` and before `srcu_read_lock(&vcpu->kvm->srcu)`. It's a refactoring, but it's a refactoring that *changes* the order of operations in a KVM ioctl. The critical thing here is that the patch *alters* the control flow of the `KVM_PRE_FAULT_MEMORY` ioctl and also introduces a helper abstraction to implement a check to be run.

Therefore, because this touches executable and reachable code in the core KVM subsystem, I must say **WorthFuzzing = TRUE**.

So, what should I focus on? Specifically, the symbols: `kvm_vcpu_pre_fault_memory`, `kvm_arch_pre_fault_allowed`, and `kvm_arch_vcpu_pre_fault_memory`. These are the key functions. They are specific to the `KVM_PRE_FAULT_MEMORY` feature, and not general, so they're important!

The required config options will be: `KVM`, `KVM_INTEL` (or `KVM_AMD`), and `VIRTUALIZATION`. Because of where `KVM_PRE_FAULT_MEMORY` is used, the option `KVM_SW_PROTECTED_VM` is a good thing to select as well.

In summary, the patch series modifies core KVM logic on both amd64 and arm64, introducing a new architecture hook and altering the control flow of the `KVM_PRE_FAULT_MEMORY` ioctl. It's touching the code paths on amd64 when KVM and Intel/AMD are selected! Therefore, this patch is a high-priority candidate for fuzzing.



22/2 2026/09/25 22:23 tool set-results
0m Args:
{
  "EnableConfigs": [
    "KVM",
    "KVM_INTEL",
    "KVM_AMD",
    "VIRTUALIZATION"
  ],
  "FocusSymbols": [
    "kvm_vcpu_pre_fault_memory",
    "kvm_arch_pre_fault_allowed",
    "kvm_arch_vcpu_pre_fault_memory"
  ],
  "Reasoning": "While the patch primarily implements KVM_PRE_FAULT_MEMORY for arm64, it also refactors generic KVM (virt/kvm/kvm_main.c) and x86 KVM (arch/x86/kvm/mmu/mmu.c). It introduces the kvm_arch_pre_fault_allowed() hook, altering the control flow of the KVM_PRE_FAULT_MEMORY vcpu ioctl on amd64 to validate prefault permissions early prior to calling vcpu_load() and taking the SRCU lock. Because reachable executable code in the KVM subsystem is modified on amd64, it warrants fuzzing.",
  "WorthFuzzing": true
}

Results:
{
  "EnableConfigs": [
    "KVM",
    "KVM_INTEL",
    "KVM_AMD",
    "VIRTUALIZATION"
  ],
  "FocusSymbols": [
    "kvm_vcpu_pre_fault_memory",
    "kvm_arch_pre_fault_allowed",
    "kvm_arch_vcpu_pre_fault_memory"
  ],
  "Reasoning": "While the patch primarily implements KVM_PRE_FAULT_MEMORY for arm64, it also refactors generic KVM (virt/kvm/kvm_main.c) and x86 KVM (arch/x86/kvm/mmu/mmu.c). It introduces the kvm_arch_pre_fault_allowed() hook, altering the control flow of the KVM_PRE_FAULT_MEMORY vcpu ioctl on amd64 to validate prefault permissions early prior to calling vcpu_load() and taking the SRCU lock. Because reachable executable code in the KVM subsystem is modified on amd64, it warrants fuzzing.",
  "WorthFuzzing": true
}

23/1 2026/09/25 22:23 action If
0m Args:
{
  "WorthFuzzing": true
}

Results:
null

24/2 2026/09/25 22:23 agent kmsan-evaluator
0m Model:
core

Args:
null

Results:
{
  "KMSANReasoning": "The patch series adds support for KVM_PRE_FAULT_MEMORY on arm64 and refactors the generic ioctl path (virt/kvm/kvm_main.c) to invoke a new architecture hook kvm_arch_pre_fault_allowed() before vcpu_load(). On the target architecture (amd64), the change simply moves the existing pre-fault allowance check (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed) from kvm_arch_vcpu_pre_fault_memory() to kvm_arch_pre_fault_allowed().\n\nNo uninitialized memory hazards are introduced:\n1. No uninitialized memory is returned to userspace: KVM_PRE_FAULT_MEMORY reads struct kvm_pre_fault_memory fully from userspace via copy_from_user() before mutating its fields and copying it back.\n2. In the arm64 implementation, all local structs (such as kvm_s2_fault_desc and kvm_s2_fault_result) are fully initialized before use.\n3. No conditional branches rely on uninitialized values, and no padding leaks or uninitialized page/heap reads are exposed.\n\nStandard bug detectors and KASAN are fully sufficient to test these changes. A dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Instruction:
You are an expert Linux kernel security engineer specializing in kernel memory
error detectors (KASAN and KMSAN). Your job is to review the provided patch series and
determine if the code changes justify spawning a dedicated KMSAN (KernelMemorySanitizer)
fuzzing session in addition to standard KASAN fuzzing.

CRITICAL DISTINCTION BETWEEN KASAN AND KMSAN:
- Standard KASAN kernel builds (upstream-apparmor-kasan.config) already enable
  a comprehensive suite of debugging tools and sanitizers, including KASAN
  (out-of-bounds accesses, use-after-free, double free, invalid free), LOCKDEP
  (locking bugs and deadlocks), UB-sanitizers, and memory corruption checks.
- KMSAN (KernelMemorySanitizer) detects reads of UNINITIALIZED memory (stack, heap,
  or page allocations) and kernel-to-user memory info-leaks.

Rule: THERE IS NO SENSE IN RUNNING A KMSAN SESSION IF A BUG CAN BE CAUGHT BY KASAN,
LOCKDEP, OR OTHER STANDARD BUG DETECTORS.
A dedicated KMSAN fuzzing session incurs significant resource costs. You must ONLY
set NeedsKMSAN=true if the code changes introduce or expose UNINITIALIZED MEMORY risks
that are detected ONLY by KMSAN.

Look holistically at the patch series and surrounding code. Even if no direct
uninitialized field accesses or new buffer allocations are added in the diff itself,
a patch may alter control flow, bounds checking, or data length calculations in ways
that change how the rest of the code operates on existing buffers (e.g. allowing
uninitialized stack/heap memory to be read, copied to user space, or used in control
flow). Do not hesitate to use your code access tools to inspect the surrounding code,
called functions, and callers.

Set NeedsKMSAN=true ONLY IF the patch introduces or modifies:
1. Kernel structures sent to user space (via copy_to_user, put_user, netlink skb
   attributes, ioctl output arguments, socket options, or BPF buffers) where fields
   or structure padding might not be fully initialized/zeroed.
2. Conditional logic or branching that depends on potentially uninitialized variables
   or struct fields.
3. Allocation or initialization of complex data structures where uninitialized fields
   could be read later in reachable code paths.
4. Bounds checks, lengths, or logic in a way that allows surrounding code to access
   uninitialized bytes of existing buffers.

Set NeedsKMSAN=false IF:
- The code changes primarily risk out-of-bounds access, array overflows, NULL pointer
  dereferences, locking deadlocks, or use-after-free bugs (these are already caught
  by KASAN, LOCKDEP, or standard bug detectors).
- All stack/heap structures touched or introduced by the patch are fully zeroed
  or initialized (e.g. using = {0}, memset, kzalloc) before being read or copied.
- The patch does not introduce any risk of uninitialized memory usage or info-leaks.

Use your code access tools to inspect the surrounding code if necessary, then provide
detailed KMSANReasoning contrasting KASAN vs KMSAN applicability for this patch.
Prefer calling several tools at the same time to save round-trips.


Use set-results tool to provide results of the analysis.
It must be called exactly once before the final reply.
Ignore results of this tool.

Prompt:
Target architecture: amd64

For your convenience, here is the diff of the changes:
commit d79cc5272139a53b1da70940b59fa09bf4f16b16
Author: syz-cluster <triage@syzkaller.com>
Date:   Fri Sep 25 22:21:57 2026 +0000

    syz-cluster: applied patch under review

diff --git a/Documentation/virt/kvm/api.rst b/Documentation/virt/kvm/api.rst
index 1a90598901c57..71b642148b7f6 100644
--- a/Documentation/virt/kvm/api.rst
+++ b/Documentation/virt/kvm/api.rst
@@ -6489,7 +6489,7 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.
 ---------------------------
 
 :Capability: KVM_CAP_PRE_FAULT_MEMORY
-:Architectures: none
+:Architectures: x86, s390, arm64
 :Type: vcpu ioctl
 :Parameters: struct kvm_pre_fault_memory (in/out)
 :Returns: 0 if at least one page is processed, < 0 on error
@@ -6497,12 +6497,15 @@ See KVM_SET_USER_MEMORY_REGION2 for additional details.
 Errors:
 
   ========== ===============================================================
+  EAGAIN     A race occurred before progress was made, but a retry may succeed.
   EINVAL     The specified `gpa` and `size` were invalid (e.g. not
              page aligned, causes an overflow, or size is zero), or the VM
              is UCONTROL (s390).
   ENOENT     The specified `gpa` is outside defined memslots.
+  ENOEXEC    The vCPU has not been initialised (arm64).
   EINTR      An unmasked signal is pending and no page was processed.
   EFAULT     The parameter address was invalid.
+  EHWPOISON  A poisoned host page was encountered.
   EOPNOTSUPP Mapping memory for a GPA is unsupported by the
              hypervisor, and/or for the current vCPU state/mode.
   EIO        unexpected error conditions (also causes a WARN)
@@ -6522,7 +6525,17 @@ Errors:
 KVM_PRE_FAULT_MEMORY populates KVM's stage-2 page tables used to map memory
 for the current vCPU state.  KVM maps memory as if the vCPU generated a
 stage-2 read page fault, e.g. faults in memory as needed, but doesn't break
-CoW.  On x86, KVM does not mark any newly created stage-2 PTE as Accessed.
+CoW.  On arm64, KVM marks newly created stage-2 PTEs as Accessed, as it
+does for any stage-2 fault, but leaves the Accessed state of existing PTEs
+unchanged.  On x86, KVM does not mark any newly created stage-2 PTE as
+Accessed, and for s390 it is not applicable.
+
+On arm64, a GPA is interpreted as an IPA, and never interpreted as the IPA
+of a nested guest. Pre-faulting only populates canonical stage-2 page
+tables.
+
+The feature is not supported on arm64 if the protected KVM (pKVM) feature
+is enabled.
 
 In the case of confidential VM types where there is an initial set up of
 private guest memory before the guest is 'finalized'/measured, this ioctl
@@ -6537,7 +6550,7 @@ When the ioctl returns, the input values are updated to point to the
 remaining range.  If `size` > 0 on return, the caller can just issue
 the ioctl again with the same `struct kvm_map_memory` argument.
 
-Shadow page tables cannot support this ioctl because they
+On x86, shadow page tables cannot support this ioctl because they
 are indexed by virtual address or nested guest physical address.
 Calling this ioctl when the guest is using shadow page tables (for
 example because it is running a nested guest with nested page tables)
diff --git a/arch/arm64/include/asm/esr.h b/arch/arm64/include/asm/esr.h
index f816f5d77f1a5..86756fc9bb5a2 100644
--- a/arch/arm64/include/asm/esr.h
+++ b/arch/arm64/include/asm/esr.h
@@ -437,6 +437,35 @@
 #ifndef __ASSEMBLER__
 #include <asm/types.h>
 
+static __always_inline bool esr_trap_is_iabt(unsigned long esr)
+{
+	return ESR_ELx_EC(esr) == ESR_ELx_EC_IABT_LOW;
+}
+
+/* Always check for S1PTW *before* using this. */
+static __always_inline bool esr_dabt_is_write(unsigned long esr)
+{
+	return esr & ESR_ELx_WNR;
+}
+
+static __always_inline bool esr_dabt_is_cm(unsigned long esr)
+{
+	return esr & ESR_ELx_CM;
+}
+
+static __always_inline bool esr_abt_is_sea(unsigned long esr)
+{
+	switch (esr & ESR_ELx_FSC) {
+	case ESR_ELx_FSC_EXTABT:
+	case ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):
+	case ESR_ELx_FSC_SECC:
+	case ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):
+		return true;
+	default:
+		return false;
+	}
+}
+
 static inline unsigned long esr_brk_comment(unsigned long esr)
 {
 	return esr & ESR_ELx_BRK64_ISS_COMMENT_MASK;
diff --git a/arch/arm64/include/asm/kvm_emulate.h b/arch/arm64/include/asm/kvm_emulate.h
index a3c1928bdf743..708a4b6d88230 100644
--- a/arch/arm64/include/asm/kvm_emulate.h
+++ b/arch/arm64/include/asm/kvm_emulate.h
@@ -336,6 +336,16 @@ static __always_inline u64 kvm_vcpu_get_esr(const struct kvm_vcpu *vcpu)
 	return vcpu->arch.fault.esr_el2;
 }
 
+static __always_inline bool esr_abt_is_s1ptw(unsigned long esr)
+{
+	return esr & ESR_ELx_S1PTW;
+}
+
+static __always_inline bool esr_abt_is_exec_fault(unsigned long esr)
+{
+	return esr_trap_is_iabt(esr) && !esr_abt_is_s1ptw(esr);
+}
+
 static inline bool guest_hyp_wfx_traps_enabled(const struct kvm_vcpu *vcpu)
 {
 	u64 esr = kvm_vcpu_get_esr(vcpu);
@@ -411,18 +421,13 @@ static __always_inline int kvm_vcpu_dabt_get_rd(const struct kvm_vcpu *vcpu)
 
 static __always_inline bool kvm_vcpu_abt_iss1tw(const struct kvm_vcpu *vcpu)
 {
-	return !!(kvm_vcpu_get_esr(vcpu) & ESR_ELx_S1PTW);
+	return esr_abt_is_s1ptw(kvm_vcpu_get_esr(vcpu));
 }
 
 /* Always check for S1PTW *before* using this. */
 static __always_inline bool kvm_vcpu_dabt_iswrite(const struct kvm_vcpu *vcpu)
 {
-	return kvm_vcpu_get_esr(vcpu) & ESR_ELx_WNR;
-}
-
-static inline bool kvm_vcpu_dabt_is_cm(const struct kvm_vcpu *vcpu)
-{
-	return !!(kvm_vcpu_get_esr(vcpu) & ESR_ELx_CM);
+	return esr_dabt_is_write(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline unsigned int kvm_vcpu_dabt_get_as(const struct kvm_vcpu *vcpu)
@@ -443,12 +448,7 @@ static __always_inline u8 kvm_vcpu_trap_get_class(const struct kvm_vcpu *vcpu)
 
 static inline bool kvm_vcpu_trap_is_iabt(const struct kvm_vcpu *vcpu)
 {
-	return kvm_vcpu_trap_get_class(vcpu) == ESR_ELx_EC_IABT_LOW;
-}
-
-static inline bool kvm_vcpu_trap_is_exec_fault(const struct kvm_vcpu *vcpu)
-{
-	return kvm_vcpu_trap_is_iabt(vcpu) && !kvm_vcpu_abt_iss1tw(vcpu);
+	return esr_trap_is_iabt(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline u8 kvm_vcpu_trap_get_fault(const struct kvm_vcpu *vcpu)
@@ -468,26 +468,9 @@ bool kvm_vcpu_trap_is_translation_fault(const struct kvm_vcpu *vcpu)
 	return esr_fsc_is_translation_fault(kvm_vcpu_get_esr(vcpu));
 }
 
-static inline
-u64 kvm_vcpu_trap_get_perm_fault_granule(const struct kvm_vcpu *vcpu)
-{
-	unsigned long esr = kvm_vcpu_get_esr(vcpu);
-
-	BUG_ON(!esr_fsc_is_permission_fault(esr));
-	return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(esr & ESR_ELx_FSC_LEVEL));
-}
-
 static __always_inline bool kvm_vcpu_abt_issea(const struct kvm_vcpu *vcpu)
 {
-	switch (kvm_vcpu_trap_get_fault(vcpu)) {
-	case ESR_ELx_FSC_EXTABT:
-	case ESR_ELx_FSC_SEA_TTW(-1) ... ESR_ELx_FSC_SEA_TTW(3):
-	case ESR_ELx_FSC_SECC:
-	case ESR_ELx_FSC_SECC_TTW(-1) ... ESR_ELx_FSC_SECC_TTW(3):
-		return true;
-	default:
-		return false;
-	}
+	return esr_abt_is_sea(kvm_vcpu_get_esr(vcpu));
 }
 
 static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)
@@ -496,9 +479,9 @@ static __always_inline int kvm_vcpu_sys_get_rt(struct kvm_vcpu *vcpu)
 	return ESR_ELx_SYS64_ISS_RT(esr);
 }
 
-static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
+static inline bool esr_abt_is_write_fault(unsigned long esr)
 {
-	if (kvm_vcpu_abt_iss1tw(vcpu)) {
+	if (esr_abt_is_s1ptw(esr)) {
 		/*
 		 * Only a permission fault on a S1PTW should be
 		 * considered as a write. Otherwise, page tables baked
@@ -511,13 +494,18 @@ static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
 		 * first), then a permission fault to allow the flags
 		 * to be set.
 		 */
-		return kvm_vcpu_trap_is_permission_fault(vcpu);
+		return esr_fsc_is_permission_fault(esr);
 	}
 
-	if (kvm_vcpu_trap_is_iabt(vcpu))
+	if (esr_trap_is_iabt(esr))
 		return false;
 
-	return kvm_vcpu_dabt_iswrite(vcpu);
+	return esr_dabt_is_write(esr);
+}
+
+static inline bool kvm_is_write_fault(struct kvm_vcpu *vcpu)
+{
+	return esr_abt_is_write_fault(kvm_vcpu_get_esr(vcpu));
 }
 
 static inline unsigned long kvm_vcpu_get_mpidr_aff(struct kvm_vcpu *vcpu)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 20d12da3d28ee..589db1d6405fa 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -859,6 +859,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,
  * @addr:	Input address for the start of the walk.
  * @ptep:	Pointer to storage for the retrieved PTE.
  * @level:	Pointer to storage for the level of the retrieved PTE.
+ * @flags:	Flags to control the page-table walk
+ *		(see struct kvm_pgtable_visit_ctx).
  *
  * The offset of @addr within a page is ignored.
  *
@@ -869,7 +871,8 @@ int kvm_pgtable_walk(struct kvm_pgtable *pgt, u64 addr, u64 size,
  * Return: 0 on success, negative error code on failure.
  */
 int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
-			 kvm_pte_t *ptep, s8 *level);
+			 kvm_pte_t *ptep, s8 *level,
+			 enum kvm_pgtable_walk_flags flags);
 
 /**
  * kvm_pgtable_stage2_pte_prot() - Retrieve the protection attributes of a
diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h
index cad60569f0619..c1831c8e421d4 100644
--- a/arch/arm64/include/asm/kvm_pkvm.h
+++ b/arch/arm64/include/asm/kvm_pkvm.h
@@ -44,9 +44,9 @@ static inline bool kvm_pkvm_ext_allowed(struct kvm *kvm, long ext)
 	case KVM_CAP_ARM_PTRAUTH_GENERIC:
 		return true;
 	case KVM_CAP_ARM_MTE:
-		return false;
 	case KVM_CAP_ARM_EAGER_SPLIT_CHUNK_SIZE:
 	case KVM_CAP_ARM_SUPPORTED_BLOCK_SIZES:
+	case KVM_CAP_PRE_FAULT_MEMORY:
 		return false;
 	default:
 		return !kvm || !kvm_vm_is_protected(kvm);
diff --git a/arch/arm64/kvm/Kconfig b/arch/arm64/kvm/Kconfig
index 449154f9a4852..71233068b7cb6 100644
--- a/arch/arm64/kvm/Kconfig
+++ b/arch/arm64/kvm/Kconfig
@@ -37,6 +37,7 @@ menuconfig KVM
 	select SCHED_INFO
 	select GUEST_PERF_EVENTS if PERF_EVENTS
 	select KVM_GUEST_MEMFD
+	select KVM_GENERIC_PRE_FAULT_MEMORY
 	help
 	  Support hosting virtualized guest machines.
 
diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c
index 31f803010c57c..367638e1e0204 100644
--- a/arch/arm64/kvm/arm.c
+++ b/arch/arm64/kvm/arm.c
@@ -410,6 +410,7 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext)
 	case KVM_CAP_COUNTER_OFFSET:
 	case KVM_CAP_ARM_WRITABLE_IMP_ID_REGS:
 	case KVM_CAP_ARM_SEA_TO_USER:
+	case KVM_CAP_PRE_FAULT_MEMORY:
 		r = 1;
 		break;
 	case KVM_CAP_SET_GUEST_DEBUG2:
diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
index a6a47c1e058b3..443ee433119ef 100644
--- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c
+++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c
@@ -540,7 +540,7 @@ static int host_stage2_adjust_range(u64 addr, struct kvm_mem_range *range)
 	int ret;
 
 	hyp_assert_lock_held(&host_mmu.lock);
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, addr, &pte, &level, 0);
 	if (ret)
 		return ret;
 
@@ -930,7 +930,7 @@ static int get_valid_guest_pte(struct pkvm_hyp_vm *vm, u64 ipa, kvm_pte_t *ptep,
 	s8 level;
 	int ret;
 
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		return ret;
 	if (guest_pte_is_poisoned(pte))
@@ -979,7 +979,7 @@ int __pkvm_vcpu_in_poison_fault(struct pkvm_hyp_vcpu *hyp_vcpu)
 	ipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(&hyp_vcpu->vcpu));
 
 	guest_lock_component(vm);
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		goto unlock;
 
@@ -1335,7 +1335,7 @@ static int host_stage2_get_guest_info(phys_addr_t phys, struct pkvm_hyp_vm **vm,
 		return -EPERM;
 	}
 
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, &level, 0);
 	if (ret)
 		return ret;
 
@@ -1564,7 +1564,7 @@ static int __check_host_shared_guest(struct pkvm_hyp_vm *vm, u64 *__phys, u64 ip
 	s8 level;
 	int ret;
 
-	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&vm->pgt, ipa, &pte, &level, 0);
 	if (ret)
 		return ret;
 	if (!kvm_pte_valid(pte))
diff --git a/arch/arm64/kvm/hyp/nvhe/mm.c b/arch/arm64/kvm/hyp/nvhe/mm.c
index 29ab5ee9d57fc..bcf2fd1ddd75d 100644
--- a/arch/arm64/kvm/hyp/nvhe/mm.c
+++ b/arch/arm64/kvm/hyp/nvhe/mm.c
@@ -486,7 +486,7 @@ static int check_page_ownership(phys_addr_t phys)
 			return -EPERM;
 	}
 
-	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, NULL);
+	ret = kvm_pgtable_get_leaf(&host_mmu.pgt, phys, &pte, NULL, 0);
 	if (ret)
 		return ret;
 
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index b74dd5ce1efd3..347eec3957d6a 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -298,12 +298,13 @@ static int leaf_walker(const struct kvm_pgtable_visit_ctx *ctx,
 }
 
 int kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,
-			 kvm_pte_t *ptep, s8 *level)
+			 kvm_pte_t *ptep, s8 *level,
+			 enum kvm_pgtable_walk_flags flags)
 {
 	struct leaf_walk_data data;
 	struct kvm_pgtable_walker walker = {
 		.cb	= leaf_walker,
-		.flags	= KVM_PGTABLE_WALK_LEAF,
+		.flags	= flags | KVM_PGTABLE_WALK_LEAF,
 		.arg	= &data,
 	};
 	int ret;
diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c
index d6187295c3736..2a281488792fb 100644
--- a/arch/arm64/kvm/mmu.c
+++ b/arch/arm64/kvm/mmu.c
@@ -5,6 +5,7 @@
  */
 
 #include <linux/acpi.h>
+#include <linux/cleanup.h>
 #include <linux/mman.h>
 #include <linux/kvm_host.h>
 #include <linux/interval_tree.h>
@@ -895,7 +896,7 @@ static int get_user_mapping_size(struct kvm *kvm, u64 addr)
 	 * IPI-ing threads).
 	 */
 	local_irq_save(flags);
-	ret = kvm_pgtable_get_leaf(&pgt, addr, &pte, &level);
+	ret = kvm_pgtable_get_leaf(&pgt, addr, &pte, &level, 0);
 	local_irq_restore(flags);
 
 	if (ret)
@@ -1606,9 +1607,9 @@ static void *get_mmu_memcache(struct kvm_vcpu *vcpu)
 		return &vcpu->arch.pkvm_memcache;
 }
 
-static int topup_mmu_memcache(struct kvm_vcpu *vcpu, void *memcache)
+static int topup_mmu_memcache(struct kvm_s2_mmu *mmu, void *memcache)
 {
-	int min_pages = kvm_mmu_cache_min_pages(vcpu->arch.hw_mmu);
+	int min_pages = kvm_mmu_cache_min_pages(mmu);
 
 	if (!is_protected_kvm_enabled())
 		return kvm_mmu_topup_memory_cache(memcache, min_pages);
@@ -1655,15 +1656,47 @@ struct kvm_s2_fault_desc {
 	struct kvm_s2_trans	*nested;
 	struct kvm_memory_slot	*memslot;
 	unsigned long		hva;
+	unsigned long		esr;
+	struct kvm_s2_mmu	*mmu;
 };
 
-static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
+struct kvm_s2_fault_result {
+	unsigned long mapping_size;
+};
+
+static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_fsc_is_permission_fault(s2fd->esr);
+}
+
+static bool kvm_s2_fault_is_exec(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_abt_is_exec_fault(s2fd->esr);
+}
+
+static bool kvm_s2_fault_is_write(const struct kvm_s2_fault_desc *s2fd)
+{
+	return esr_abt_is_write_fault(s2fd->esr);
+}
+
+static u64 kvm_s2_perm_fault_granule(const struct kvm_s2_fault_desc *s2fd)
+{
+	u64 level;
+
+	if (!kvm_s2_fault_is_perm(s2fd))
+		return 0;
+	level = s2fd->esr & ESR_ELx_FSC_LEVEL;
+	return BIT(ARM64_HW_PGTABLE_LEVEL_SHIFT(level));
+}
+
+static int gmem_abort(const struct kvm_s2_fault_desc *s2fd,
+		      struct kvm_s2_fault_result *result)
 {
 	bool write_fault, exec_fault;
-	bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
+	bool perm_fault = kvm_s2_fault_is_perm(s2fd);
 	enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;
 	enum kvm_pgtable_prot prot = KVM_PGTABLE_PROT_R;
-	struct kvm_pgtable *pgt = s2fd->vcpu->arch.hw_mmu->pgt;
+	struct kvm_pgtable *pgt = s2fd->mmu->pgt;
 	struct kvm_guest_s2_mapping *mapping = NULL;
 	unsigned long mmu_seq;
 	struct page *page;
@@ -1675,7 +1708,7 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 
 	if (!perm_fault) {
 		memcache = get_mmu_memcache(s2fd->vcpu);
-		ret = topup_mmu_memcache(s2fd->vcpu, memcache);
+		ret = topup_mmu_memcache(s2fd->mmu, memcache);
 		if (ret)
 			return ret;
 		if (kvm_is_nested_s2_mmu(kvm, pgt->mmu)) {
@@ -1690,8 +1723,8 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 	else
 		gfn = s2fd->fault_ipa >> PAGE_SHIFT;
 
-	write_fault = kvm_is_write_fault(s2fd->vcpu);
-	exec_fault = kvm_vcpu_trap_is_exec_fault(s2fd->vcpu);
+	write_fault = kvm_s2_fault_is_write(s2fd);
+	exec_fault = kvm_s2_fault_is_exec(s2fd);
 
 	VM_WARN_ON_ONCE(write_fault && exec_fault);
 
@@ -1701,8 +1734,10 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 
 	ret = kvm_gmem_get_pfn(kvm, s2fd->memslot, gfn, &pfn, &page, NULL);
 	if (ret) {
-		kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE,
-					      write_fault, exec_fault, false);
+		/* If result is non-NULL this is a synthetic fault. */
+		if (!result)
+			kvm_prepare_memory_fault_exit(s2fd->vcpu, s2fd->fault_ipa, PAGE_SIZE,
+						      write_fault, exec_fault, false);
 		kfree(mapping);
 		return ret;
 	}
@@ -1757,7 +1792,13 @@ static int gmem_abort(const struct kvm_s2_fault_desc *s2fd)
 	if ((prot & KVM_PGTABLE_PROT_W) && !ret)
 		mark_page_dirty_in_slot(kvm, s2fd->memslot, gfn);
 
-	return ret != -EAGAIN ? ret : 0;
+	if (ret == -EAGAIN)
+		return result ? ret : 0;
+
+	if (result && !ret)
+		result->mapping_size = PAGE_SIZE;
+
+	return ret;
 }
 
 struct kvm_s2_fault_vma_info {
@@ -1779,7 +1820,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 {
 	unsigned int flags = FOLL_HWPOISON | FOLL_LONGTERM | FOLL_WRITE;
 	struct kvm_vcpu *vcpu = s2fd->vcpu;
-	struct kvm_pgtable *pgt = vcpu->arch.hw_mmu->pgt;
+	struct kvm_pgtable *pgt = s2fd->mmu->pgt;
 	struct mm_struct *mm = current->mm;
 	struct kvm *kvm = vcpu->kvm;
 	void *hyp_memcache;
@@ -1787,7 +1828,7 @@ static int pkvm_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	int ret;
 
 	hyp_memcache = get_mmu_memcache(vcpu);
-	ret = topup_mmu_memcache(vcpu, hyp_memcache);
+	ret = topup_mmu_memcache(s2fd->mmu, hyp_memcache);
 	if (ret)
 		return -ENOMEM;
 
@@ -1910,11 +1951,6 @@ static short kvm_s2_resolve_vma_size(const struct kvm_s2_fault_desc *s2fd,
 	return vma_shift;
 }
 
-static bool kvm_s2_fault_is_perm(const struct kvm_s2_fault_desc *s2fd)
-{
-	return kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
-}
-
 static int kvm_s2_fault_get_vma_info(const struct kvm_s2_fault_desc *s2fd,
 				     struct kvm_s2_fault_vma_info *s2vi)
 {
@@ -1980,13 +2016,11 @@ static int kvm_s2_fault_pin_pfn(const struct kvm_s2_fault_desc *s2fd,
 		return ret;
 
 	s2vi->pfn = __kvm_faultin_pfn(s2fd->memslot, get_canonical_gfn(s2fd, s2vi),
-				      kvm_is_write_fault(s2fd->vcpu) ? FOLL_WRITE : 0,
+				      kvm_s2_fault_is_write(s2fd) ? FOLL_WRITE : 0,
 				      &s2vi->map_writable, &s2vi->page);
 	if (unlikely(is_error_noslot_pfn(s2vi->pfn))) {
-		if (s2vi->pfn == KVM_PFN_ERR_HWPOISON) {
-			kvm_send_hwpoison_signal(s2fd->hva, __ffs(s2vi->vma_pagesize));
-			return 0;
-		}
+		if (s2vi->pfn == KVM_PFN_ERR_HWPOISON)
+			return -EHWPOISON;
 		return -EFAULT;
 	}
 
@@ -2038,7 +2072,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 {
 	struct kvm *kvm = s2fd->vcpu->kvm;
 
-	if (kvm_vcpu_trap_is_exec_fault(s2fd->vcpu) && s2vi->map_non_cacheable)
+	if (kvm_s2_fault_is_exec(s2fd) && s2vi->map_non_cacheable)
 		return -ENOEXEC;
 
 	/*
@@ -2047,7 +2081,7 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 	 * and trigger the exception here. Since the memslot is valid, inject
 	 * the fault back to the guest.
 	 */
-	if (esr_fsc_is_excl_atomic_fault(kvm_vcpu_get_esr(s2fd->vcpu))) {
+	if (esr_fsc_is_excl_atomic_fault(s2fd->esr)) {
 		kvm_inject_dabt_excl_atomic(s2fd->vcpu, kvm_vcpu_get_hfar(s2fd->vcpu));
 		return 1;
 	}
@@ -2056,13 +2090,13 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 
 	if (s2vi->map_writable && (s2vi->device ||
 				   !memslot_is_logging(s2fd->memslot) ||
-				   kvm_is_write_fault(s2fd->vcpu)))
+				   kvm_s2_fault_is_write(s2fd)))
 		*prot |= KVM_PGTABLE_PROT_W;
 
 	if (s2fd->nested)
 		*prot = adjust_nested_fault_perms(s2fd->nested, *prot);
 
-	if (kvm_vcpu_trap_is_exec_fault(s2fd->vcpu))
+	if (kvm_s2_fault_is_exec(s2fd))
 		*prot |= KVM_PGTABLE_PROT_X;
 
 	if (s2vi->map_non_cacheable)
@@ -2086,7 +2120,8 @@ static int kvm_s2_fault_compute_prot(const struct kvm_s2_fault_desc *s2fd,
 static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 			    const struct kvm_s2_fault_vma_info *s2vi,
 			    enum kvm_pgtable_prot prot,
-			    void *memcache)
+			    void *memcache,
+			    struct kvm_s2_fault_result *result)
 {
 	enum kvm_pgtable_walk_flags flags = KVM_PGTABLE_WALK_SHARED;
 	struct kvm_guest_s2_mapping *mapping = NULL;
@@ -2100,7 +2135,7 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 	gfn_t gfn;
 	int ret;
 
-	if (kvm_is_nested_s2_mmu(kvm, s2fd->vcpu->arch.hw_mmu)) {
+	if (kvm_is_nested_s2_mmu(kvm, s2fd->mmu)) {
 		mapping = kmalloc_obj(struct kvm_guest_s2_mapping,
 				      GFP_KERNEL_ACCOUNT);
 		if (!mapping) {
@@ -2110,13 +2145,12 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 	}
 
 	kvm_fault_lock(kvm);
-	pgt = s2fd->vcpu->arch.hw_mmu->pgt;
+	pgt = s2fd->mmu->pgt;
 	ret = -EAGAIN;
 	if (mmu_invalidate_retry(kvm, s2vi->mmu_seq))
 		goto out_unlock;
 
-	perm_fault_granule = (kvm_s2_fault_is_perm(s2fd) ?
-			      kvm_vcpu_trap_get_perm_fault_granule(s2fd->vcpu) : 0);
+	perm_fault_granule = kvm_s2_perm_fault_granule(s2fd);
 	mapping_size = s2vi->vma_pagesize;
 	pfn = s2vi->pfn;
 	gfn = s2vi->gfn;
@@ -2188,14 +2222,19 @@ static int kvm_s2_fault_map(const struct kvm_s2_fault_desc *s2fd,
 		mark_page_dirty_in_slot(kvm, s2fd->memslot,
 					gpa_to_gfn(canonical_ipa));
 
-	if (ret != -EAGAIN)
-		return ret;
-	return 0;
+	if (ret == -EAGAIN)
+		return result ? ret : 0;
+
+	if (result && !ret)
+		result->mapping_size = mapping_size;
+
+	return ret;
 }
 
-static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
+static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd,
+			  struct kvm_s2_fault_result *result)
 {
-	bool perm_fault = kvm_vcpu_trap_is_permission_fault(s2fd->vcpu);
+	bool perm_fault = kvm_s2_fault_is_perm(s2fd);
 	struct kvm_s2_fault_vma_info s2vi = {};
 	enum kvm_pgtable_prot prot;
 	void *memcache;
@@ -2213,7 +2252,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	memcache = get_mmu_memcache(s2fd->vcpu);
 	if (!perm_fault || memslot_is_logging(s2fd->memslot) ||
 	    is_protected_kvm_enabled()) {
-		ret = topup_mmu_memcache(s2fd->vcpu, memcache);
+		ret = topup_mmu_memcache(s2fd->mmu, memcache);
 		if (ret)
 			return ret;
 	}
@@ -2223,6 +2262,13 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 	 * get block mapping for device MMIO region.
 	 */
 	ret = kvm_s2_fault_pin_pfn(s2fd, &s2vi);
+	if (ret == -EHWPOISON) {
+		/* If result is specified, let the caller handle this. */
+		if (result)
+			return -EHWPOISON;
+		kvm_send_hwpoison_signal(s2fd->hva, __ffs(s2vi.vma_pagesize));
+		return 0;
+	}
 	if (ret != 1)
 		return ret;
 
@@ -2232,7 +2278,7 @@ static int user_mem_abort(const struct kvm_s2_fault_desc *s2fd)
 		return ret;
 	}
 
-	return kvm_s2_fault_map(s2fd, &s2vi, prot, memcache);
+	return kvm_s2_fault_map(s2fd, &s2vi, prot, memcache, result);
 }
 
 /* Resolve the access fault by making the page young again. */
@@ -2342,7 +2388,8 @@ int kvm_handle_guest_sea(struct kvm_vcpu *vcpu)
 int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 {
 	struct kvm_s2_trans nested_trans, *nested = NULL;
-	unsigned long esr;
+	unsigned long esr = kvm_vcpu_get_esr(vcpu);
+	struct kvm_s2_mmu *mmu = vcpu->arch.hw_mmu;
 	phys_addr_t fault_ipa; /* The address we faulted on */
 	phys_addr_t ipa; /* Always the IPA in the L1 guest phys space */
 	struct kvm_memory_slot *memslot;
@@ -2351,11 +2398,9 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	gfn_t gfn;
 	int ret, idx;
 
-	if (kvm_vcpu_abt_issea(vcpu))
+	if (esr_abt_is_sea(esr))
 		return kvm_handle_guest_sea(vcpu);
 
-	esr = kvm_vcpu_get_esr(vcpu);
-
 	/*
 	 * The fault IPA should be reliable at this point as we're not dealing
 	 * with an SEA.
@@ -2364,7 +2409,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	if (KVM_BUG_ON(ipa == INVALID_GPA, vcpu->kvm))
 		return -EFAULT;
 
-	is_iabt = kvm_vcpu_trap_is_iabt(vcpu);
+	is_iabt = esr_trap_is_iabt(esr);
 
 	if (esr_fsc_is_translation_fault(esr)) {
 		/* Beyond sanitised PARange (which is the IPA limit) */
@@ -2374,14 +2419,14 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		}
 
 		/* Falls between the IPA range and the PARange? */
-		if (fault_ipa >= BIT_ULL(VTCR_EL2_IPA(vcpu->arch.hw_mmu->vtcr))) {
+		if (fault_ipa >= BIT_ULL(VTCR_EL2_IPA(mmu->vtcr))) {
 			fault_ipa |= FAR_TO_FIPA_OFFSET(kvm_vcpu_get_hfar(vcpu));
 
 			return kvm_inject_sea(vcpu, is_iabt, fault_ipa);
 		}
 	}
 
-	trace_kvm_guest_fault(*vcpu_pc(vcpu), kvm_vcpu_get_esr(vcpu),
+	trace_kvm_guest_fault(*vcpu_pc(vcpu), esr,
 			      kvm_vcpu_get_hfar(vcpu), fault_ipa);
 
 	/* Check the stage-2 fault is trans. fault or write fault */
@@ -2389,10 +2434,10 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	    !esr_fsc_is_permission_fault(esr) &&
 	    !esr_fsc_is_access_flag_fault(esr) &&
 	    !esr_fsc_is_excl_atomic_fault(esr)) {
-		kvm_err("Unsupported FSC: EC=%#x xFSC=%#lx ESR_EL2=%#lx\n",
-			kvm_vcpu_trap_get_class(vcpu),
-			(unsigned long)kvm_vcpu_trap_get_fault(vcpu),
-			(unsigned long)kvm_vcpu_get_esr(vcpu));
+		kvm_err("Unsupported FSC: EC=%#lx xFSC=%#lx ESR_EL2=%#lx\n",
+			ESR_ELx_EC(esr),
+			(unsigned long)(esr & ESR_ELx_FSC),
+			(unsigned long)esr);
 		return -EFAULT;
 	}
 
@@ -2411,8 +2456,8 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	 * nothing to walk and we treat it as a 1:1 before going through the
 	 * canonical translation.
 	 */
-	if (kvm_is_nested_s2_mmu(vcpu->kvm,vcpu->arch.hw_mmu) &&
-	    vcpu->arch.hw_mmu->nested_stage2_enabled) {
+	if (kvm_is_nested_s2_mmu(vcpu->kvm, mmu) &&
+	    mmu->nested_stage2_enabled) {
 		u32 esr;
 
 		ret = kvm_walk_nested_s2(vcpu, fault_ipa, &nested_trans);
@@ -2441,7 +2486,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	gfn = ipa >> PAGE_SHIFT;
 	memslot = gfn_to_memslot(vcpu->kvm, gfn);
 	hva = gfn_to_hva_memslot_prot(memslot, gfn, &writable);
-	write_fault = kvm_is_write_fault(vcpu);
+	write_fault = esr_abt_is_write_fault(esr);
 	if (kvm_is_error_hva(hva) || (write_fault && !writable)) {
 		/*
 		 * The guest has put either its instructions or its page-tables
@@ -2454,7 +2499,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 			goto out;
 		}
 
-		if (kvm_vcpu_abt_iss1tw(vcpu)) {
+		if (esr_abt_is_s1ptw(esr)) {
 			ret = kvm_inject_sea_dabt(vcpu, kvm_vcpu_get_hfar(vcpu));
 			goto out_unlock;
 		}
@@ -2469,7 +2514,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		 * So let's assume that the guest is just being
 		 * cautious, and skip the instruction.
 		 */
-		if (kvm_is_error_hva(hva) && kvm_vcpu_dabt_is_cm(vcpu)) {
+		if (kvm_is_error_hva(hva) && esr_dabt_is_cm(esr)) {
 			kvm_incr_pc(vcpu);
 			ret = 1;
 			goto out_unlock;
@@ -2487,7 +2532,7 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 	}
 
 	/* Userspace should not be able to register out-of-bounds IPAs */
-	VM_BUG_ON(ipa >= kvm_phys_size(vcpu->arch.hw_mmu));
+	VM_BUG_ON(ipa >= kvm_phys_size(mmu));
 
 	if (esr_fsc_is_access_flag_fault(esr)) {
 		handle_access_fault(vcpu, fault_ipa);
@@ -2501,19 +2546,20 @@ int kvm_handle_guest_abort(struct kvm_vcpu *vcpu)
 		.nested		= nested,
 		.memslot	= memslot,
 		.hva		= hva,
+		.esr		= esr,
+		.mmu		= mmu,
 	};
 
 	if (kvm_vm_is_protected(vcpu->kvm)) {
 		ret = pkvm_mem_abort(&s2fd);
 	} else {
-		VM_WARN_ON_ONCE(kvm_vcpu_trap_is_permission_fault(vcpu) &&
-				!write_fault &&
-				!kvm_vcpu_trap_is_exec_fault(vcpu));
+		VM_WARN_ON_ONCE(kvm_s2_fault_is_perm(&s2fd) && !write_fault &&
+				!kvm_s2_fault_is_exec(&s2fd));
 
 		if (kvm_slot_has_gmem(memslot))
-			ret = gmem_abort(&s2fd);
+			ret = gmem_abort(&s2fd, NULL);
 		else
-			ret = user_mem_abort(&s2fd);
+			ret = user_mem_abort(&s2fd, NULL);
 	}
 
 	if (ret == 0)
@@ -2890,3 +2936,146 @@ void kvm_toggle_cache(struct kvm_vcpu *vcpu, bool was_enabled)
 
 	trace_kvm_toggle_cache(*vcpu_pc(vcpu), was_enabled, now_enabled);
 }
+
+/*
+ * Try to walk to the specified GPA in canonical mmu - if unmapped returns 0, if
+ * mapped returns the granule size, otherwise returns an error.
+ */
+static long kvm_walk_s2(struct kvm_pgtable *pgt,
+			gpa_t gpa, s8 *level)
+{
+	struct kvm *kvm = kvm_s2_mmu_to_kvm(pgt->mmu);
+	kvm_pte_t pte;
+	long ret;
+
+	guard(read_lock)(&kvm->mmu_lock);
+
+	ret = kvm_pgtable_get_leaf(pgt, gpa, &pte, level,
+				   KVM_PGTABLE_WALK_SHARED);
+	if (ret)
+		return ret;
+	/* Unpopulated, must fault. */
+	if (!kvm_pte_valid(pte))
+		return 0;
+	return kvm_granule_size(*level);
+}
+
+/* Synthesised data abort at specified page table level. */
+#define PRE_FAULT_ESR(level)				\
+	 ((ESR_ELx_EC_DABT_LOW << ESR_ELx_EC_SHIFT) |	\
+	  ESR_ELx_IL | ESR_ELx_FSC_FAULT_L(level))
+
+/* Retrieve either a read-only or a read/write hva. */
+static hva_t gfn_to_hva_memslot_read(struct kvm_memory_slot *slot, gfn_t gfn)
+{
+	return gfn_to_hva_memslot_prot(slot, gfn, /*writable=*/NULL);
+}
+
+static long __pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,
+			   gpa_t gpa, struct kvm_memory_slot *memslot, s8 level)
+{
+	const bool is_gmem = kvm_slot_has_gmem(memslot);
+	const gfn_t gfn = gpa_to_gfn(gpa);
+	const hva_t hva = is_gmem ? 0 : gfn_to_hva_memslot_read(memslot, gfn);
+	const struct kvm_s2_fault_desc s2fd = {
+		.vcpu		= vcpu,
+		.fault_ipa	= gpa,
+		.nested		= NULL,
+		.memslot	= memslot,
+		.hva		= hva,
+		.esr		= PRE_FAULT_ESR(level),
+		.mmu		= mmu,
+	};
+	struct kvm_s2_fault_result result = {};
+	long ret;
+
+	if (kvm_is_error_hva(hva))
+		return -EFAULT;
+
+	if (is_gmem)
+		ret = gmem_abort(&s2fd, &result);
+	else
+		ret = user_mem_abort(&s2fd, &result);
+	if (IS_ERR_VALUE(ret))
+		return ret;
+	return result.mapping_size;
+}
+
+static long pre_fault_s2(struct kvm_s2_mmu *mmu, struct kvm_vcpu *vcpu,
+			 gpa_t gpa, struct kvm_memory_slot *memslot)
+{
+	s8 level;
+	long ret;
+
+	/* Try a walk first. */
+	ret = kvm_walk_s2(mmu->pgt, gpa, &level);
+	if (ret)
+		return ret;
+	/* OK, have to fault page in. */
+	return __pre_fault_s2(mmu, vcpu, gpa, memslot, level);
+}
+
+static unsigned long
+pre_fault_bytes_consumed(gpa_t gpa, unsigned long granule_size,
+			 unsigned long bytes_remaining)
+{
+	/* Granules are always a power-of-2. */
+	const unsigned long granule_bytes_remaining =
+		granule_size - (gpa % granule_size);
+
+	return min(granule_bytes_remaining, bytes_remaining);
+}
+
+/* If you lose the race this many times, time to give up. */
+#define MAX_PRE_FAULT_RETRIES 3
+
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (is_protected_kvm_enabled())
+		return -EOPNOTSUPP;
+	if (!kvm_vcpu_initialized(vcpu))
+		return -ENOEXEC;
+
+	return 0;
+}
+
+/**
+ * kvm_arch_vcpu_pre_fault_memory - pre-fault stage-2 page tables for the
+ * specified GPA.
+ * @vcpu:	The VCPU pointer
+ * @range:	{gpa, size, flags} tuple
+ *
+ * The mapping performed is always best-effort - faulting in is necessarily
+ * racey. The ranges faulted in are canonical, nested page tables are ignored.
+ *
+ * @range->gpa specifies the GPA to pre-fault, @range->size specifies how many
+ * bytes remain to be pre-faulted and @range->flags is reserved and must be 0.
+ *
+ * Returns: the number of bytes the pre-fault consumed, or an error.
+ */
+long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
+				    struct kvm_pre_fault_memory *range)
+{
+	struct kvm *kvm = vcpu->kvm;
+	const u64 bytes_remaining = range->size;
+	struct kvm_s2_mmu *mmu = &kvm->arch.mmu; /* Canonical. */
+	struct kvm_memory_slot *memslot;
+	const gpa_t gpa = range->gpa;
+	int num_retries = 0;
+	long ret;
+
+	memslot = gfn_to_memslot(kvm, gpa_to_gfn(gpa));
+	if (!memslot)
+		return -ENOENT;
+	/* SRCU must be released for progress and only userland can do that. */
+	if (memslot->flags & KVM_MEMSLOT_INVALID)
+		return -EAGAIN;
+
+	do {
+		ret = pre_fault_s2(mmu, vcpu, gpa, memslot);
+	} while (ret == -EAGAIN && num_retries++ < MAX_PRE_FAULT_RETRIES);
+
+	if (IS_ERR_VALUE(ret))
+		return ret;
+	return pre_fault_bytes_consumed(gpa, ret, bytes_remaining);
+}
diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c
index ec754a865a006..cd868af63b1b8 100644
--- a/arch/arm64/kvm/nested.c
+++ b/arch/arm64/kvm/nested.c
@@ -662,7 +662,7 @@ static u8 get_guest_mapping_ttl(struct kvm_s2_mmu *mmu, u64 addr)
 		return 0;
 
 	tmp &= ~(sz - 1);
-	if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL))
+	if (kvm_pgtable_get_leaf(mmu->pgt, tmp, &pte, NULL, 0))
 		goto again;
 	if (!(pte & PTE_VALID))
 		goto again;
diff --git a/arch/s390/kvm/s390/s390.c b/arch/s390/kvm/s390/s390.c
index 5c73f43782a74..47fe032444f45 100644
--- a/arch/s390/kvm/s390/s390.c
+++ b/arch/s390/kvm/s390/s390.c
@@ -5784,6 +5784,14 @@ void kvm_arch_commit_memory_region(struct kvm *kvm, struct kvm_memory_slot *old,
 	s390_kvm_mmu_commit_memory_region(kvm, old, new, change);
 }
 
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (kvm_is_ucontrol(vcpu->kvm))
+		return -EINVAL;
+
+	return 0;
+}
+
 /**
  * kvm_arch_vcpu_pre_fault_memory() -- pre-fault and link gmap dat tables
  * @vcpu: the vcpu that shall appear to have generated the fault-in.
@@ -5810,9 +5818,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, struct kvm_pre_fault_
 	gpa_t end;
 	int rc;
 
-	if (kvm_is_ucontrol(vcpu->kvm))
-		return -EINVAL;
-
 	rc = kvm_s390_faultin_gfn(vcpu, NULL, &f);
 	if (rc == PGM_ADDRESSING)
 		return -ENOENT;
diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c
index 064ecc33b9267..c35fd2868c20b 100644
--- a/arch/x86/kvm/mmu/mmu.c
+++ b/arch/x86/kvm/mmu/mmu.c
@@ -5086,6 +5086,14 @@ static int kvm_tdp_page_prefault(struct kvm_vcpu *vcpu, gpa_t gpa,
 	}
 }
 
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (!vcpu->kvm->arch.pre_fault_allowed)
+		return -EOPNOTSUPP;
+
+	return 0;
+}
+
 long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				    struct kvm_pre_fault_memory *range)
 {
@@ -5095,9 +5103,6 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 	u64 end;
 	int r;
 
-	if (!vcpu->kvm->arch.pre_fault_allowed)
-		return -EOPNOTSUPP;
-
 	if (kvm_is_gfn_alias(vcpu->kvm, gpa_to_gfn(range->gpa)))
 		return -EINVAL;
 
diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h
index 03bfc92864b6e..39a4d345aeca8 100644
--- a/include/linux/kvm_host.h
+++ b/include/linux/kvm_host.h
@@ -2643,6 +2643,7 @@ void kvm_arch_gmem_invalidate_range(struct kvm *kvm, struct kvm_gfn_range *range
 #endif
 
 #ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu);
 long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				    struct kvm_pre_fault_memory *range);
 #endif
diff --git a/tools/testing/selftests/kvm/Makefile.kvm b/tools/testing/selftests/kvm/Makefile.kvm
index 6a1482e3a286b..908bdc7cf4f58 100644
--- a/tools/testing/selftests/kvm/Makefile.kvm
+++ b/tools/testing/selftests/kvm/Makefile.kvm
@@ -178,6 +178,7 @@ TEST_GEN_PROGS_arm64 += arm64/debug-exceptions
 TEST_GEN_PROGS_arm64 += arm64/hello_el2
 TEST_GEN_PROGS_arm64 += arm64/host_sve
 TEST_GEN_PROGS_arm64 += arm64/hypercalls
+TEST_GEN_PROGS_arm64 += arm64/nv_pre_fault_memory_test
 TEST_GEN_PROGS_arm64 += arm64/external_aborts
 TEST_GEN_PROGS_arm64 += arm64/mmio_sign_ext
 TEST_GEN_PROGS_arm64 += arm64/page_fault_test
@@ -205,6 +206,7 @@ TEST_GEN_PROGS_arm64 += guest_memfd_test
 TEST_GEN_PROGS_arm64 += mmu_stress_test
 TEST_GEN_PROGS_arm64 += rseq_test
 TEST_GEN_PROGS_arm64 += steal_time
+TEST_GEN_PROGS_arm64 += pre_fault_memory_test
 
 TEST_GEN_PROGS_s390 = $(TEST_GEN_PROGS_COMMON)
 TEST_GEN_PROGS_s390 += s390/memop
diff --git a/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c
new file mode 100644
index 0000000000000..09c1db3038566
--- /dev/null
+++ b/tools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c
@@ -0,0 +1,158 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * nv_pre_fault_memory_test - Test KVM_PRE_FAULT_MEMORY on a vCPU whose
+ * last-run context is nested.
+ *
+ * The guest enters vEL2, sets up its EL2 translation configuration into the
+ * real EL1 registers then ERETs to vEL1 and exits to userspace so its vCPU
+ * last-run context is nested backed by a shadow stage 2 MMU.
+ *
+ * Assert that pre-faulting ignores that and targets the canonical stage-2
+ * page tables only.
+ */
+#include "kvm_util.h"
+#include "processor.h"
+#include "test_util.h"
+#include "ucall.h"
+
+#include <asm/sysreg.h>
+#include <linux/sizes.h>
+
+#define TEST_MEM_SLOT		10
+#define TEST_MEM_SIZE		SZ_2M
+#define TEST_MEM_GPA		SZ_1G
+
+static void guest_el1_code(void)
+{
+	u64 offset;
+
+	GUEST_ASSERT_EQ(get_current_el(), 1);
+
+	/* Exit to userspace with the vEL1 (nested) context live. */
+	GUEST_SYNC(1);
+
+	/*
+	 * Touch the prefaulted range. vstage-2 is disabled, so the shadow
+	 * stage-2 is a 1:1 view of the canonical IPA space.
+	 */
+	for (offset = 0; offset < TEST_MEM_SIZE; offset += SZ_4K)
+		READ_ONCE(*(u64 *)(TEST_MEM_GPA + offset));
+
+	GUEST_DONE();
+}
+
+static void guest_code(void)
+{
+	u64 sp;
+
+	GUEST_ASSERT_EQ(get_current_el(), 2);
+
+	/*
+	 * Mirror the EL2 translation regime into the real EL1 registers so
+	 * that vEL1 runs on the test's stage-1 page tables. With E2H=1, the
+	 * _EL1 accessors read the EL2 registers, and the _EL12 accessors
+	 * write the real EL1 registers.
+	 */
+	write_sysreg_s(read_sysreg(sctlr_el1), SYS_SCTLR_EL12);
+	write_sysreg_s(read_sysreg(tcr_el1), SYS_TCR_EL12);
+	write_sysreg_s(read_sysreg(ttbr0_el1), SYS_TTBR0_EL12);
+	write_sysreg_s(read_sysreg(mair_el1), SYS_MAIR_EL12);
+	write_sysreg_s(read_sysreg(cpacr_el1), SYS_CPACR_EL12);
+
+	/* Run vEL1 on the same stack. */
+	asm volatile("mov %0, sp" : "=r"(sp));
+	write_sysreg(sp, sp_el1);
+
+	/*
+	 * Drop TGE so that vEL1 is a nested context rather than host EL0.
+	 * KVM backs it with a shadow stage-2 MMU even though vstage-2 is
+	 * disabled (HCR_EL2.VM=0).
+	 */
+	write_sysreg(read_sysreg(hcr_el2) & ~HCR_EL2_TGE, hcr_el2);
+	isb();
+
+	write_sysreg(PSR_MODE_EL1h | PSR_F_BIT | PSR_I_BIT | PSR_A_BIT |
+		     PSR_D_BIT, spsr_el2);
+	write_sysreg((u64)guest_el1_code, elr_el2);
+	asm volatile("eret");
+
+	GUEST_ASSERT(false);
+}
+
+static void pre_fault(struct kvm_vcpu *vcpu, u64 gpa, u64 size)
+{
+	struct kvm_pre_fault_memory range = {
+		.gpa = gpa,
+		.size = size,
+	};
+	int ret;
+
+	do {
+		ret = __vcpu_ioctl(vcpu, KVM_PRE_FAULT_MEMORY, &range);
+	} while ((!ret && range.size) ||
+		 (ret < 0 && (errno == EINTR || errno == EAGAIN)));
+
+	TEST_ASSERT(!ret, "KVM_PRE_FAULT_MEMORY failed, ret: %d errno: %d",
+		    ret, errno);
+	TEST_ASSERT_EQ(range.size, 0);
+}
+
+int main(void)
+{
+	struct kvm_vcpu_init init;
+	struct kvm_vcpu *vcpu;
+	struct kvm_vm *vm;
+	struct ucall uc;
+	u64 npages;
+
+	TEST_REQUIRE(test_supports_el2());
+	TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));
+
+	vm = vm_create(1);
+
+	kvm_get_default_vcpu_target(vm, &init);
+	init.features[0] |= BIT(KVM_ARM_VCPU_HAS_EL2);
+	vcpu = aarch64_vcpu_add(vm, 0, &init, guest_code);
+	kvm_arch_vm_finalize_vcpus(vm);
+
+	npages = TEST_MEM_SIZE / vm->page_size;
+	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, TEST_MEM_GPA,
+				    TEST_MEM_SLOT, npages, 0);
+	virt_map(vm, TEST_MEM_GPA, TEST_MEM_GPA, npages);
+
+	/* Run the guest until it has ERET'd from vEL2 to vEL1. */
+	vcpu_run(vcpu);
+	switch (get_ucall(vcpu, &uc)) {
+	case UCALL_SYNC:
+		TEST_ASSERT_EQ(uc.args[1], 1);
+		break;
+	case UCALL_ABORT:
+		REPORT_GUEST_ASSERT(uc);
+		break;
+	default:
+		TEST_FAIL("Unhandled ucall: %ld", uc.cmd);
+	}
+
+	/*
+	 * The vCPU's last-run context is vEL1, so its hw_mmu is a shadow
+	 * stage-2 MMU.
+	 *
+	 * Pre-faulting must ignore that and populate the canonical stage-2.
+	 */
+	pre_fault(vcpu, TEST_MEM_GPA, TEST_MEM_SIZE);
+
+	/* Resume at vEL1 and touch the prefaulted range. */
+	vcpu_run(vcpu);
+	switch (get_ucall(vcpu, &uc)) {
+	case UCALL_DONE:
+		break;
+	case UCALL_ABORT:
+		REPORT_GUEST_ASSERT(uc);
+		break;
+	default:
+		TEST_FAIL("Unhandled ucall: %ld", uc.cmd);
+	}
+
+	kvm_vm_free(vm);
+	return 0;
+}
diff --git a/tools/testing/selftests/kvm/pre_fault_memory_test.c b/tools/testing/selftests/kvm/pre_fault_memory_test.c
index c57631aab3d38..3082fe09b95fa 100644
--- a/tools/testing/selftests/kvm/pre_fault_memory_test.c
+++ b/tools/testing/selftests/kvm/pre_fault_memory_test.c
@@ -12,19 +12,29 @@
 #include <processor.h>
 #include <pthread.h>
 #include <ucall_common.h>
+#include <guest_modes.h>
 
 /* Arbitrarily chosen values */
-#define TEST_SIZE		(SZ_2M + PAGE_SIZE)
-#define TEST_NPAGES		(TEST_SIZE / PAGE_SIZE)
+#define TEST_BASE_SIZE		SZ_2M
 #define TEST_SLOT		10
 
+/* Storage of test info to share with guest code */
+struct test_config {
+	u64 page_size;
+	u64 test_size;
+	u64 test_num_pages;
+};
+
+static struct test_config test_config;
+
 static void guest_code(u64 base_gva)
 {
 	volatile u64 val __used;
+	struct test_config *config = &test_config;
 	int i;
 
-	for (i = 0; i < TEST_NPAGES; i++) {
-		u64 *src = (u64 *)(base_gva + i * PAGE_SIZE);
+	for (i = 0; i < config->test_num_pages; i++) {
+		u64 *src = (u64 *)(base_gva + i * config->page_size);
 
 		val = *src;
 	}
@@ -36,6 +46,7 @@ struct slot_worker_data {
 	struct kvm_vm *vm;
 	gpa_t gpa;
 	u32 flags;
+	enum vm_mem_backing_src_type mem_backing_src;
 	bool worker_ready;
 	bool prefault_ready;
 	bool recreate_slot;
@@ -56,14 +67,16 @@ static void *delete_slot_worker(void *__data)
 	while (!READ_ONCE(data->recreate_slot))
 		cpu_relax();
 
-	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, data->gpa,
-				    TEST_SLOT, TEST_NPAGES, data->flags);
+	vm_userspace_mem_region_add(vm, data->mem_backing_src, data->gpa,
+				    TEST_SLOT, test_config.test_num_pages, data->flags);
 
 	return NULL;
 }
 
 static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
-			     u64 size, u64 expected_left, bool private)
+			     u64 size, u64 expected_left,
+			     enum vm_mem_backing_src_type mem_backing_src,
+			     bool private)
 {
 	struct kvm_pre_fault_memory range = {
 		.gpa = base_gpa + offset,
@@ -74,6 +87,7 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 		.vm = vcpu->vm,
 		.gpa = base_gpa,
 		.flags = private ? KVM_MEM_GUEST_MEMFD : 0,
+		.mem_backing_src = mem_backing_src,
 	};
 	bool slot_recreated = false;
 	pthread_t slot_worker;
@@ -150,8 +164,8 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 	/*
 	 * Assert success if prefaulting the entire range should succeed, i.e.
 	 * complete with no bytes remaining.  Otherwise prefaulting should have
-	 * failed due to ENOENT (due to RET_PF_EMULATE for emulated MMIO when
-	 * no memslot exists).
+	 * failed due to ENOENT (no memslot exists for the GPA; on x86 this
+	 * surfaces via RET_PF_EMULATE).
 	 */
 	if (!expected_left)
 		TEST_ASSERT_VM_VCPU_IOCTL(!ret, KVM_PRE_FAULT_MEMORY, ret, vcpu->vm);
@@ -160,39 +174,82 @@ static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,
 					  KVM_PRE_FAULT_MEMORY, ret, vcpu->vm);
 }
 
-static void __test_pre_fault_memory(unsigned long vm_type, bool private)
+struct test_params {
+	unsigned long vm_type;
+	bool private;
+	enum vm_mem_backing_src_type mem_backing_src;
+};
+
+static void __test_pre_fault_memory(enum vm_guest_mode guest_mode, void *arg)
 {
-	gpa_t gpa, gva, alignment, guest_page_size;
+	gpa_t gpa, gva, alignment, guest_page_size, host_page_size;
+	gpa_t backing_src_pagesz, mem_page_size;
+	struct test_params *p = arg;
 	const struct vm_shape shape = {
-		.mode = VM_MODE_DEFAULT,
-		.type = vm_type,
+		.mode = guest_mode,
+		.type = p->vm_type,
 	};
 	struct kvm_vcpu *vcpu;
+	struct kvm_run *run;
 	struct kvm_vm *vm;
 	struct ucall uc;
 
+	pr_info("Testing guest mode: %s\n", vm_guest_mode_string(guest_mode));
+	pr_info("Testing memory backing src type: %s\n",
+		vm_mem_backing_src_alias(p->mem_backing_src)->name);
+
 	vm = vm_create_shape_with_one_vcpu(shape, &vcpu, guest_code);
 
-	alignment = guest_page_size = vm_guest_mode_params[VM_MODE_DEFAULT].page_size;
-	gpa = (vm->max_gfn - TEST_NPAGES) * guest_page_size;
+	guest_page_size = vm_guest_mode_params[guest_mode].page_size;
+	host_page_size = getpagesize();
+	backing_src_pagesz = get_backing_src_pagesz(p->mem_backing_src);
+	mem_page_size = max(host_page_size, backing_src_pagesz);
+
+	test_config.page_size = guest_page_size;
+	test_config.test_size = align_up(TEST_BASE_SIZE + test_config.page_size,
+					 mem_page_size);
+	test_config.test_num_pages = vm_calc_num_guest_pages(vm->mode, test_config.test_size);
+
+	gpa = (vm->max_gfn - test_config.test_num_pages) * test_config.page_size;
 	alignment = SZ_2M;
+	alignment = max(alignment, mem_page_size);
 	gpa = align_down(gpa, alignment);
 	gva = gpa & ((1ULL << (vm->va_bits - 1)) - 1);
 
-	vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, gpa, TEST_SLOT,
-				    TEST_NPAGES, private ? KVM_MEM_GUEST_MEMFD : 0);
-	virt_map(vm, gva, gpa, TEST_NPAGES);
+	vm_userspace_mem_region_add(vm, p->mem_backing_src,
+				    gpa, TEST_SLOT, test_config.test_num_pages,
+				    p->private ? KVM_MEM_GUEST_MEMFD : 0);
+	virt_map(vm, gva, gpa, test_config.test_num_pages);
+
+	if (p->private)
+		vm_mem_set_private(vm, gpa, test_config.test_size);
+
+	pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,
+			 p->mem_backing_src, p->private);
+	/* Retry the same range after the first prefault attempt. */
+	pre_fault_memory(vcpu, gpa, 0, test_config.test_size, 0,
+			 p->mem_backing_src, p->private);
+	pre_fault_memory(vcpu, gpa,
+			 test_config.test_size - host_page_size,
+			 host_page_size * 2, host_page_size,
+			 p->mem_backing_src, p->private);
+	pre_fault_memory(vcpu, gpa, test_config.test_size,
+			 host_page_size, host_page_size,
+			 p->mem_backing_src, p->private);
 
-	if (private)
-		vm_mem_set_private(vm, gpa, TEST_SIZE);
+	vcpu_args_set(vcpu, 1, gva);
 
-	pre_fault_memory(vcpu, gpa, 0, SZ_2M, 0, private);
-	pre_fault_memory(vcpu, gpa, SZ_2M, PAGE_SIZE * 2, PAGE_SIZE, private);
-	pre_fault_memory(vcpu, gpa, TEST_SIZE, PAGE_SIZE, PAGE_SIZE, private);
+	/* Export the shared variables to the guest. */
+	sync_global_to_guest(vm, test_config);
 
-	vcpu_args_set(vcpu, 1, gva);
 	vcpu_run(vcpu);
 
+	run = vcpu->run;
+	TEST_ASSERT(run->exit_reason == UCALL_EXIT_REASON,
+		    "Wanted %s, got exit reason: %u (%s)",
+		    exit_reason_str(UCALL_EXIT_REASON),
+		    run->exit_reason, exit_reason_str(run->exit_reason));
+
 	switch (get_ucall(vcpu, &uc)) {
 	case UCALL_ABORT:
 		REPORT_GUEST_ASSERT(uc);
@@ -207,24 +264,61 @@ static void __test_pre_fault_memory(unsigned long vm_type, bool private)
 	kvm_vm_free(vm);
 }
 
-static void test_pre_fault_memory(unsigned long vm_type, bool private)
+static void test_pre_fault_memory(unsigned long vm_type, enum vm_mem_backing_src_type backing_src,
+				  bool private)
 {
+	struct test_params p = {
+		.vm_type = vm_type,
+		.private = private,
+		.mem_backing_src = backing_src,
+	};
+
 	if (vm_type && !(kvm_check_cap(KVM_CAP_VM_TYPES) & BIT(vm_type))) {
 		pr_info("Skipping tests for vm_type 0x%lx\n", vm_type);
 		return;
 	}
 
-	__test_pre_fault_memory(vm_type, private);
+	for_each_guest_mode(__test_pre_fault_memory, &p);
+}
+
+static void help(char *name)
+{
+	puts("");
+	printf("usage: %s [-h] [-m mode] [-s mem-type]\n", name);
+	puts("");
+	guest_modes_help();
+	backing_src_help("-s");
+	puts("");
 }
 
 int main(int argc, char *argv[])
 {
+	enum vm_mem_backing_src_type backing = DEFAULT_VM_MEM_SRC;
+	int opt;
+
+	guest_modes_append_default();
+
+	while ((opt = getopt(argc, argv, "hm:s:")) != -1) {
+		switch (opt) {
+		case 'm':
+			guest_modes_cmdline(optarg);
+			break;
+		case 's':
+			backing = parse_backing_src_type(optarg);
+			break;
+		case 'h':
+		default:
+			help(argv[0]);
+			exit(0);
+		}
+	}
+
 	TEST_REQUIRE(kvm_check_cap(KVM_CAP_PRE_FAULT_MEMORY));
 
-	test_pre_fault_memory(0, false);
+	test_pre_fault_memory(0, backing, false);
 #ifdef __x86_64__
-	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, false);
-	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, true);
+	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, false);
+	test_pre_fault_memory(KVM_X86_SW_PROTECTED_VM, backing, true);
 #endif
 	return 0;
 }
diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index 65eb26a0520d8..41356577bee79 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -4353,7 +4353,7 @@ static int kvm_vcpu_ioctl_get_stats_fd(struct kvm_vcpu *vcpu)
 static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 				     struct kvm_pre_fault_memory *range)
 {
-	int idx;
+	int idx, err;
 	long r;
 	u64 full_size;
 
@@ -4365,6 +4365,14 @@ static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,
 	    range->gpa + range->size <= range->gpa)
 		return -EINVAL;
 
+	/*
+	 * Certain architectures (e.g. arm64) need to reject the ioctl 'early'
+	 * before vcpu_load().
+	 */
+	err = kvm_arch_pre_fault_allowed(vcpu);
+	if (err)
+		return err;
+
 	vcpu_load(vcpu);
 	idx = srcu_read_lock(&vcpu->kvm->srcu);
 

25/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 24453
							output: 36
							thoughts: 1157
						

Thoughts:
**Analyzing Potential Uninitialized Memory Risks in the Patch**

Okay, so the task is to determine if this patch introduces any uninitialized memory risks significant enough to warrant a KMSAN fuzzing session. The target architecture is amd64, but the changes span multiple architectures, with a significant focus on arm64 and some touches to x86 and s390.

First, I need to get a handle on the changes. The patch modifies quite a few files: documentation, architecture-specific headers and code (arm64, x86, s390), generic KVM code, and self-tests. The bulk of the work seems centered around implementing `KVM_CAP_PRE_FAULT_MEMORY` on arm64. It looks like fault handling is being refactored for arm64's KVM MMU, with some corresponding adjustments to the generic KVM code, x86, and s390.

Specifically, I see a new `kvm_arch_pre_fault_allowed` function being added and used. This function seems to be an architecture-specific check that's called *before* `vcpu_load()`. This is interesting – it's a gatekeeper, likely to prevent certain architectures from using the pre-fault memory feature if they're not ready for it. x86, s390, and arm64 now have implementations of this function, with arm64 having some involved changes around `struct kvm_s2_fault_desc` too.

On x86, it simply checks a flag, and s390 seems to reject the feature under `ucontrol` mode. The arm64 implementation will be more involved, of course.

The core of the functionality seems to reside in the `kvm_vcpu_pre_fault_memory` function, defined in `virt/kvm/kvm_main.c`. It takes a `struct kvm_pre_fault_memory` which contains the guest physical address (GPA) and size to pre-fault. My main concern is how this structure interacts with user space. I need to understand if there are any opportunities for uninitialized memory to be leaked from the kernel to userspace through this pre-faulting mechanism. I should check how the callers interact with it. Are there any data copies or memory accesses that might be vulnerable?

Therefore, I will now focus on:

1.  **`struct kvm_pre_fault_memory`**:  I need a detailed look at its definition and where it's populated.  Is there any chance uninitialized memory is being passed to it from userspace?
2.  **`kvm_vcpu_pre_fault_memory` callers**: Understanding the context in which this function is invoked is crucial. How is the `struct kvm_pre_fault_memory` populated and used within these calls?  Any potential data leakage in these contexts is a red flag.
3.  **Architecture-specific implementations**: I need to dig into the arm64, x86, and s390 implementations of `kvm_arch_pre_fault_allowed` in particular, as their differences might highlight potential areas for concern.  Are they handling memory initialization or validation correctly?
4.  **`kvm_s2_fault_desc` and helpers:** Because arm64 has significant changes relating to fault handling and this structure. I need to review this to understand whether any changes here can result in uninitialized memory reads/writes.



26/3 2026/09/25 22:23 tool grepper
0m Args:
{
  "Expression": "kvm_vcpu_pre_fault_memory",
  "PathPrefix": "virt/kvm/"
}

Results:
{
  "Output": "virt/kvm/kvm_main.c=4327=static int kvm_vcpu_ioctl_get_stats_fd(struct kvm_vcpu *vcpu)\n--\nvirt/kvm/kvm_main.c-4352-#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\nvirt/kvm/kvm_main.c:4353:static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\nvirt/kvm/kvm_main.c-4354-\t\t\t\t     struct kvm_pre_fault_memory *range)\n--\nvirt/kvm/kvm_main.c=4433=static long kvm_vcpu_ioctl(struct file *filp,\n--\nvirt/kvm/kvm_main.c-4665-\t\t\tbreak;\nvirt/kvm/kvm_main.c:4666:\t\tr = kvm_vcpu_pre_fault_memory(vcpu, \u0026range);\nvirt/kvm/kvm_main.c-4667-\t\t/* Pass back leftover range. */\n"
}

27/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 26681
							output: 40
							thoughts: 35
						

Thoughts:
Okay, let's dive into this. Here's my take on those thoughts:

**Dissecting `kvm_vcpu_ioctl` in `kvm_main.c`**

Alright, my brain is zeroing in on `kvm_vcpu_ioctl`, specifically around line 4660 in `virt/kvm/kvm_main.c`. That's where the core of the vCPU interaction with the KVM subsystem likely resides. I'm immediately thinking, what specific functionality is being triggered around that point? What are the key IOCTL codes that are being handled, and what operations are being performed on the vCPU context? I'll probably be looking for the crucial switch statements or function calls that are dispatched based on the IOCTL input. Is this where the critical path lies for handling guest memory access, register manipulation, or maybe even instruction injection? I need to understand the control flow here, because this is probably a crucial entry point for a lot of vCPU operations. I'll need to remember to check how this function handles error conditions and locking, as I know that’s a frequent source of performance issues or subtle bugs. Time to get my hands dirty and start digging in, line by line, and maybe correlate it with the Intel SDM or AMD BSP to identify the behavior.



28/3 2026/09/25 22:23 tool read-file
0m Args:
{
  "File": "virt/kvm/kvm_main.c",
  "FirstLine": 4655,
  "LineCount": 25
}

Results:
{
  "Contents": "4655:\t\tcase KVM_GET_STATS_FD: {\n4656:\t\t\tr = kvm_vcpu_ioctl_get_stats_fd(vcpu);\n4657:\t\t\tbreak;\n4658:\t\t}\n4659:\t#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\n4660:\t\tcase KVM_PRE_FAULT_MEMORY: {\n4661:\t\t\tstruct kvm_pre_fault_memory range;\n4662:\t\n4663:\t\t\tr = -EFAULT;\n4664:\t\t\tif (copy_from_user(\u0026range, argp, sizeof(range)))\n4665:\t\t\t\tbreak;\n4666:\t\t\tr = kvm_vcpu_pre_fault_memory(vcpu, \u0026range);\n4667:\t\t\t/* Pass back leftover range. */\n4668:\t\t\tif (copy_to_user(argp, \u0026range, sizeof(range)))\n4669:\t\t\t\tr = -EFAULT;\n4670:\t\t\tbreak;\n4671:\t\t}\n4672:\t#endif\n4673:\t\tdefault:\n4674:\t\t\tr = kvm_arch_vcpu_ioctl(filp, ioctl, arg);\n4675:\t\t}\n4676:\tout:\n4677:\t\tmutex_unlock(\u0026vcpu-\u003emutex);\n4678:\t\tkfree(fpu);\n4679:\t\tkfree(kvm_sregs);\n"
}

29/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 27423
							output: 23
							thoughts: 15
						

Thoughts:
Okay, let's dive into this. Here's my take on those thoughts, as if they're my own:

**Analyzing `struct kvm_pre_fault_memory`**

Alright, let's take a look at this `struct kvm_pre_fault_memory`. It's clearly a critical data structure, especially given its name – anything dealing with pre-fault handling is going to be central to performance and reliability in the KVM hypervisor. I need to understand exactly what information this struct is conveying and how it's being used within the fault handling path. My gut feeling is that it's likely involved in optimization; perhaps pre-fetching data, speculative access, or perhaps even tracking memory regions that are likely to be accessed. I expect it holds information needed for the hypervisor to efficiently manage the virtual machine's memory, especially when a page fault occurs. I'm going to want to check all the fields and trace the points where it is defined, used, and modified. Then I should follow any calling relationships to related functions. This should give me a solid handle on how it's working. I'll need to go beyond a superficial glance at the field names and really dig into the implications of how they interact with each other. This is crucial for understanding how the memory fault path behaves in the KVM.



30/3 2026/09/25 22:23 tool grepper
0m Args:
{
  "Expression": "struct kvm_pre_fault_memory"
}

Results:
{
  "Output": "Documentation/virt/kvm/api.rst=6486=See KVM_SET_USER_MEMORY_REGION2 for additional details.\n--\nDocumentation/virt/kvm/api.rst-6493-:Type: vcpu ioctl\nDocumentation/virt/kvm/api.rst:6494::Parameters: struct kvm_pre_fault_memory (in/out)\nDocumentation/virt/kvm/api.rst-6495-:Returns: 0 if at least one page is processed, \u003c 0 on error\n--\nDocumentation/virt/kvm/api.rst=6497=Errors:\n--\nDocumentation/virt/kvm/api.rst-6515-\nDocumentation/virt/kvm/api.rst:6516:  struct kvm_pre_fault_memory {\nDocumentation/virt/kvm/api.rst-6517-\t/* in/out */\n--\narch/arm64/kvm/mmu.c=3056=long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\narch/arm64/kvm/mmu.c:3057:\t\t\t\t    struct kvm_pre_fault_memory *range)\narch/arm64/kvm/mmu.c-3058-{\n--\narch/s390/kvm/s390/s390.c=5787=int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)\n--\narch/s390/kvm/s390/s390.c-5814- */\narch/s390/kvm/s390/s390.c:5815:long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, struct kvm_pre_fault_memory *range)\narch/s390/kvm/s390/s390.c-5816-{\n--\narch/x86/kvm/mmu/mmu.c=5097=long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\narch/x86/kvm/mmu/mmu.c:5098:\t\t\t\t    struct kvm_pre_fault_memory *range)\narch/x86/kvm/mmu/mmu.c-5099-{\n--\ninclude/linux/kvm_host.h=2647=long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\ninclude/linux/kvm_host.h:2648:\t\t\t\t    struct kvm_pre_fault_memory *range);\ninclude/linux/kvm_host.h-2649-#endif\n--\ninclude/uapi/linux/kvm.h=1662=struct kvm_create_guest_memfd {\n--\ninclude/uapi/linux/kvm.h-1667-\ninclude/uapi/linux/kvm.h:1668:#define KVM_PRE_FAULT_MEMORY\t_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)\ninclude/uapi/linux/kvm.h-1669-\ninclude/uapi/linux/kvm.h:1670:struct kvm_pre_fault_memory {\ninclude/uapi/linux/kvm.h-1671-\t__u64 gpa;\n--\ntools/include/uapi/linux/kvm.h=1658=struct kvm_create_guest_memfd {\n--\ntools/include/uapi/linux/kvm.h-1663-\ntools/include/uapi/linux/kvm.h:1664:#define KVM_PRE_FAULT_MEMORY\t_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)\ntools/include/uapi/linux/kvm.h-1665-\ntools/include/uapi/linux/kvm.h:1666:struct kvm_pre_fault_memory {\ntools/include/uapi/linux/kvm.h-1667-\t__u64 gpa;\n--\ntools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c=82=static void pre_fault(struct kvm_vcpu *vcpu, u64 gpa, u64 size)\ntools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c-83-{\ntools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c:84:\tstruct kvm_pre_fault_memory range = {\ntools/testing/selftests/kvm/arm64/nv_pre_fault_memory_test.c-85-\t\t.gpa = gpa,\n--\ntools/testing/selftests/kvm/pre_fault_memory_test.c=76=static void pre_fault_memory(struct kvm_vcpu *vcpu, u64 base_gpa, u64 offset,\n--\ntools/testing/selftests/kvm/pre_fault_memory_test.c-80-{\ntools/testing/selftests/kvm/pre_fault_memory_test.c:81:\tstruct kvm_pre_fault_memory range = {\ntools/testing/selftests/kvm/pre_fault_memory_test.c-82-\t\t.gpa = base_gpa + offset,\n--\nvirt/kvm/kvm_main.c=4353=static int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\nvirt/kvm/kvm_main.c:4354:\t\t\t\t     struct kvm_pre_fault_memory *range)\nvirt/kvm/kvm_main.c-4355-{\n--\nvirt/kvm/kvm_main.c=4433=static long kvm_vcpu_ioctl(struct file *filp,\n--\nvirt/kvm/kvm_main.c-4660-\tcase KVM_PRE_FAULT_MEMORY: {\nvirt/kvm/kvm_main.c:4661:\t\tstruct kvm_pre_fault_memory range;\nvirt/kvm/kvm_main.c-4662-\n"
}

31/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 29356
							output: 40
							thoughts: 25
						

Thoughts:
Okay, let's dive into this. Here's my thought process laid out:

**Examining `include/uapi/linux/kvm.h` - Near Line 1670**

Right, let's get down to brass tacks. I need to refresh my memory on something KVM-related, specifically within the userspace API definition. I'm going to jump into the `include/uapi/linux/kvm.h` header file. It's the definitive source for the user-space interface. Knowing the details in here is crucial for interacting with KVM at a lower level, so I’ll be carefully examining the code. I'm thinking I'll be interested around line 1670 - it’s probably a struct or a constant I need to reference. Let's see what’s there, it could relate to… well, who knows, let's find out! I'll be looking for something that relates to whatever issue I'm tackling right now.



32/3 2026/09/25 22:23 tool read-file
0m Args:
{
  "File": "include/uapi/linux/kvm.h",
  "FirstLine": 1665,
  "LineCount": 20
}

Results:
{
  "Contents": "1665:\t\t__u64 reserved[6];\n1666:\t};\n1667:\t\n1668:\t#define KVM_PRE_FAULT_MEMORY\t_IOWR(KVMIO, 0xd5, struct kvm_pre_fault_memory)\n1669:\t\n1670:\tstruct kvm_pre_fault_memory {\n1671:\t\t__u64 gpa;\n1672:\t\t__u64 size;\n1673:\t\t__u64 flags;\n1674:\t\t__u64 padding[5];\n1675:\t};\n1676:\t\n1677:\t#endif /* __LINUX_KVM_H */\n"
}

33/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 29837
							output: 40
							thoughts: 221
						

Thoughts:
**Analyzing the KVM Pre-Fault Memory Handling**

Okay, so I'm looking at this KVM code and my initial thought is, this `struct kvm_pre_fault_memory range` is getting a full copy from userspace using `copy_from_user`. So, right off the bat, whatever `range` holds is *already* initialized by the user. Excellent! No worries about uninitialized memory here.

Then, `kvm_vcpu_pre_fault_memory(vcpu, &range)` gets called, and my spidey-sense is telling me this function likely modifies `range.gpa` and `range.size`. That's normal behavior for this type of function, and it makes sense in the context. Crucially, the code *then* copies the *entire* updated `range` back to userspace via `copy_to_user`. I'm glad to see they're doing a full copy back, ensuring any potential changes made internally are reflected in userspace. And, thinking about it a bit more, since `copy_from_user` takes care of the whole size, there's no way padding could introduce an issue here. Good design, good design.

The important point is that this ioctl handler within `kvm_vcpu_ioctl` was already in the tree. The patch just tweaked `kvm_vcpu_pre_fault_memory` to call `kvm_arch_pre_fault_allowed(vcpu)` *before* `vcpu_load(vcpu)`. That's it! So, the real question I need to answer is about what the implications are of that call to `kvm_arch_pre_fault_allowed(vcpu)`. I guess that's my next step. Let's see what `kvm_vcpu_pre_fault_memory` is *really* doing...



34/3 2026/09/25 22:23 tool read-file
0m Args:
{
  "File": "virt/kvm/kvm_main.c",
  "FirstLine": 4350,
  "LineCount": 45
}

Results:
{
  "Contents": "4350:\t}\n4351:\t\n4352:\t#ifdef CONFIG_KVM_GENERIC_PRE_FAULT_MEMORY\n4353:\tstatic int kvm_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu,\n4354:\t\t\t\t\t     struct kvm_pre_fault_memory *range)\n4355:\t{\n4356:\t\tint idx, err;\n4357:\t\tlong r;\n4358:\t\tu64 full_size;\n4359:\t\n4360:\t\tif (range-\u003eflags)\n4361:\t\t\treturn -EINVAL;\n4362:\t\n4363:\t\tif (!PAGE_ALIGNED(range-\u003egpa) ||\n4364:\t\t    !PAGE_ALIGNED(range-\u003esize) ||\n4365:\t\t    range-\u003egpa + range-\u003esize \u003c= range-\u003egpa)\n4366:\t\t\treturn -EINVAL;\n4367:\t\n4368:\t\t/*\n4369:\t\t * Certain architectures (e.g. arm64) need to reject the ioctl 'early'\n4370:\t\t * before vcpu_load().\n4371:\t\t */\n4372:\t\terr = kvm_arch_pre_fault_allowed(vcpu);\n4373:\t\tif (err)\n4374:\t\t\treturn err;\n4375:\t\n4376:\t\tvcpu_load(vcpu);\n4377:\t\tidx = srcu_read_lock(\u0026vcpu-\u003ekvm-\u003esrcu);\n4378:\t\n4379:\t\tfull_size = range-\u003esize;\n4380:\t\tdo {\n4381:\t\t\tif (signal_pending(current)) {\n4382:\t\t\t\tr = -EINTR;\n4383:\t\t\t\tbreak;\n4384:\t\t\t}\n4385:\t\n4386:\t\t\tr = kvm_arch_vcpu_pre_fault_memory(vcpu, range);\n4387:\t\t\tif (WARN_ON_ONCE(r == 0 || r == -EIO))\n4388:\t\t\t\tbreak;\n4389:\t\n4390:\t\t\tif (r \u003c 0)\n4391:\t\t\t\tbreak;\n4392:\t\n4393:\t\t\trange-\u003esize -= r;\n4394:\t\t\trange-\u003egpa += r;\n"
}

35/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 31127
							output: 43
							thoughts: 396
						

Thoughts:
**Analyzing Architecture Implementations and Changes**

Okay, let's dive into these architecture implementations. The target is amd64, so my focus is immediately drawn there. I need to understand what's changed specifically for amd64, that x86-based architecture.

Alright, looking at `arch/x86/kvm/mmu/mmu.c`, I see a new addition: `kvm_arch_pre_fault_allowed`. Interesting. It's essentially a check:

```c
+int kvm_arch_pre_fault_allowed(struct kvm_vcpu *vcpu)
+{
+	if (!vcpu->kvm->arch.pre_fault_allowed)
+		return -EOPNOTSUPP;
+
+	return 0;
+}
```

It appears to be a check, and it appears the check was *moved* from `kvm_arch_vcpu_pre_fault_memory`. It's now called *before* `vcpu_load`. So, it's just a refactoring, effectively. The code now checks `vcpu->kvm->arch.pre_fault_allowed` before proceeding with the fault handling, which makes sense. Everything else in the x86 MMU code seems identical. No big surprises there.

Now, even though my primary concern is amd64, let's be thorough. I should glance at the arm64 changes, just in case there's something relevant. I'm always looking for potential interactions or overlooked implications.

Let's see what's changed for arm64. I'm looking at `arch/arm64/kvm/mmu.c` now. I see: `struct kvm_s2_fault_result result = {};` is initialized with `= {}`. Also, `s8 level;` passed by pointer to `kvm_walk_s2` -> `kvm_pgtable_get_leaf`, where `struct leaf_walk_data data;` is defined. This is interesting. Now, I need to check `leaf_walker` in `arch/arm64/kvm/hyp/pgtable.c`.



36/3 2026/09/25 22:23 tool read-file
0m Args:
{
  "File": "arch/arm64/kvm/hyp/pgtable.c",
  "FirstLine": 280,
  "LineCount": 40
}

Results:
{
  "Contents": " 280:\t\n 281:\t\treturn r;\n 282:\t}\n 283:\t\n 284:\tstruct leaf_walk_data {\n 285:\t\tkvm_pte_t\tpte;\n 286:\t\ts8\t\tlevel;\n 287:\t};\n 288:\t\n 289:\tstatic int leaf_walker(const struct kvm_pgtable_visit_ctx *ctx,\n 290:\t\t\t       enum kvm_pgtable_walk_flags visit)\n 291:\t{\n 292:\t\tstruct leaf_walk_data *data = ctx-\u003earg;\n 293:\t\n 294:\t\tdata-\u003epte   = ctx-\u003eold;\n 295:\t\tdata-\u003elevel = ctx-\u003elevel;\n 296:\t\n 297:\t\treturn 0;\n 298:\t}\n 299:\t\n 300:\tint kvm_pgtable_get_leaf(struct kvm_pgtable *pgt, u64 addr,\n 301:\t\t\t\t kvm_pte_t *ptep, s8 *level,\n 302:\t\t\t\t enum kvm_pgtable_walk_flags flags)\n 303:\t{\n 304:\t\tstruct leaf_walk_data data;\n 305:\t\tstruct kvm_pgtable_walker walker = {\n 306:\t\t\t.cb\t= leaf_walker,\n 307:\t\t\t.flags\t= flags | KVM_PGTABLE_WALK_LEAF,\n 308:\t\t\t.arg\t= \u0026data,\n 309:\t\t};\n 310:\t\tint ret;\n 311:\t\n 312:\t\tret = kvm_pgtable_walk(pgt, ALIGN_DOWN(addr, PAGE_SIZE),\n 313:\t\t\t\t       PAGE_SIZE, \u0026walker);\n 314:\t\tif (!ret) {\n 315:\t\t\tif (ptep)\n 316:\t\t\t\t*ptep  = data.pte;\n 317:\t\t\tif (level)\n 318:\t\t\t\t*level = data.level;\n 319:\t\t}\n"
}

37/3 2026/09/25 22:23 llm kmsan-evaluator
0m Model:
gemini-3.8-flash

Args:
null

Results:
null

Tokens:
							input: 32604
							output: 305
							thoughts: 956
						

Thoughts:
**KMSAN Assessment: No New Risks Detected**

Okay, let's break this down. My analysis indicates that `NeedsKMSAN` is `false`. There are no apparent new avenues for uninitialized memory exposure.

The core of the patch involves refactoring and extending pre-faulting logic, specifically for stage-2 page tables. On `amd64`, the primary change is the factoring out of a capability check (`kvm_arch_pre_fault_allowed()`) from `kvm_arch_vcpu_pre_fault_memory()`. This allows the check to occur before a `vcpu_load()` call. The ARM64 and s390 changes are more substantial, implementing or updating pre-faulting logic for their respective architectures, along with associated selftests and documentation.

Crucially, from a memory safety standpoint, I see no introduction of risks. The `kvm_pre_fault_memory` ioctl argument is fully populated from user space *before* it's read by the kernel, thus avoiding an issue there. All local structs relevant to the pre-faulting operations, such as `struct kvm_s2_fault_result` and `struct kvm_s2_fault_desc`, are explicitly initialized, either using `{}` or with field-by-field initialization. The code doesn't introduce any new data length manipulations or buffer bound changes that could expose uninitialized memory. Moreover, on the x86 side, it's essentially a refactoring of an early capability check. I can't find a single instance where an uninitialized value might leak to user space.

Therefore, given the explicit initialization of all relevant data structures and the limited scope of the changes, this patch doesn't warrant a dedicated KMSAN fuzzing session. Standard sanitizers and KASAN should be sufficient to detect any memory-related issues.
`set-results`



38/3 2026/09/25 22:23 tool set-results
0m Args:
{
  "KMSANReasoning": "The patch series adds support for KVM_PRE_FAULT_MEMORY on arm64 and refactors the generic ioctl path (virt/kvm/kvm_main.c) to invoke a new architecture hook kvm_arch_pre_fault_allowed() before vcpu_load(). On the target architecture (amd64), the change simply moves the existing pre-fault allowance check (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed) from kvm_arch_vcpu_pre_fault_memory() to kvm_arch_pre_fault_allowed().\n\nNo uninitialized memory hazards are introduced:\n1. No uninitialized memory is returned to userspace: KVM_PRE_FAULT_MEMORY reads struct kvm_pre_fault_memory fully from userspace via copy_from_user() before mutating its fields and copying it back.\n2. In the arm64 implementation, all local structs (such as kvm_s2_fault_desc and kvm_s2_fault_result) are fully initialized before use.\n3. No conditional branches rely on uninitialized values, and no padding leaks or uninitialized page/heap reads are exposed.\n\nStandard bug detectors and KASAN are fully sufficient to test these changes. A dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

Results:
{
  "KMSANReasoning": "The patch series adds support for KVM_PRE_FAULT_MEMORY on arm64 and refactors the generic ioctl path (virt/kvm/kvm_main.c) to invoke a new architecture hook kvm_arch_pre_fault_allowed() before vcpu_load(). On the target architecture (amd64), the change simply moves the existing pre-fault allowance check (!vcpu-\u003ekvm-\u003earch.pre_fault_allowed) from kvm_arch_vcpu_pre_fault_memory() to kvm_arch_pre_fault_allowed().\n\nNo uninitialized memory hazards are introduced:\n1. No uninitialized memory is returned to userspace: KVM_PRE_FAULT_MEMORY reads struct kvm_pre_fault_memory fully from userspace via copy_from_user() before mutating its fields and copying it back.\n2. In the arm64 implementation, all local structs (such as kvm_s2_fault_desc and kvm_s2_fault_result) are fully initialized before use.\n3. No conditional branches rely on uninitialized values, and no padding leaks or uninitialized page/heap reads are exposed.\n\nStandard bug detectors and KASAN are fully sufficient to test these changes. A dedicated KMSAN fuzzing session is not warranted.",
  "NeedsKMSAN": false
}

LLM Calls Summary:
Total Calls Total Tokens Avg Tokens Total Duration (Seconds) Avg Duration (Seconds)
Tool Calls Summary:
Total Calls Total Duration (Seconds) Avg Duration (Seconds)