Nullify have_run_cpus when freeing the mask, particularly in the error path of __sev_guest_init(), so that KVM doesn't have to subtly use sev->active to track whether or not the mask has been freed. As pointed out by Sashiko, blindly freeing the mask in sev_vm_destroy() results in a double-free if the mask is freed if __sev_guest_init() fails. Throw the logic in a helper as nullifying the pointer is frustratingly difficult and weird due to have_run_cpus being a single-entry array when CPUMASK_OFFSTACK=n. Deliberately don't use CPUMASK_VAR_NULL, as it's not directly assignable when the cpumask is on-stack, e.g. requires using a local variable and a memcpy(), which is beyond ridiculous. Furthermore, while clearing the on-stack bitmask is an unnecessary and arguably unwanted side effect, KVM absolutely relies on '0' being the "null" value given that the struct is zero-allocated. Opportunistically add an alloc() helper to pair with free(); there are just enough call sites to make doing so worthwhile. Fixes: 12c1f6e03f94 ("KVM: SEV: Free have_run_cpus during VM destruction even if VM is no longer SEV") Cc: stable@vger.kernel.org Reported-by: Sashiko Bot Closes: https://lore.kernel.org/all/20260923165349.CAAF01F000FF@smtp.kernel.org Signed-off-by: Sean Christopherson --- arch/x86/kvm/svm/sev.c | 33 ++++++++++++++++++++++----------- 1 file changed, 22 insertions(+), 11 deletions(-) diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c index 3448d56520c6..d3a2e6a51efc 100644 --- a/arch/x86/kvm/svm/sev.c +++ b/arch/x86/kvm/svm/sev.c @@ -488,6 +488,20 @@ static void snp_guest_req_cleanup(struct kvm *kvm) sev->guest_resp_buf = NULL; } +static int sev_alloc_have_run_cpus(struct kvm_sev_info *sev) +{ + if (!zalloc_cpumask_var(&sev->have_run_cpus, GFP_KERNEL_ACCOUNT)) + return -ENOMEM; + + return 0; +} + +static void sev_free_have_run_cpus(struct kvm_sev_info *sev) +{ + free_cpumask_var(sev->have_run_cpus); + memset(&sev->have_run_cpus, 0, sizeof(sev->have_run_cpus)); +} + static int __sev_guest_init(struct kvm *kvm, struct kvm_sev_cmd *argp, struct kvm_sev_init *data, unsigned long vm_type) @@ -545,10 +559,9 @@ static int __sev_guest_init(struct kvm *kvm, struct kvm_sev_cmd *argp, if (ret) goto e_free_asid; - if (!zalloc_cpumask_var(&sev->have_run_cpus, GFP_KERNEL_ACCOUNT)) { - ret = -ENOMEM; + ret = sev_alloc_have_run_cpus(sev); + if (ret) goto e_free_asid; - } /* This needs to happen after SEV/SNP firmware initialization. */ if (snp_active) { @@ -566,7 +579,7 @@ static int __sev_guest_init(struct kvm *kvm, struct kvm_sev_cmd *argp, return 0; e_free: - free_cpumask_var(sev->have_run_cpus); + sev_free_have_run_cpus(sev); e_free_asid: argp->error = init_args.error; sev_asid_free(sev); @@ -2191,10 +2204,9 @@ int sev_vm_move_enc_context_from(struct kvm *kvm, unsigned int source_fd) * does not, i.e. KVM could skip flushes if memory is reclaimed from * the old VM but not the new VM. */ - if (!zalloc_cpumask_var(&dst_sev->have_run_cpus, GFP_KERNEL_ACCOUNT)) { - ret = -ENOMEM; + ret = sev_alloc_have_run_cpus(dst_sev); + if (ret) goto out_source_vcpu; - } sev_migrate_from(kvm, source_kvm); kvm_vm_dead(source_kvm); @@ -2888,10 +2900,9 @@ int sev_vm_copy_enc_context_from(struct kvm *kvm, unsigned int source_fd) } mirror_sev = to_kvm_sev_info(kvm); - if (!zalloc_cpumask_var(&mirror_sev->have_run_cpus, GFP_KERNEL_ACCOUNT)) { - ret = -ENOMEM; + ret = sev_alloc_have_run_cpus(mirror_sev); + if (ret) goto e_unlock; - } /* * The mirror kvm holds an enc_context_owner ref so its asid can't @@ -2984,7 +2995,7 @@ void sev_vm_destroy(struct kvm *kvm) * Free the mask even if the VM is not *currently* an SEV VM, as it may * have been an SEV VM prior to intra-host migration. */ - free_cpumask_var(sev->have_run_cpus); + sev_free_have_run_cpus(sev); if (!sev_guest(kvm)) return; -- 2.56.0.rc1.315.gc6ed9934b7-goog Explicitly flush caches after intra-host migration before clearing "SEV active" on the source VM, as doing cache maintenance afterwards creates a tiny window where memory reclaim could return memory to the host without performing a cache flush, e.g. as pointed out by Sashiko: CPU1 in sev_migrate_from(): src->active = false; CPU2 running concurrent unmap: Since active is false, the automatic cache flush in sev_guest_memory_reclaimed is skipped. The host frees and reallocates the page. CPU1 in sev_migrate_from(): sev_writeback_caches(src_kvm); Executes a hardware cache flush (wbnoinvd), which writes the guest's old dirty ciphertext over the new page owner's data. Fixes: 93de2a6a4b91 ("KVM: SEV: Do cache maintenance on the source VM during intra-host migration") Cc: stable@vger.kernel.org Reported-by: Sashiko Bot Closes: https://lore.kernel.org/all/20260923165304.1662E1F000FF@smtp.kernel.org Signed-off-by: Sean Christopherson --- arch/x86/kvm/svm/sev.c | 13 +++++++------ 1 file changed, 7 insertions(+), 6 deletions(-) diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c index d3a2e6a51efc..0c1ebb16cec6 100644 --- a/arch/x86/kvm/svm/sev.c +++ b/arch/x86/kvm/svm/sev.c @@ -2045,6 +2045,13 @@ static void sev_migrate_from(struct kvm *dst_kvm, struct kvm *src_kvm) struct kvm_sev_info *mirror; unsigned long i; + /* + * Do cache maintenance on the source VM *before* clearing "SEV active", + * as memory reclaim flows won't trigger cache maintenance on the VM + * once it's no longer an SEV VM. + */ + sev_writeback_caches(src_kvm); + dst->active = true; dst->asid = src->asid; dst->handle = src->handle; @@ -2058,12 +2065,6 @@ static void sev_migrate_from(struct kvm *dst_kvm, struct kvm *src_kvm) src->pages_locked = 0; src->es_active = false; - /* - * Do cache maintenance on the source VM as it is no longer an SEV VM, - * i.e. memory reclaim flows won't trigger cache maintenance on the VM. - */ - sev_writeback_caches(src_kvm); - list_cut_before(&dst->regions_list, &src->regions_list, &src->regions_list); mutex_lock(&sev_mirror_lock); -- 2.56.0.rc1.315.gc6ed9934b7-goog