From: Jim Mattson When a vCPU is destroyed while L2 is active, KVM synthesizes a nested VM-Exit, which flushes the cached shadow VMCS12 back to guest memory: vmx_vcpu_free() |-> nested_vmx_free_vcpu() |-> vmx_leave_nested() |-> nested_vmx_vmexit(vcpu, -1, 0, 0) |-> nested_flush_cached_shadow_vmcs12() |-> kvm_write_guest_cached() |-> __copy_to_user(ghc->hva, ...) Accessing user memory via a memslot during VM destruction is broken, as there are no guarantees that current->mm == kvm->mm when the VM is dying, because the last reference to the VM can be put from a different process than the original creating processes. And even if the original process does put the final reference, during process exit, do_exit() calls exit_mm() before closing file descriptors, so vCPU destruction runs with current->mm == NULL on a borrowed lazy TLB active_mm. If the borrowed address space has a writable mapping at the to-be-written userspace address, KVM will corrupt an unrelated task's memory since uaccess APIs, including __copy_to_user(), don't sanity check current->mm (and *can't* sanity perform KVM's current->mm == kvm->mm check since that is firmly a KVM-only concept). Hack-a-fix the nVMX flow even though KVM now protects against bad uaccess reads/writes in the core APIs, as doing so will allow adding even more sanity checks in KVM's APIs to help detect other buggy code. Add a TODO to call out that checking if KVM can do a uaccess for the VM is a hack; nVMX really needs to stop abusing __nested_vmx_vmexit() when destroying a vCPU. Fixes: 61ada7488ffd ("KVM: nVMX: Cache shadow vmcs12 on VMEntry and flush to memory on VMExit") Signed-off-by: Jim Mattson [sean: key off __kvm_can_do_uaccess(), add TODO] Signed-off-by: Sean Christopherson --- arch/x86/kvm/vmx/nested.c | 7 ++++++- include/linux/kvm_host.h | 8 ++++++-- 2 files changed, 12 insertions(+), 3 deletions(-) diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c index 151873407abd..8e31eba4d9fa 100644 --- a/arch/x86/kvm/vmx/nested.c +++ b/arch/x86/kvm/vmx/nested.c @@ -5134,8 +5134,13 @@ void __nested_vmx_vmexit(struct kvm_vcpu *vcpu, u32 vm_exit_reason, * Otherwise, this flush will dirty guest memory at a * point it is already assumed by user-space to be * immutable. + * + * TODO: Drop the explicit check on being able to access guest + * memory once KVM no longer abuses the nested VM-Exit + * flow when destroying a vCPU. */ - nested_flush_cached_shadow_vmcs12(vcpu, vmcs12); + if (__kvm_can_do_uaccess(vcpu->kvm)) + nested_flush_cached_shadow_vmcs12(vcpu, vmcs12); } else { /* * The only expected VM-instruction error is "VM entry with diff --git a/include/linux/kvm_host.h b/include/linux/kvm_host.h index 0ee81754d730..0c58a4945595 100644 --- a/include/linux/kvm_host.h +++ b/include/linux/kvm_host.h @@ -1350,10 +1350,14 @@ int kvm_write_guest_offset_cached(struct kvm *kvm, struct gfn_to_hva_cache *ghc, int kvm_gfn_to_hva_cache_init(struct kvm *kvm, struct gfn_to_hva_cache *ghc, gpa_t gpa, unsigned long len); +static __always_inline __must_check bool __kvm_can_do_uaccess(struct kvm *kvm) +{ + return current->mm == kvm->mm && refcount_read(&kvm->users_count); +} + static __always_inline __must_check bool kvm_can_do_uaccess(struct kvm *kvm) { - return !WARN_ON_ONCE(current->mm != kvm->mm || - !refcount_read(&kvm->users_count)); + return !WARN_ON_ONCE(!__kvm_can_do_uaccess(kvm)); } #define BUILD_KVM_COPY_USER_WRAPPER(fn, to_user, from_user) \ -- 2.56.0.rc1.315.gc6ed9934b7-goog