Stops in the Windows hypervisor

When VBS runs, the GDB stub reports what each vCPU executed when it halted. An idle vCPU is usually inside the Windows hypervisor, with the hypervisor’s own CR3. ntoseye names such a stop by the image in which the vCPU stopped. The context shows hypervisor, or VTL1 for the secure kernel. As in WinDbg, the module name of the hypervisor image (hvix64.exe) is hv, so code and stack frames show hv+0x…. While the context of the hypervisor is selected, expressions accept hv and hv+<offset> like any other module name.

hvix64 has no public symbols, but ntoseye names some of its code: each hypercall handler by the TLFS name of the lowest call code it serves (hv!HvCallGetVpRegisters), or HvCall and the code when the TLFS does not name it (hv!HvCall0004), the handler of every unimplemented code hv!HvCallUnimplemented, and the VM-exit entry point from the eVMCS pages hv!VmExitEntry. k, u, ln, x hv!*, and expressions use these names. A name covers only its own function, as the image’s .pdata bounds it, so code in the other functions still shows hv+0x…. A leaf function has no .pdata entry, so its name ends where the next function that .pdata lists begins.

Microsoft does not publish symbols or images for the hypervisor, and its address space does not map its unwind data (.pdata). With a copy of the running build’s hvix64.exe, ntoseye unwinds its stacks from the file’s unwind data, and reads a function’s prolog where that data leaves out its stack allocation, as it does for a few assembly functions from build 22621 on (how): copy C:\Windows\System32\hvix64.exe from the guest and run .fetchimage /f hvix64.exe once per hypervisor build, or put the file in a local directory on the symbol path. The file also bounds the names that ntoseye gives the hypervisor’s functions. Without the file, ntoseye unwinds each frame by reading its function’s prolog ([prolog]), which agrees with the real unwind data at 99.8% to 99.9% of call sites in the builds it was measured on, and it falls back to a stack scan ([scan]) where the prolog does not decide. A walk ends at the VM-exit entry point, where the hypervisor’s stack begins, and the scan takes only addresses in the hypervisor’s image, because NT is not mapped in its address space, each once (how).

Because the CR3 of the hypervisor does not map NT memory, ntoseye inspects such a stop at the point where NT left off (below). bp and the bugcheck trap work in all address spaces in which the vCPU can stop (how).

NT and the secure kernel call the hypervisor through the hypercall page that each one’s HvcallCodeVa points to. ntoseye names that page’s code as the module hvcall: the hypercall itself (hvcall!Hypercall), and VTL call and VTL return for x64 and x86 callers (hvcall!VtlCall64, hvcall!VtlReturn64, hvcall!VtlCall32, hvcall!VtlReturn32), which it finds by their code. A saved state that left off in a hypercall then reads saved VTL0 hvcall!Hypercall, and VTL1 waiting in a VTL return reads VTL1 hvcall!VtlReturn64+0xd. VTL0’s page is named at the first stop in the hypervisor, and VTL1’s once the secure kernel’s symbols are loaded.

Because Microsoft does not publish symbols for the hypervisor, ntoseye names only some of its code, and its stacks are exact only with a copy of its file (how).

Important

When Windows runs its own hypervisor (VBS, Hyper-V, WSL2), the gdb backend steps without the trap flag, because the trap flag can freeze such a guest (how). Steps and resumes from breakpoints work as usual, including through syscall, sysret, iretq, int, hypercalls, and far transfers.

The one difference is that a step sometimes ends early, in an interrupt handler or in another thread that the handler switched to. This happens in about 2 of 1,000 steps in kernel code, and ntoseye then shows a notice. Use g to resume.

Step-until walks (pa, ta, pc, tc, and the SDK’s until= and run_to(step=)) and call traces (wt and trace_calls()) do not end early at such a point. They run until their thread has executed the instruction, and then continue.

Where NT left off under the hypervisor

When a vCPU halts in the Windows hypervisor, the hypervisor keeps the state of that vCPU’s VTLs in its own memory. ntoseye reads this state from Enlightened VMCS pages, whose layout the Hyper-V TLFS defines, so this method does not depend on a hypervisor build. The hypervisor uses these pages only if the VM gives the hv-evmcs enlightenment (KVM/QEMU setup). A plain nested VMCS has a CPU-private format, and KVM keeps it out of guest memory.

ntoseye then inspects a stop in the hypervisor at the point where VTL0 left off. The stop header shows this point (saved VTL0 nt!HalProcessorIdle+0xf) with its registers, code, and stack, ~ shows it on a line below the vCPU, and r, k, u, memory reads, and expressions use the NT state on that processor. The same applies when you switch to such a vCPU with ~Ns (~) or with a bare .thread, and when a DAP client gets the stack of each vCPU that is halted in the hypervisor.

.cxr goes back to the registers, stack, and address space of the hypervisor, and .vtlcxr selects the NT context again. .vtlcxr also shows the saved state of each VTL and the reason why that VTL last entered the hypervisor (HLT, VMCALL, …). For a VMCALL whose registers ntoseye has, it decodes the hypercall from RCX: its call code and TLFS name, and whether it is a fast or a rep call (hypercall 0x000b HvCallSendSyntheticClusterIpi fast). The stop header shows the same, and !hvcall shows the call’s input (how).

A processor that runs a guest partition’s VP enters the hypervisor for that VP’s exits, not the root’s. At such a stop, the stop header, ~, and .vtlcxr name the VP that the hypervisor serves (serving partition 0x7 VP 2), with where it left off, its last exit, and the hypercall it made.

The other clients show the same: the MCP trailer and the thread names of DAP and gdbserver clients add where each VTL left off with its hypercall and the VP that the processor serves (p01.02 [hypervisor] hv!HvCallFlushVirtualAddressList (VTL0 hvcall!Hypercall (hypercall 0x0003 HvCallFlushVirtualAddressList rep 0/12))), and the SDK has cpu.serving and saved.hypercall (Python SDK).

The saved context has RIP, RSP, flags, control registers, and segment registers from the eVMCS. The eVMCS does not hold the other general-purpose registers: the hypervisor’s VM-exit entry code saves them itself. ntoseye reads that code to find where (how) and adds RAX to R15 to the context of the current VTL, so r, expressions, and stack walks that need a frame pointer have them. This is experimental. The registers are missing, and .vtlcxr says why, for a VTL that is not the current one, while the vCPU is on the entry point or still saving them, and when the entry code does not save them in one place before it first branches. .vtlcxr 1 selects the saved state of VTL1 instead, in the secure kernel’s address space, whose symbols it loads first as .vtl 1 does, so k walks the secure kernel’s stack from where VTL1 left off, usually its VTL return (hvcall!VtlReturn64+0xd, then securekernel!SkpReturnFromNormalMode). Like a VTL1 stop, this context is read-only, and .vtlcxr goes back to the VTL0 state. ntoseye resolves the RIP of VTL1 after it finds the secure kernel (.vtl 1).

Because such a vCPU runs hypervisor code, ntoseye does not step it, and t, p, gu, wt, and the other step commands give an error. g resumes the vCPU. The registers of the vCPU are read-only in all selected contexts. If there is no saved state, which happens when the VM does not have hv-evmcs or on an AMD host, the stop stays in the context of the hypervisor.

A vCPU can also stop on the first instruction of the hypervisor’s VM-exit handler (the eVMCS host_rip). KVM writes the saved state when it enters the hypervisor, and a stop from outside, such as a break-in, can fall between a VM exit and that entry, so the saved state may still describe the previous exit. Under load this is common. ntoseye marks such a state (may be one exit behind) in the stop header, ~, and .vtlcxr, and does not select it or start stacks from it on its own. .vtlcxr still selects it when you ask. The guest’s general-purpose registers at that point are the vCPU’s own registers. A breakpoint on host_rip fires only after KVM has written the state, so at a stop on such a breakpoint the state is current, and ntoseye selects it with the vCPU’s own general-purpose registers.

The NT thread that runs on that processor also starts from the same saved state everywhere that ntoseye shows this thread:

  • !thread shows the stack of the thread from this point (k-stack (saved VTL0 context)).

  • .thread selects this state as the register context of the thread.

  • The SDK’s Thread.backtrace() also starts from this state.

  • For running threads, the stacks that !running -t, !stacks, !process, and !analyze -hang show also start from this state.

ntoseye finds the Enlightened VMCS pages with a scan of host RAM, once per boot, the first time that a stop, ~, or .vtlcxr finds a vCPU in the hypervisor. For an 8 GiB guest, the scan takes about 0.6 s. The saved state of VTL1 is recognized by the secure kernel’s image, so the first of these stops also finds the secure kernel, as the first .vtl 1 does. If ntoseye does not find it then, it does not look again during that boot, and it does not show VTL1’s saved state. If the scan found no pages for a vCPU, which a stop early in the boot can cause, ntoseye scans once more for that vCPU. If a saved state fails validation, ntoseye does not show it (how).

ntoseye does not support AMD hosts for this feature, because the Windows hypervisor uses eVMCS only on Intel (VMX). On AMD, its nested state is a VMCB.