Battle-tested config for passing through a Sapphire Pulse RX 9070 XT to a Windows 10/11 guest on a single-GPU, no-iGPU system. Everything below was debugged across dozens of reboots — if it says "don't do X," someone already tried it and got burned.
| CPU | AMD Ryzen 5 5600X (Zen 3, no iGPU) |
| GPU | AMD Radeon RX 9070 XT (Navi 48, RDNA 4) — 1002:7550 + 1002:ab40 |
| Board | ASUS ROG STRIX B550-F GAMING |
| Host | CachyOS (Arch-based), kernel linux-cachyos 7.0.x |
| Bootloader | Limine + mkinitcpio |
| Desktop | KDE Plasma (Wayland) via plasmalogin |
| Guest | Windows 10/11 on Q35 + OVMF (UEFI) |
- No reset bug on Navi 48 (unlike RDNA 1-3). The card resets cleanly via PSP on kernel ≥ 6.13.
- Do NOT install
vendor-resetDKMS. It tops out at RDNA 1 and won't compile on modern kernels. - Do NOT pre-bind the GPU to
vfio-pciat boot. Unlike older AMD cards, the 9070 XT must be initialized byamdgpufirst. Pre-binding causes the "stuck in D3" error that plagues every forum thread. Letamdgpuown the card at boot, then hand it tovfio-pciat VM start. - Do NOT use
disable_idle_d3=1. It makes D3 issues worse on Navi 48, not better. - BAR2 must be resized to 8 MB before binding to
vfio-pci. The Windows AMD driver chokes on BAR2 > 8 MB and throws Code 43.
| Setting | Value |
|---|---|
| SVM Mode | Enabled |
| IOMMU | Enabled (not Auto) |
| Above 4G Decoding | Enabled |
| Re-Size BAR Support | Enabled |
| CSM | Disabled (UEFI only) |
ReBAR + Above 4G are mandatory for the 9070 XT to enumerate its 16 GiB BAR0 cleanly.
CachyOS uses Limine. Edit /etc/default/limine, then sudo limine-update.
iommu=pt initcall_blacklist=sysfb_init video=efifb:off pcie_aspm=off pci=noaer
iommu=pt— passthrough mode (lower DMA overhead). On Ryzen with BIOS IOMMU=Enabled,amd_iommu=onis implicit and not needed in the cmdline.initcall_blacklist=sysfb_init— stopssimpledrmfrom grabbing the GPU early, which blocks the unbind later.pcie_aspm=off— disables PCIe Active State Power Management; prevents the GPU from entering low-power link states that can interfere with the bind/unbind dance.pci=noaer— suppresses PCIe Advanced Error Reporting noise during GPU hand-off.- Do NOT add
pcie_acs_overrideunless your IOMMU groups actually need it (check §4 below).
options kvm_amd nested=1
options kvm ignore_msrs=1
options kvm report_ignored_msrs=0
nested=1— enables nested SVM so the guest can run its own hypervisor (e.g. for RE work). If you don't need nested virt, you can omit this.ignore_msrs=1— prevents KVM from injecting #GP on unimplemented MSR accesses (Windows reads many model-specific registers KVM doesn't emulate).
MODULES=(vfio_pci vfio vfio_iommu_type1)
HOOKS=(base systemd autodetect microcode kms modconf block keyboard sd-vconsole plymouth filesystems)
The HOOKS line shown here uses systemd-based initramfs (CachyOS default). If you use the traditional busybox hooks, replace
systemd/sd-vconsolewithudev/consolefont/keymap.
vfio_virqfdwas folded intovfioin kernel 5.x. Don't list it.
Do not create a /etc/modprobe.d/vfio.conf with options vfio-pci ids=1002:7550,1002:ab40. That pre-binds the GPU to vfio at boot, which is exactly what breaks Navi 48. The hookscripts handle the bind/unbind dance.
Rebuild:
sudo mkinitcpio -P
sudo limine-updatesudo pacman -S qemu-full libvirt virt-manager dnsmasq edk2-ovmf swtpm
sudo systemctl enable --now libvirtd
sudo usermod -aG libvirt,input,kvm "$USER"for d in /sys/kernel/iommu_groups/*/devices/*; do
n=${d#*/iommu_groups/}; n=${n%%/*}
printf 'IOMMU Group %s ' "$n"; lspci -nns "${d##*/}"
done | sort -VYou want 1002:7550 (VGA) and 1002:ab40 (HDA) alone or together in a group. CachyOS ships the ACS override patch in the stock kernel if you need it.
This is the core issue everyone hits. The error looks like:
kvm: vfio: Unable to power on device, stuck in D3
When amdgpu unbinds from the GPU, the kernel's runtime PM + ACPI drops the card into D3cold within ~1 second. In D3cold, the PCIe config space returns 0xFFFF — the device is electrically powered off. When vfio-pci tries to bind or QEMU tries to access it, it reads 0xFFFF everywhere and reports "stuck in D3."
In your start script, before unbinding amdgpu:
echo 0 > /sys/bus/pci/devices/$GPU/d3cold_allowed
echo "on" > /sys/bus/pci/devices/$GPU/power/controlThis prevents the kernel from dropping the card into D3cold during the driverless window between amdgpu-unbind and vfio-pci-bind.
With managed='yes', libvirt's nodedev-detach fires a PM reset (D3hot→D0 cycle) that triggers D3cold on Navi 48. With managed='no', libvirt skips the detach — your script has already bound vfio-pci manually, and QEMU just uses the device as-is.
This runs when the VM starts. It tears down the host desktop, locks the GPU in D0, unbinds amdgpu, resizes BAR2, binds vfio-pci, and starts the VM.
#!/usr/bin/env bash
set -uo pipefail
VGA=$(lspci -D -d 1002:7550 | awk '{print $1}')
HDA=$(lspci -D -d 1002:ab40 | awk '{print $1}')
# 1. Stop the display manager (adapt to your DM: sddm, gdm, plasmalogin, etc.)
if systemctl is-active --quiet plasmalogin.service 2>/dev/null; then
systemctl stop plasmalogin.service || true
sleep 1
fi
# 2. Drop VT consoles and framebuffer
for vt in /sys/class/vtconsole/vtcon*/bind; do
[ -w "$vt" ] && echo 0 > "$vt" 2>/dev/null
done
for d in /sys/bus/platform/drivers/efi-framebuffer/efi-framebuffer.0 \
/sys/bus/platform/drivers/simple-framebuffer/simple-framebuffer.0; do
[ -e "$d" ] || continue
basename "$d" > "$(dirname "$d")/unbind" 2>/dev/null || true
done
# 3. Lock GPU in D0 BEFORE unbinding amdgpu (THE critical step)
echo 0 > /sys/bus/pci/devices/$VGA/d3cold_allowed
echo "on" > /sys/bus/pci/devices/$VGA/power/control
# 4. Unbind amdgpu and snd_hda_intel by PCI address
# Uses the device's own driver symlink, not the driver's global unbind —
# avoids races if multiple devices share the same driver.
echo "$VGA" > /sys/bus/pci/devices/$VGA/driver/unbind 2>/dev/null || true
echo "$HDA" > /sys/bus/pci/devices/$HDA/driver/unbind 2>/dev/null || true
sleep 1
# 5. Resize BAR2 to 8 MB (Navi 48 requirement)
echo 3 > /sys/bus/pci/devices/$VGA/resource2_resize
# 6. Bind to vfio-pci manually (managed='no' in XML)
echo vfio-pci > /sys/bus/pci/devices/$VGA/driver_override
echo "$VGA" > /sys/bus/pci/drivers/vfio-pci/bind
echo vfio-pci > /sys/bus/pci/devices/$HDA/driver_override
echo "$HDA" > /sys/bus/pci/drivers/vfio-pci/bind
# 7. (Optional) Hugepages + performance governor
echo 3 > /proc/sys/vm/drop_caches
echo 1 > /proc/sys/vm/compact_memory
echo 8192 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages
for c in /sys/devices/system/cpu/cpu*/cpufreq; do
echo performance > "$c/scaling_governor" 2>/dev/null || true
done
# 8. Start the VM
virsh define /path/to/your-runtime.xml > /dev/null
virsh start win10gaming || exit 1Important: Do not use
modprobe -r amdgpuinstead of unbinding by PCI address. If the module is used by anything (console, another GPU), it will refuse. The by-address unbind is the proven path from every working Proxmox/Unraid/Arch setup.
This runs when the VM shuts down. The critical detail everyone gets wrong: you must explicitly bind amdgpu by PCI address, not just modprobe amdgpu. The module is already loaded (it drove the console before the VM started), so modprobe is a no-op and the GPU sits driverless.
#!/usr/bin/env bash
VGA=$(lspci -D -d 1002:7550 | awk '{print $1}')
HDA=$(lspci -D -d 1002:ab40 | awk '{print $1}')
UPSTREAM=$(basename "$(readlink -f /sys/bus/pci/devices/$VGA/../)")
gpu_bound() { [ -e /dev/dri/card0 ] || [ -e /dev/dri/card1 ]; }
for attempt in 1 2 3; do
# Unbind from vfio-pci
echo "$VGA" > /sys/bus/pci/drivers/vfio-pci/unbind 2>/dev/null || true
echo "$HDA" > /sys/bus/pci/drivers/vfio-pci/unbind 2>/dev/null || true
# Clear driver_override (echo "", not ": >" — truncation doesn't trigger the write handler)
echo "" > /sys/bus/pci/devices/$VGA/driver_override 2>/dev/null || true
echo "" > /sys/bus/pci/devices/$HDA/driver_override 2>/dev/null || true
# On retry, do a secondary bus reset via the upstream bridge
if [ "$attempt" -gt 1 ] && [ -n "$UPSTREAM" ]; then
cur=$(setpci -s "$UPSTREAM" BRIDGE_CONTROL.w)
setpci -s "$UPSTREAM" BRIDGE_CONTROL.w=$(printf '%04x' $((0x$cur | 0x40)))
sleep 1
setpci -s "$UPSTREAM" BRIDGE_CONTROL.w="$cur"
sleep 1
fi
# Force D0
setpci -s "${VGA#0000:}" CAP_PM+4.w=0x0000 2>/dev/null || true
sleep 1
# Bind amdgpu BY PCI ADDRESS (modprobe alone does nothing)
modprobe amdgpu 2>/dev/null || true
echo "$VGA" > /sys/bus/pci/drivers/amdgpu/bind 2>/dev/null || true
sleep 3
if gpu_bound; then
break
fi
done
# Rebind audio
echo "$HDA" > /sys/bus/pci/drivers/snd_hda_intel/bind 2>/dev/null || true
# Restore VT consoles
for vt in /sys/class/vtconsole/vtcon*; do
[ -e "$vt/bind" ] && echo 1 > "$vt/bind" 2>/dev/null || true
done
# Restore governor + EPP + free hugepages
for c in /sys/devices/system/cpu/cpu*/cpufreq; do
echo powersave > "$c/scaling_governor" 2>/dev/null || true
[ -w "$c/energy_performance_preference" ] && \
echo balance_performance > "$c/energy_performance_preference" 2>/dev/null || true
done
echo 0 > /sys/kernel/mm/hugepages/hugepages-2048kB/nr_hugepages 2>/dev/null || true
# Restart display manager only if GPU actually came back
if gpu_bound; then
systemctl start display-manager.service # or plasmalogin.service, sddm, gdm, etc.
else
echo "GPU reclaim failed. SSH in and run reclaim manually, or reboot."
fiWhen the start script sets driver_override=vfio-pci, that value persists after vfio-pci unbinds. If you don't clear it, modprobe amdgpu and even explicit amdgpu/bind writes will be rejected — the kernel routes all probe attempts to vfio-pci because the override is still set.
You must write an empty string to driver_override to clear it:
echo "" > /sys/bus/pci/devices/$VGA/driver_overrideDo not use shell truncation (: > file). It truncates the sysfs file to zero bytes but doesn't invoke the kernel's write handler, so the in-memory override value stays set. This is the most subtle bug in the whole setup and the one that cost the most debugging time.
<hostdev mode='subsystem' type='pci' managed='no'>
<driver name='vfio'/>
<source>
<address domain='0x0000' bus='0x__' slot='0x00' function='0x0'/>
</source>
<address type='pci' domain='0x0000' bus='0x05' slot='0x00' function='0x0' multifunction='on'/>
</hostdev>
<hostdev mode='subsystem' type='pci' managed='no'>
<driver name='vfio'/>
<source>
<address domain='0x0000' bus='0x__' slot='0x00' function='0x1'/>
</source>
<address type='pci' domain='0x0000' bus='0x05' slot='0x00' function='0x1'/>
</hostdev>Warning: The GPU's PCI bus number can change across kernel versions (e.g.
0x09on one kernel,0x0bon another). Either resolve the address dynamically in your start script vialspci -D -d 1002:7550and sed-substitute it into a template XML at runtime, or hardcode it and remember to update when you switch kernels.
managed='no'— your script handles the bind/unbind dance, not libvirt.- No
<rom bar='off'/>— the VBIOS comes from ACPI VFCT. Withrombar=0, OVMF can't find the GOP driver and the monitor stays black. - No
<rom file='...'/>— not needed on kernel ≥ 6.13 with Navi 48.
<cpu mode='host-passthrough' check='none' migratable='off'>
<topology sockets='1' dies='1' clusters='1' cores='4' threads='2'/>
<cache mode='passthrough'/>
<feature policy='require' name='topoext'/>
<feature policy='disable' name='hypervisor'/>
</cpu>Plus a QEMU commandline override (because libvirt's -cpu gets overridden by qemu:commandline):
<qemu:commandline xmlns:qemu='http://libvirt.org/schemas/domain/qemu/1.0'>
<qemu:arg value='-cpu'/>
<qemu:arg value='host,migratable=off,topoext=on,svm=on,npt=on,nrip-save=on,vmcb-clean=on,flushbyasid=on,decodeassists=on,pause-filter=on,pfthreshold=on,v-vmsave-vmload=on,vgif=on,avx2=on,bmi1=on,bmi2=on,fma=on,monitor=off,hypervisor=off,hv-passthrough=on,kvm=off,pmu=off,kvm-hint-dedicated=on,host-cache-info=on,l3-cache=off,ucode-rev=0xYOURREV,hv-vendor-id=AuthenticAMD'/>
<qemu:arg value='-overcommit'/>
<qemu:arg value='cpu-pm=on'/>
<qemu:arg value='-fw_cfg'/>
<qemu:arg value='opt/ovmf/X-PciMmio64Mb,string=65536'/>
<qemu:arg value='-machine'/>
<qemu:arg value='x-oem-id=ALASKA,x-oem-table-id=A M I '/>
</qemu:commandline>Critical: if you use
-cpuinqemu:commandline, it completely overrides libvirt's generated-cpuline. All features from the XML<cpu>block are thrown away. You must put everything in theqemu:commandlineversion —hypervisor=off,hv-passthrough=on, all SVM flags if you need nested virt, everything.
Get your host's microcode revision:
grep microcode /proc/cpuinfo | head -1Add it to the -cpu line as ucode-rev=0xYOURREV. Without this, Windows caches the microcode revision on first boot and BSODs with MICROCODE_REVISION_MISMATCH if you ever boot a different host kernel that loads a different microcode update.
<qemu:arg value='-machine'/>
<qemu:arg value='x-oem-id=ALASKA,x-oem-table-id=A M I '/>Replaces the default BOCHS/BXPC OEM IDs in ACPI tables. No QEMU rebuild needed — this is a runtime flag since ~QEMU 8.x.
<qemu:arg value='-fw_cfg'/>
<qemu:arg value='opt/ovmf/X-PciMmio64Mb,string=65536'/>Expands OVMF's 64-bit MMIO window to 64 GiB so ReBAR can map the 9070 XT's BARs. Without this, ReBAR collapses to 256 MiB in-guest.
| Vector | Fix |
|---|---|
| CPUID hypervisor bit | hypervisor=off in -cpu |
| KVM signature | kvm=off + <kvm><hidden state='on'/></kvm> |
| Hyper-V vendor ID | hv-vendor-id=AuthenticAMD in -cpu |
| SMBIOS/DMI strings | <sysinfo type='smbios'> with values from dmidecode |
| MAC OUI | Set <mac address='XX:XX:XX:...'/> with a real OEM OUI, not 52:54:00 |
| Virtio device IDs | Use SATA/AHCI for disks, e1000e for network (no 1af4:) |
| ACPI OEM IDs | -machine x-oem-id=ALASKA,x-oem-table-id=A M I |
| Microcode revision | ucode-rev=0xYOURREV in -cpu |
After boot, verify in the guest:
- Task Manager → CPU → "Virtual machine" should say No
msinfo32→ no "hypervisor has been detected"- Device Manager → no
Red Hat VirtIOdevices
If Windows BSODs with this on boot, the cause is Windows' VBS/Hyper-V stack checking the CPU microcode revision. Fix by disabling it offline:
Mount the guest disk:
sudo modprobe nbd max_part=8
sudo qemu-nbd --connect=/dev/nbd0 -f raw /path/to/disk.img
sudo mount -t ntfs3 /dev/nbd0p3 /mnt/win
sudo mount -t vfat /dev/nbd0p1 /mnt/winefiDisable VBS + HVCI in SYSTEM hive (using hivexsh -w):
cd \ControlSet001\Control\DeviceGuard
setval 3
CachedDrtmAuthIndex
dword:0
RequireMicrosoftSignedBootChain
dword:0
EnableVirtualizationBasedSecurity
dword:0
commit
cd \ControlSet001\Control\DeviceGuard\Scenarios
add HypervisorEnforcedCodeIntegrity
cd HypervisorEnforcedCodeIntegrity
setval 2
Enabled
dword:0
WasEnabledBy
dword:0
commit
hivexsh gotcha:
setvalvalue names must be BARE (no quotes).setval Nreplaces ALL values at the current node with the next N name/value pairs — if the node has other values you want to keep, include them all.
Disable hypervisor launch in BCD. First find your default loader GUID:
hivexget /mnt/winefi/EFI/Microsoft/Boot/BCD '\Objects\{9dea862c-5cdd-4e70-acc1-f32b344d4795}\Elements\23000003'
Then set hypervisorlaunchtype (element 250000f0) to off:
cd \Objects\{YOUR-LOADER-GUID}\Elements
add 250000f0
cd 250000f0
setval 1
Element
hex:3:00,00,00,00,00,00,00,00
commit
Then pin ucode-rev in the -cpu line so it never happens again.
virsh destroywhile the GPU is active — force-kills QEMU mid-operation, corrupts the GPU's SMU firmware. The card becomes completely unresponsive (trn=2 ACK should not assert). Only a host reboot recovers it. Always shut down from inside Windows.- Rebooting the guest is safe —
<on_reboot>restart</on_reboot>does an internal QEMU reset. The GPU stays onvfio-pci, the revert script does NOT fire. - Guest BSOD — let it sit or auto-reboot. Do not
virsh destroyit. If you must kill it, SSH in and tryvirsh shutdown win10gaming --mode acpifirst.
If your onboard audio controller shares a PCI reset domain with another function (common on AMD B550/X570), QEMU 11.0.1+ will refuse to start with:
vfio: Cannot reset device, depends on group XX which is not owned
Workaround: don't include the audio device in the XML. Instead, hot-plug it after the VM starts:
virsh attach-device win10gaming /tmp/audio-hostdev.xml --liveThis bypasses QEMU's startup reset-domain check. The device attaches fine after QEMU is already running.
- Proxmox forums — "PVE9 + AMD 9070xt now work" — uzumo's hookscript pattern (amdgpu unbind → BAR2 resize → vfio bind)
- Proxmox forums — "WORKING Amd RX 9070xt support" — confirmed:
disable_idle_d3=1makes things worse - Reddit r/VFIO — "AMD Radeon RX 9070 (XT) Reset Bug" — confirmed: let amdgpu init first, don't pre-bind vfio
- Reddit r/VFIO — "RX 9070 XT passthrough on Proxmox" by drlokey — dual-GPU writeup confirming the "opposite rules" for RDNA 4 vs older AMD
- Level1Techs — "VFIO Pass-through working on 9070XT"
- GitHub Gist — gdesatrigraha "VFIO AMD RX 9070 XT GPU Passthrough"
- Arch Wiki — "PCI passthrough via OVMF"
I fell in love with you