On Wed, 20 May 2026 01:59:41 -0300, Desnes Nunes wrote: > On Mon, May 18, 2026 at 3:33 AM Michal Pecio wrote: > > > The chip IOMMU faults shortly after setting USBCMD.RUN = 1. > > > Such fault is expected to cause HSE assertion and usually it does. > > > You will probably find that HSE is already set while Enable Slot > > > is being queued, even if it was clear in xhci_gen_setup(). > > I've just read HSE at these places and confirmed that HSE was already > set even before queuing the enable slot trb, even though it was > previously clear in xhci_gen_setup(). Makes sense, thanks for checking. > The fault addresses do not appear in the main log, nor anywhere other > than the DMAR fault addr messages in the crashkernel's log. > > However, by comparing the previous log messages from the past kernel, > to the ones I saw with the new kernel I built today, I noticed the > same 8K displacement from the fault addr. Maybe an iommu driver bug > clue? If the bug is deterministic it should be fairly easy to nail it down. Attached xhci debugfs patch adds a list of almost all memory the xHC is allowed to access, including (if I havevn't missed something) all mappings it is allowed to access before any slots are enabled. Please apply, reboot and then: zip -r before.zip /sys/kernel/debug/usb/xhci/0000:80:14.0 # trigger crash kexec and the bug zip -r after.zip /sys/kernel/debug/usb/xhci/0000:80:14.0 Note: if you need to use tar instead, copy the directory and then archive the copy, because tar doesn't work on debugfs directly. And since the bug may be an out of bounds access by the HW, if you don't mind running slightly experimental patch to a critical subsystem, please also apply the DMA guard pages patch. I've been using it for a few months without issues, but YMMV. It helps determine which mapping is accessed OOB. Note: DMA guard pages may casue USB to stop working before kexec if it's a HW bug masked by memory layout, or begin to work after kexec in case of some IOMMU subsystem issues. > PS: there was a big iommu PR a few days ago - all the results from on > this email were performed with a recent 7.1.0-rc3 kernel checked out > at 30e0ff6d6a83. Well, so those IOMMU changes neither broke it nor fixed it. Knowing xHCI I frankly suspect that the problem is here. Regards, Michal