* [PATCH v5 01/17] x86/alternative: Support alt_replace_call() with instructions after call
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
@ 2026-09-11 8:41 ` Juergen Gross
2026-09-11 8:41 ` [PATCH v5 02/17] coco/tdx: Rename MSR access helpers Juergen Gross
` (16 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:41 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin
Today alt_replace_call() requires the initial indirect call not to be
followed by any further instructions, including padding NOPs. In case
any replacement is longer than 6 bytes, a subsequent replacement of
the indirect call with a direct one will result in a crash.
Fix that by crashing only if the original instruction is less than
6 bytes long or not a known indirect call.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
arch/x86/kernel/alternative.c | 5 +++--
1 file changed, 3 insertions(+), 2 deletions(-)
diff --git a/arch/x86/kernel/alternative.c b/arch/x86/kernel/alternative.c
index 91b1cdd16569..21b344761ee2 100644
--- a/arch/x86/kernel/alternative.c
+++ b/arch/x86/kernel/alternative.c
@@ -525,6 +525,7 @@ noinstr void BUG_func(void)
}
EXPORT_SYMBOL(BUG_func);
+#define CALL_RIP_INSTRLEN 6
#define CALL_RIP_REL_OPCODE 0xff
#define CALL_RIP_REL_MODRM 0x15
@@ -542,7 +543,7 @@ static unsigned int alt_replace_call(u8 *instr, u8 *insn_buff, struct alt_instr
BUG();
}
- if (a->instrlen != 6 ||
+ if (a->instrlen < CALL_RIP_INSTRLEN ||
instr[0] != CALL_RIP_REL_OPCODE ||
instr[1] != CALL_RIP_REL_MODRM) {
pr_err("ALT_FLAG_DIRECT_CALL set for unrecognized indirect call\n");
@@ -554,7 +555,7 @@ static unsigned int alt_replace_call(u8 *instr, u8 *insn_buff, struct alt_instr
#ifdef CONFIG_X86_64
/* ff 15 00 00 00 00 call *0x0(%rip) */
/* target address is stored at "next instruction + disp". */
- target = *(void **)(instr + a->instrlen + disp);
+ target = *(void **)(instr + CALL_RIP_INSTRLEN + disp);
#else
/* ff 15 00 00 00 00 call *0x0 */
/* target address is stored at disp. */
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 02/17] coco/tdx: Rename MSR access helpers
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
2026-09-11 8:41 ` [PATCH v5 01/17] x86/alternative: Support alt_replace_call() with instructions after call Juergen Gross
@ 2026-09-11 8:41 ` Juergen Gross
2026-09-11 8:41 ` [PATCH v5 03/17] x86/msr: Minimize usage of native_*() msr access functions Juergen Gross
` (15 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:41 UTC (permalink / raw)
To: linux-kernel, x86, linux-coco, kvm
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Kiryl Shutsemau, Rick Edgecombe
In order to avoid a name clash with some general MSR access helpers
after a future MSR infrastructure rework, rename the TDX specific
helpers.
Signed-off-by: Juergen Gross <jgross@suse.com>
Reviewed-by: Kiryl Shutsemau <kas@kernel.org>
Reviewed-by: H. Peter Anvin (Intel) <hpa@zytor.com>
Reviewed-by: Rick Edgecombe <rick.p.edgecombe@intel.com>
---
arch/x86/coco/tdx/tdx.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/arch/x86/coco/tdx/tdx.c b/arch/x86/coco/tdx/tdx.c
index f904a636d449..29aa57fb4244 100644
--- a/arch/x86/coco/tdx/tdx.c
+++ b/arch/x86/coco/tdx/tdx.c
@@ -469,7 +469,7 @@ static void __cpuidle tdx_safe_halt(void)
raw_local_irq_enable();
}
-static int read_msr(struct pt_regs *regs, struct ve_info *ve)
+static int tdx_read_msr(struct pt_regs *regs, struct ve_info *ve)
{
struct tdx_module_args args = {
.r10 = TDX_HYPERCALL_STANDARD,
@@ -490,7 +490,7 @@ static int read_msr(struct pt_regs *regs, struct ve_info *ve)
return ve_instr_len(ve);
}
-static int write_msr(struct pt_regs *regs, struct ve_info *ve)
+static int tdx_write_msr(struct pt_regs *regs, struct ve_info *ve)
{
struct tdx_module_args args = {
.r10 = TDX_HYPERCALL_STANDARD,
@@ -841,9 +841,9 @@ static int virt_exception_kernel(struct pt_regs *regs, struct ve_info *ve)
case EXIT_REASON_HLT:
return handle_halt(ve);
case EXIT_REASON_MSR_READ:
- return read_msr(regs, ve);
+ return tdx_read_msr(regs, ve);
case EXIT_REASON_MSR_WRITE:
- return write_msr(regs, ve);
+ return tdx_write_msr(regs, ve);
case EXIT_REASON_CPUID:
return handle_cpuid(regs, ve);
case EXIT_REASON_EPT_VIOLATION:
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 03/17] x86/msr: Minimize usage of native_*() msr access functions
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
2026-09-11 8:41 ` [PATCH v5 01/17] x86/alternative: Support alt_replace_call() with instructions after call Juergen Gross
2026-09-11 8:41 ` [PATCH v5 02/17] coco/tdx: Rename MSR access helpers Juergen Gross
@ 2026-09-11 8:41 ` Juergen Gross
2026-09-11 8:41 ` [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up Juergen Gross
` (14 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:41 UTC (permalink / raw)
To: linux-kernel, x86, linux-hyperv, kvm
Cc: Juergen Gross, K. Y. Srinivasan, Haiyang Zhang, Wei Liu,
Dexuan Cui, Long Li, Thomas Gleixner, Ingo Molnar,
Borislav Petkov, Dave Hansen, H. Peter Anvin, Paolo Bonzini,
Vitaly Kuznetsov, Sean Christopherson, Boris Ostrovsky,
xen-devel
In order to prepare for some MSR access function reorg work, switch
most users of native_{read|write}_msr[_safe]() to the more generic
rdmsr*()/wrmsr*() variants.
For now this will have some intermediate performance impact with
paravirtualization configured when running on bare metal, but this
is a prereq change for the planned direct inlining of the rdmsr/wrmsr
instructions with this configuration.
The main reason for this switch is the planned move of the MSR trace
function invocation from the native_*() functions to the generic
rdmsr*()/wrmsr*() variants. Without this switch the users of the
native_*() functions would lose the related tracing entries.
Note that the Xen related MSR access functions will not be switched,
as these will be handled after the move of the trace hooks.
Signed-off-by: Juergen Gross <jgross@suse.com>
Acked-by: Sean Christopherson <seanjc@google.com>
Acked-by: Wei Liu <wei.liu@kernel.org>
Reviewed-by: H. Peter Anvin (Intel) <hpa@zytor.com>
---
arch/x86/hyperv/ivm.c | 2 +-
arch/x86/kernel/cpu/mshyperv.c | 4 ++--
arch/x86/kernel/kvmclock.c | 2 +-
arch/x86/kvm/svm/svm.c | 16 ++++++++--------
arch/x86/xen/pmu.c | 4 ++--
5 files changed, 14 insertions(+), 14 deletions(-)
diff --git a/arch/x86/hyperv/ivm.c b/arch/x86/hyperv/ivm.c
index 2ce4dfe53472..a74f121f2a02 100644
--- a/arch/x86/hyperv/ivm.c
+++ b/arch/x86/hyperv/ivm.c
@@ -328,7 +328,7 @@ int hv_snp_boot_ap(u32 apic_id, unsigned long start_ip, unsigned int cpu)
savesegment(ds, vmsa->ds.selector);
hv_populate_vmcb_seg(vmsa->ds, vmsa->gdtr.base);
- vmsa->efer = native_read_msr(MSR_EFER);
+ vmsa->efer = rdmsrq(MSR_EFER);
vmsa->cr4 = native_read_cr4();
vmsa->cr3 = __native_read_cr3();
diff --git a/arch/x86/kernel/cpu/mshyperv.c b/arch/x86/kernel/cpu/mshyperv.c
index e1388ed27384..53ac4ef53929 100644
--- a/arch/x86/kernel/cpu/mshyperv.c
+++ b/arch/x86/kernel/cpu/mshyperv.c
@@ -114,7 +114,7 @@ u64 hv_para_get_synic_register(unsigned int reg)
{
if (WARN_ON(!ms_hyperv.paravisor_present || !hv_is_synic_msr(reg)))
return ~0ULL;
- return native_read_msr(reg);
+ return rdmsrq(reg);
}
/*
@@ -124,7 +124,7 @@ void hv_para_set_synic_register(unsigned int reg, u64 val)
{
if (WARN_ON(!ms_hyperv.paravisor_present || !hv_is_synic_msr(reg)))
return;
- native_write_msr(reg, val);
+ wrmsrq(reg, val);
}
u64 hv_get_msr(unsigned int reg)
diff --git a/arch/x86/kernel/kvmclock.c b/arch/x86/kernel/kvmclock.c
index cb3d0ca1fa22..6ddef8b5426a 100644
--- a/arch/x86/kernel/kvmclock.c
+++ b/arch/x86/kernel/kvmclock.c
@@ -219,7 +219,7 @@ static void kvm_setup_secondary_clock(void)
void kvmclock_disable(void)
{
if (msr_kvm_system_time)
- native_write_msr(msr_kvm_system_time, 0);
+ wrmsrq(msr_kvm_system_time, 0);
}
static void __init kvmclock_init_mem(void)
diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c
index 91f5a5344529..5e5bbecb8020 100644
--- a/arch/x86/kvm/svm/svm.c
+++ b/arch/x86/kvm/svm/svm.c
@@ -412,12 +412,12 @@ static void svm_init_erratum_383(void)
return;
/* Use _safe variants to not break nested virtualization */
- if (native_read_msr_safe(MSR_AMD64_DC_CFG, &val))
+ if (rdmsrq_safe(MSR_AMD64_DC_CFG, &val))
return;
val |= (1ULL << 47);
- native_write_msr_safe(MSR_AMD64_DC_CFG, val);
+ wrmsrq_safe(MSR_AMD64_DC_CFG, val);
erratum_383_found = true;
}
@@ -470,8 +470,8 @@ static void svm_init_os_visible_workarounds(void)
return;
if (!this_cpu_has(X86_FEATURE_OSVW) ||
- native_read_msr_safe(MSR_AMD64_OSVW_ID_LENGTH, &len) ||
- native_read_msr_safe(MSR_AMD64_OSVW_STATUS, &status))
+ rdmsrq_safe(MSR_AMD64_OSVW_ID_LENGTH, &len) ||
+ rdmsrq_safe(MSR_AMD64_OSVW_STATUS, &status))
len = status = 0;
if (status == READ_ONCE(osvw_status) && len >= READ_ONCE(osvw_len))
@@ -2112,7 +2112,7 @@ static bool is_erratum_383(void)
if (!erratum_383_found)
return false;
- if (native_read_msr_safe(MSR_IA32_MC0_STATUS, &value))
+ if (rdmsrq_safe(MSR_IA32_MC0_STATUS, &value))
return false;
/* Bit 62 may or may not be set for this mce */
@@ -2123,11 +2123,11 @@ static bool is_erratum_383(void)
/* Clear MCi_STATUS registers */
for (i = 0; i < 6; ++i)
- native_write_msr_safe(MSR_IA32_MCx_STATUS(i), 0);
+ wrmsrq_safe(MSR_IA32_MCx_STATUS(i), 0);
- if (!native_read_msr_safe(MSR_IA32_MCG_STATUS, &value)) {
+ if (!rdmsrq_safe(MSR_IA32_MCG_STATUS, &value)) {
value &= ~(1ULL << 2);
- native_write_msr_safe(MSR_IA32_MCG_STATUS, value);
+ wrmsrq_safe(MSR_IA32_MCG_STATUS, value);
}
/* Flush tlb to evict multi-match entries */
diff --git a/arch/x86/xen/pmu.c b/arch/x86/xen/pmu.c
index 5f50a3ee08f5..37512df8b8f2 100644
--- a/arch/x86/xen/pmu.c
+++ b/arch/x86/xen/pmu.c
@@ -324,7 +324,7 @@ static u64 xen_amd_read_pmc(int counter)
u64 val;
msr = amd_counters_base + (counter * amd_msr_step);
- native_read_msr_safe(msr, &val);
+ rdmsrq_safe(msr, &val);
return val;
}
@@ -350,7 +350,7 @@ static u64 xen_intel_read_pmc(int counter)
else
msr = MSR_IA32_PERFCTR0 + counter;
- native_read_msr_safe(msr, &val);
+ rdmsrq_safe(msr, &val);
return val;
}
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (2 preceding siblings ...)
2026-09-11 8:41 ` [PATCH v5 03/17] x86/msr: Minimize usage of native_*() msr access functions Juergen Gross
@ 2026-09-11 8:41 ` Juergen Gross
2026-09-23 20:26 ` Shreshth Srivastava
2026-09-11 8:41 ` [PATCH v5 05/17] x86/hyperv: Switch from __rdmsr() to native_rdmsrq() Juergen Gross
` (13 subsequent siblings)
17 siblings, 1 reply; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:41 UTC (permalink / raw)
To: linux-kernel, x86, virtualization
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Ajay Kaher, Alexey Makhalov,
Broadcom internal kernel review list
In order to prepare paravirt inlining of the MSR access instructions
move the calls of MSR trace functions one function level up.
Introduce {read|write}_msr[_safe]() helpers allowing to have common
definitions in msr.h doing the trace calls.
Signed-off-by: Juergen Gross <jgross@suse.com>
Reviewed-by: H. Peter Anvin (Intel) <hpa@zytor.com>
---
V4:
- some modifications removed due to rebase
---
arch/x86/include/asm/msr.h | 79 ++++++++++++++++++++++-----------
arch/x86/include/asm/paravirt.h | 8 ++--
2 files changed, 57 insertions(+), 30 deletions(-)
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index 3b33d432bc24..266298b3d201 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -95,14 +95,7 @@ static __always_inline void native_wrmsrq(u32 msr, u64 val)
static inline u64 native_read_msr(u32 msr)
{
- u64 val;
-
- val = __rdmsr(msr);
-
- if (tracepoint_enabled(read_msr))
- do_trace_read_msr(msr, val, 0);
-
- return val;
+ return __rdmsr(msr);
}
static inline int native_read_msr_safe(u32 msr, u64 *p)
@@ -115,8 +108,6 @@ static inline int native_read_msr_safe(u32 msr, u64 *p)
_ASM_EXTABLE_TYPE_REG(1b, 2b, EX_TYPE_RDMSR_SAFE, %[err])
: [err] "=r" (err), EAX_EDX_RET(val, low, high)
: "c" (msr));
- if (tracepoint_enabled(read_msr))
- do_trace_read_msr(msr, EAX_EDX_VAL(val, low, high), err);
*p = EAX_EDX_VAL(val, low, high);
@@ -127,9 +118,6 @@ static inline int native_read_msr_safe(u32 msr, u64 *p)
static inline void notrace native_write_msr(u32 msr, u64 val)
{
native_wrmsrq(msr, val);
-
- if (tracepoint_enabled(write_msr))
- do_trace_write_msr(msr, val, 0);
}
/* Can be uninlined because referenced by paravirt */
@@ -143,8 +131,6 @@ static inline int notrace native_write_msr_safe(u32 msr, u64 val)
: [err] "=a" (err)
: "c" (msr), "0" ((u32)val), "d" ((u32)(val >> 32))
: "memory");
- if (tracepoint_enabled(write_msr))
- do_trace_write_msr(msr, val, err);
return err;
}
@@ -165,36 +151,77 @@ static inline u64 native_read_pmc(int counter)
#include <asm/paravirt.h>
#else
#include <linux/errno.h>
-
-/* Access to machine-specific registers (available on 586 and better only) */
-
-static __always_inline u64 rdmsrq(u32 msr)
+static __always_inline u64 read_msr(u32 msr)
{
return native_read_msr(msr);
}
-static inline void wrmsrq(u32 msr, u64 val)
+static __always_inline int read_msr_safe(u32 msr, u64 *p)
+{
+ return native_read_msr_safe(msr, p);
+}
+
+static __always_inline void write_msr(u32 msr, u64 val)
{
native_write_msr(msr, val);
}
-/* wrmsr with exception handling */
-static inline int wrmsrq_safe(u32 msr, u64 val)
+static __always_inline int write_msr_safe(u32 msr, u64 val)
{
return native_write_msr_safe(msr, val);
}
+static __always_inline u64 rdpmc(int counter)
+{
+ return native_read_pmc(counter);
+}
+#endif /* !CONFIG_PARAVIRT_XXL */
+
+/* Access to machine-specific registers (available on 586 and better only) */
+
+static __always_inline u64 rdmsrq(u32 msr)
+{
+ u64 val = read_msr(msr);
+
+ if (tracepoint_enabled(read_msr))
+ do_trace_read_msr(msr, val, 0);
+
+ return val;
+}
+
+/* rdmsr with exception handling */
static inline int rdmsrq_safe(u32 msr, u64 *p)
{
- return native_read_msr_safe(msr, p);
+ int err;
+
+ err = read_msr_safe(msr, p);
+
+ if (tracepoint_enabled(read_msr))
+ do_trace_read_msr(msr, *p, err);
+
+ return err;
}
-static __always_inline u64 rdpmc(int counter)
+static inline void wrmsrq(u32 msr, u64 val)
{
- return native_read_pmc(counter);
+ write_msr(msr, val);
+
+ if (tracepoint_enabled(write_msr))
+ do_trace_write_msr(msr, val, 0);
}
-#endif /* !CONFIG_PARAVIRT_XXL */
+/* wrmsr with exception handling */
+static inline int wrmsrq_safe(u32 msr, u64 val)
+{
+ int err;
+
+ err = write_msr_safe(msr, val);
+
+ if (tracepoint_enabled(write_msr))
+ do_trace_write_msr(msr, val, err);
+
+ return err;
+}
/* Instruction opcode for WRMSRNS supported in binutils >= 2.40 */
#define ASM_WRMSRNS _ASM_BYTES(0x0f,0x01,0xc6)
diff --git a/arch/x86/include/asm/paravirt.h b/arch/x86/include/asm/paravirt.h
index 19442bc3af37..a5a1fc4c88d1 100644
--- a/arch/x86/include/asm/paravirt.h
+++ b/arch/x86/include/asm/paravirt.h
@@ -150,22 +150,22 @@ static inline int paravirt_write_msr_safe(u32 msr, u64 val)
return PVOP_CALL2(int, pv_ops, cpu.write_msr_safe, msr, val);
}
-static __always_inline u64 rdmsrq(u32 msr)
+static __always_inline u64 read_msr(u32 msr)
{
return paravirt_read_msr(msr);
}
-static inline void wrmsrq(u32 msr, u64 val)
+static inline void write_msr(u32 msr, u64 val)
{
paravirt_write_msr(msr, val);
}
-static inline int wrmsrq_safe(u32 msr, u64 val)
+static inline int write_msr_safe(u32 msr, u64 val)
{
return paravirt_write_msr_safe(msr, val);
}
-static __always_inline int rdmsrq_safe(u32 msr, u64 *p)
+static __always_inline int read_msr_safe(u32 msr, u64 *p)
{
return paravirt_read_msr_safe(msr, p);
}
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up
2026-09-11 8:41 ` [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up Juergen Gross
@ 2026-09-23 20:26 ` Shreshth Srivastava
2026-09-25 13:15 ` Jürgen Groß
0 siblings, 1 reply; 23+ messages in thread
From: Shreshth Srivastava @ 2026-09-23 20:26 UTC (permalink / raw)
To: Juergen Gross, linux-kernel, x86, virtualization
Cc: tglx, mingo, bp, dave.hansen, hpa, ajay.kaher, alexey.makhalov,
bcm-kernel-feedback-list
On 11.09.26 10:41, Juergen Gross wrote:
> In order to prepare paravirt inlining of the MSR access instructions
> move the calls of MSR trace functions one function level up.
> Introduce {read|write}_msr[_safe]() helpers allowing to have common
> definitions in msr.h doing the trace calls.
Hi Juergen,
read_msr() and write_msr() get a single wrapper below the
CONFIG_PARAVIRT_XXL ifdef holding the tracepoint, so it fires for either
implementation. rdpmc() moved into the same ifdef but kept the old
arrangement: still defined twice, once per arm, with do_trace_rdpmc()
still called from native_read_pmc() above the ifdef.
With CONFIG_PARAVIRT_XXL=y that leaves rdpmc() on a path with no trace
call:
rdpmc() => paravirt-msr.h, PVOP_CALL1(pv_ops_msr,
read_pmc)
xen_read_pmc() => Xen PV sets pv_ops_msr.read_pmc to this
reads the value out of the Xen shared PMU page
=> returns without ever calling
native_read_pmc(), which is where
do_trace_rdpmc() sits
So msr:rdpmc doesn't fire under Xen PV, while msr:read_msr and
msr:write_msr now do.
Was that deliberate? If it wasn't, here is a diff that treats rdpmc the
same way as the other six: rename both definitions to read_pmc(), matching
the pv_ops_msr member they dispatch to, and add one rdpmc() below the
endif holding the tracepoint. native_read_pmc() is then an untraced
primitive next to native_rdmsrq() and native_wrmsrq(). Would something
like this help?
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index eba325ecfe4c..6f50b703fba9 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -305,8 +305,7 @@ static __always_inline u64 native_read_pmc(int counter)
EAX_EDX_DECLARE_ARGS(val, low, high);
asm volatile("rdpmc" : EAX_EDX_RET(val, low, high) : "c" (counter));
- if (tracepoint_enabled(rdpmc))
- do_trace_rdpmc(counter, EAX_EDX_VAL(val, low, high), 0);
+
return EAX_EDX_VAL(val, low, high);
}
@@ -343,7 +342,7 @@ static __always_inline int write_msrns_safe(u32 msr, u64 val)
return native_wrmsrns_safe(msr, val);
}
-static __always_inline u64 rdpmc(int counter)
+static __always_inline u64 read_pmc(int counter)
{
return native_read_pmc(counter);
}
@@ -413,6 +412,16 @@ static __always_inline int wrmsrns_safe(u32 msr, u64 val)
return err;
}
+static __always_inline u64 rdpmc(int counter)
+{
+ u64 val = read_pmc(counter);
+
+ if (tracepoint_enabled(rdpmc))
+ do_trace_rdpmc(counter, val, 0);
+
+ return val;
+}
+
static __always_inline void sync_cpu_after_wrmsrns(void)
{
if (cpu_feature_enabled(X86_FEATURE_WRMSRNS))
diff --git a/arch/x86/include/asm/paravirt-msr.h b/arch/x86/include/asm/paravirt-msr.h
index ba3ee64446db..47220bf16cf3 100644
--- a/arch/x86/include/asm/paravirt-msr.h
+++ b/arch/x86/include/asm/paravirt-msr.h
@@ -172,7 +172,7 @@ static __always_inline int write_msrns_safe(u32 msr, u64 val)
return err ? -EIO : 0;
}
-static __always_inline u64 rdpmc(int counter)
+static __always_inline u64 read_pmc(int counter)
{
return PVOP_CALL1(u64, pv_ops_msr, read_pmc, counter);
}
I have no Xen PV guest, so that path is untested. What I did check is that
the __tracepoint_rdpmc relocation shows up in arch/x86/events/core.o with
CONFIG_PARAVIRT_XXL=y, where it previously did not, and that x86_64
defconfig still builds clean with gcc and clang.
Thanks,
Shreshth
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up
2026-09-23 20:26 ` Shreshth Srivastava
@ 2026-09-25 13:15 ` Jürgen Groß
0 siblings, 0 replies; 23+ messages in thread
From: Jürgen Groß @ 2026-09-25 13:15 UTC (permalink / raw)
To: Shreshth Srivastava, linux-kernel, x86, virtualization
Cc: tglx, mingo, bp, dave.hansen, hpa, ajay.kaher, alexey.makhalov,
bcm-kernel-feedback-list
[-- Attachment #1.1.1: Type: text/plain, Size: 4122 bytes --]
On 23.09.26 22:26, Shreshth Srivastava wrote:
> On 11.09.26 10:41, Juergen Gross wrote:
>> In order to prepare paravirt inlining of the MSR access instructions
>> move the calls of MSR trace functions one function level up.
>> Introduce {read|write}_msr[_safe]() helpers allowing to have common
>> definitions in msr.h doing the trace calls.
>
> Hi Juergen,
>
> read_msr() and write_msr() get a single wrapper below the
> CONFIG_PARAVIRT_XXL ifdef holding the tracepoint, so it fires for either
> implementation. rdpmc() moved into the same ifdef but kept the old
> arrangement: still defined twice, once per arm, with do_trace_rdpmc()
> still called from native_read_pmc() above the ifdef.
>
> With CONFIG_PARAVIRT_XXL=y that leaves rdpmc() on a path with no trace
> call:
>
> rdpmc() => paravirt-msr.h, PVOP_CALL1(pv_ops_msr,
> read_pmc)
> xen_read_pmc() => Xen PV sets pv_ops_msr.read_pmc to this
> reads the value out of the Xen shared PMU page
> => returns without ever calling
> native_read_pmc(), which is where
> do_trace_rdpmc() sits
>
> So msr:rdpmc doesn't fire under Xen PV, while msr:read_msr and
> msr:write_msr now do.
>
> Was that deliberate? If it wasn't, here is a diff that treats rdpmc the
> same way as the other six: rename both definitions to read_pmc(), matching
> the pv_ops_msr member they dispatch to, and add one rdpmc() below the
> endif holding the tracepoint. native_read_pmc() is then an untraced
> primitive next to native_rdmsrq() and native_wrmsrq(). Would something
> like this help?
My patch doesn't change anything in this regard, as the Xen PV case didn't
write trace entries for read_pmc() before.
OTOH I agree that this is more like an oversight than a design decision,
so I'll do something along the lines you are suggesting below.
Thanks for the review!
>
> diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
> index eba325ecfe4c..6f50b703fba9 100644
> --- a/arch/x86/include/asm/msr.h
> +++ b/arch/x86/include/asm/msr.h
> @@ -305,8 +305,7 @@ static __always_inline u64 native_read_pmc(int counter)
> EAX_EDX_DECLARE_ARGS(val, low, high);
>
> asm volatile("rdpmc" : EAX_EDX_RET(val, low, high) : "c" (counter));
> - if (tracepoint_enabled(rdpmc))
> - do_trace_rdpmc(counter, EAX_EDX_VAL(val, low, high), 0);
> +
> return EAX_EDX_VAL(val, low, high);
> }
>
> @@ -343,7 +342,7 @@ static __always_inline int write_msrns_safe(u32 msr, u64 val)
> return native_wrmsrns_safe(msr, val);
> }
>
> -static __always_inline u64 rdpmc(int counter)
> +static __always_inline u64 read_pmc(int counter)
> {
> return native_read_pmc(counter);
> }
> @@ -413,6 +412,16 @@ static __always_inline int wrmsrns_safe(u32 msr, u64 val)
> return err;
> }
>
> +static __always_inline u64 rdpmc(int counter)
> +{
> + u64 val = read_pmc(counter);
> +
> + if (tracepoint_enabled(rdpmc))
> + do_trace_rdpmc(counter, val, 0);
> +
> + return val;
> +}
> +
> static __always_inline void sync_cpu_after_wrmsrns(void)
> {
> if (cpu_feature_enabled(X86_FEATURE_WRMSRNS))
> diff --git a/arch/x86/include/asm/paravirt-msr.h b/arch/x86/include/asm/paravirt-msr.h
> index ba3ee64446db..47220bf16cf3 100644
> --- a/arch/x86/include/asm/paravirt-msr.h
> +++ b/arch/x86/include/asm/paravirt-msr.h
> @@ -172,7 +172,7 @@ static __always_inline int write_msrns_safe(u32 msr, u64 val)
> return err ? -EIO : 0;
> }
>
> -static __always_inline u64 rdpmc(int counter)
> +static __always_inline u64 read_pmc(int counter)
> {
> return PVOP_CALL1(u64, pv_ops_msr, read_pmc, counter);
> }
>
>
> I have no Xen PV guest, so that path is untested. What I did check is that
> the __tracepoint_rdpmc relocation shows up in arch/x86/events/core.o with
> CONFIG_PARAVIRT_XXL=y, where it previously did not, and that x86_64
> defconfig still builds clean with gcc and clang.
Juergen
[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]
^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v5 05/17] x86/hyperv: Switch from __rdmsr() to native_rdmsrq()
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (3 preceding siblings ...)
2026-09-11 8:41 ` [PATCH v5 04/17] x86/msr: Move MSR trace calls one function level up Juergen Gross
@ 2026-09-11 8:41 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 06/17] x86/opcode: Add immediate form MSR instructions Juergen Gross
` (12 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:41 UTC (permalink / raw)
To: linux-kernel, x86, linux-hyperv
Cc: Juergen Gross, K. Y. Srinivasan, Haiyang Zhang, Wei Liu,
Dexuan Cui, Long Li, Thomas Gleixner, Ingo Molnar,
Borislav Petkov, Dave Hansen, H. Peter Anvin, kernel test robot
The __rdmsr() helper will be changed soon, so don't use it directly
outside of msr.h. Switch to native_rdmsrq() in HyperV related code.
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202602182222.WEBLSQRj-lkp@intel.com/
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V4:
- new patch (kernel test robot)
---
arch/x86/hyperv/hv_crash.c | 6 +++---
1 file changed, 3 insertions(+), 3 deletions(-)
diff --git a/arch/x86/hyperv/hv_crash.c b/arch/x86/hyperv/hv_crash.c
index 5ffcc23255de..28ee76e18d9b 100644
--- a/arch/x86/hyperv/hv_crash.c
+++ b/arch/x86/hyperv/hv_crash.c
@@ -217,9 +217,9 @@ static void hv_hvcrash_ctxt_save(void)
native_store_gdt(&ctxt->gdtr);
store_idt(&ctxt->idtr);
- ctxt->gsbase = __rdmsr(MSR_GS_BASE);
- ctxt->efer = __rdmsr(MSR_EFER);
- ctxt->pat = __rdmsr(MSR_IA32_CR_PAT);
+ ctxt->gsbase = native_rdmsrq(MSR_GS_BASE);
+ ctxt->efer = native_rdmsrq(MSR_EFER);
+ ctxt->pat = native_rdmsrq(MSR_IA32_CR_PAT);
}
/* Add trampoline page to the kernel pagetable for transition to kernel PT */
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 06/17] x86/opcode: Add immediate form MSR instructions
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (4 preceding siblings ...)
2026-09-11 8:41 ` [PATCH v5 05/17] x86/hyperv: Switch from __rdmsr() to native_rdmsrq() Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 07/17] x86/extable: Add support for " Juergen Gross
` (11 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Xin Li (Intel)
Add the instruction opcodes used by the immediate form WRMSRNS/RDMSR
to x86-opcode-map.
Signed-off-by: Xin Li (Intel) <xin@zytor.com>
Signed-off-by: Juergen Gross <jgross@suse.com>
Reviewed-by: H. Peter Anvin (Intel) <hpa@zytor.com>
---
V2:
- new patch, taken from the RFC v2 MSR refactor series by Xin Li
---
arch/x86/lib/x86-opcode-map.txt | 5 +++--
tools/arch/x86/lib/x86-opcode-map.txt | 5 +++--
2 files changed, 6 insertions(+), 4 deletions(-)
diff --git a/arch/x86/lib/x86-opcode-map.txt b/arch/x86/lib/x86-opcode-map.txt
index 2a4e69ecc2de..62f7c90b9a83 100644
--- a/arch/x86/lib/x86-opcode-map.txt
+++ b/arch/x86/lib/x86-opcode-map.txt
@@ -844,7 +844,7 @@ f1: MOVBE My,Gy | MOVBE Mw,Gw (66) | CRC32 Gd,Ey (F2) | CRC32 Gd,Ew (66&F2)
f2: ANDN Gy,By,Ey (v)
f3: Grp17 (1A)
f5: BZHI Gy,Ey,By (v) | PEXT Gy,By,Ey (F3),(v) | PDEP Gy,By,Ey (F2),(v) | WRUSSD/Q My,Gy (66)
-f6: ADCX Gy,Ey (66) | ADOX Gy,Ey (F3) | MULX By,Gy,rDX,Ey (F2),(v) | WRSSD/Q My,Gy
+f6: ADCX Gy,Ey (66) | ADOX Gy,Ey (F3) | MULX By,Gy,rDX,Ey (F2),(v) | WRSSD/Q My,Gy | RDMSR Rq,Gq (F2),(11B) | WRMSRNS Gq,Rq (F3),(11B)
f7: BEXTR Gy,Ey,By (v) | SHLX Gy,Ey,By (66),(v) | SARX Gy,Ey,By (F3),(v) | SHRX Gy,Ey,By (F2),(v)
f8: MOVDIR64B Gv,Mdqq (66) | ENQCMD Gv,Mdqq (F2) | ENQCMDS Gv,Mdqq (F3) | URDMSR Rq,Gq (F2),(11B) | UWRMSR Gq,Rq (F3),(11B)
f9: MOVDIRI My,Gy
@@ -1019,7 +1019,7 @@ f1: CRC32 Gy,Ey (es) | CRC32 Gy,Ey (66),(es) | INVVPID Gy,Mdq (F3),(ev)
f2: INVPCID Gy,Mdq (F3),(ev)
f4: TZCNT Gv,Ev (es) | TZCNT Gv,Ev (66),(es)
f5: LZCNT Gv,Ev (es) | LZCNT Gv,Ev (66),(es)
-f6: Grp3_1 Eb (1A),(ev)
+f6: Grp3_1 Eb (1A),(ev) | RDMSR Rq,Gq (F2),(11B),(ev) | WRMSRNS Gq,Rq (F3),(11B),(ev)
f7: Grp3_2 Ev (1A),(es)
f8: MOVDIR64B Gv,Mdqq (66),(ev) | ENQCMD Gv,Mdqq (F2),(ev) | ENQCMDS Gv,Mdqq (F3),(ev) | URDMSR Rq,Gq (F2),(11B),(ev) | UWRMSR Gq,Rq (F3),(11B),(ev)
f9: MOVDIRI My,Gy (ev)
@@ -1108,6 +1108,7 @@ EndTable
Table: VEX map 7
Referrer:
AVXcode: 7
+f6: RDMSR Rq,Id (F2),(v1),(11B) | WRMSRNS Id,Rq (F3),(v1),(11B)
f8: URDMSR Rq,Id (F2),(v1),(11B) | UWRMSR Id,Rq (F3),(v1),(11B)
EndTable
diff --git a/tools/arch/x86/lib/x86-opcode-map.txt b/tools/arch/x86/lib/x86-opcode-map.txt
index 2a4e69ecc2de..62f7c90b9a83 100644
--- a/tools/arch/x86/lib/x86-opcode-map.txt
+++ b/tools/arch/x86/lib/x86-opcode-map.txt
@@ -844,7 +844,7 @@ f1: MOVBE My,Gy | MOVBE Mw,Gw (66) | CRC32 Gd,Ey (F2) | CRC32 Gd,Ew (66&F2)
f2: ANDN Gy,By,Ey (v)
f3: Grp17 (1A)
f5: BZHI Gy,Ey,By (v) | PEXT Gy,By,Ey (F3),(v) | PDEP Gy,By,Ey (F2),(v) | WRUSSD/Q My,Gy (66)
-f6: ADCX Gy,Ey (66) | ADOX Gy,Ey (F3) | MULX By,Gy,rDX,Ey (F2),(v) | WRSSD/Q My,Gy
+f6: ADCX Gy,Ey (66) | ADOX Gy,Ey (F3) | MULX By,Gy,rDX,Ey (F2),(v) | WRSSD/Q My,Gy | RDMSR Rq,Gq (F2),(11B) | WRMSRNS Gq,Rq (F3),(11B)
f7: BEXTR Gy,Ey,By (v) | SHLX Gy,Ey,By (66),(v) | SARX Gy,Ey,By (F3),(v) | SHRX Gy,Ey,By (F2),(v)
f8: MOVDIR64B Gv,Mdqq (66) | ENQCMD Gv,Mdqq (F2) | ENQCMDS Gv,Mdqq (F3) | URDMSR Rq,Gq (F2),(11B) | UWRMSR Gq,Rq (F3),(11B)
f9: MOVDIRI My,Gy
@@ -1019,7 +1019,7 @@ f1: CRC32 Gy,Ey (es) | CRC32 Gy,Ey (66),(es) | INVVPID Gy,Mdq (F3),(ev)
f2: INVPCID Gy,Mdq (F3),(ev)
f4: TZCNT Gv,Ev (es) | TZCNT Gv,Ev (66),(es)
f5: LZCNT Gv,Ev (es) | LZCNT Gv,Ev (66),(es)
-f6: Grp3_1 Eb (1A),(ev)
+f6: Grp3_1 Eb (1A),(ev) | RDMSR Rq,Gq (F2),(11B),(ev) | WRMSRNS Gq,Rq (F3),(11B),(ev)
f7: Grp3_2 Ev (1A),(es)
f8: MOVDIR64B Gv,Mdqq (66),(ev) | ENQCMD Gv,Mdqq (F2),(ev) | ENQCMDS Gv,Mdqq (F3),(ev) | URDMSR Rq,Gq (F2),(11B),(ev) | UWRMSR Gq,Rq (F3),(11B),(ev)
f9: MOVDIRI My,Gy (ev)
@@ -1108,6 +1108,7 @@ EndTable
Table: VEX map 7
Referrer:
AVXcode: 7
+f6: RDMSR Rq,Id (F2),(v1),(11B) | WRMSRNS Id,Rq (F3),(v1),(11B)
f8: URDMSR Rq,Id (F2),(v1),(11B) | UWRMSR Id,Rq (F3),(v1),(11B)
EndTable
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 07/17] x86/extable: Add support for immediate form MSR instructions
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (5 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 06/17] x86/opcode: Add immediate form MSR instructions Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 08/17] x86/msr: Make wrmsrns() a first class citizen Juergen Gross
` (10 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Dave Hansen, Andy Lutomirski, Peter Zijlstra,
Thomas Gleixner, Ingo Molnar, Borislav Petkov, H. Peter Anvin,
Xin Li (Intel)
Signed-off-by: Xin Li (Intel) <xin@zytor.com>
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V2:
- new patch, taken from the RFC v2 MSR refactor series by Xin Li
V3:
- use instruction decoder (Peter Zijlstra)
V4:
- don't assume %rax for immediate form (Andrew Cooper)
V5:
- drop stale comment (H. Peter Anvin)
---
arch/x86/mm/extable.c | 41 +++++++++++++++++++++++++++++++++++------
1 file changed, 35 insertions(+), 6 deletions(-)
diff --git a/arch/x86/mm/extable.c b/arch/x86/mm/extable.c
index fa6eebe6a69e..9976ed01ce62 100644
--- a/arch/x86/mm/extable.c
+++ b/arch/x86/mm/extable.c
@@ -166,25 +166,54 @@ static bool ex_handler_uaccess(const struct exception_table_entry *fixup,
static bool ex_handler_msr(const struct exception_table_entry *fixup,
struct pt_regs *regs, bool wrmsr, bool safe, int reg)
{
+ unsigned long *regptr;
+ struct insn insn;
+ bool imm_insn;
+ u32 msr;
+
+ imm_insn = insn_decode_kernel(&insn, (void *)regs->ip) &&
+ insn.vex_prefix.nbytes;
+ msr = imm_insn ? insn.immediate.value : (u32)regs->cx;
+ regptr = imm_insn ? insn_get_modrm_reg_ptr(&insn, regs) : ®s->ax;
+ if (unlikely(!regptr)) {
+ pr_err("Inconsistent %sMSR access instruction data at rIP: 0x%lx (%pS)!\n",
+ wrmsr ? "WR" : "RD", regs->ip, (void *)regs->ip);
+ show_stack_regs(regs);
+ goto out;
+ }
+
if (__ONCE_LITE_IF(!safe && wrmsr)) {
- pr_warn("unchecked MSR access error: WRMSR to 0x%x (tried to write 0x%08x%08x) at rIP: 0x%lx (%pS)\n",
- (unsigned int)regs->cx, (unsigned int)regs->dx,
- (unsigned int)regs->ax, regs->ip, (void *)regs->ip);
+ u64 msr_val = *regptr;
+
+ if (!imm_insn) {
+ /*
+ * On processors that support the Intel 64 architecture, the
+ * high-order 32 bits of each of RAX and RDX are ignored.
+ */
+ msr_val &= 0xffffffff;
+ msr_val |= (u64)regs->dx << 32;
+ }
+
+ pr_warn("unchecked MSR access error: WRMSR to 0x%x (tried to write 0x%016llx) at rIP: 0x%lx (%pS)\n",
+ msr, msr_val, regs->ip, (void *)regs->ip);
show_stack_regs(regs);
}
if (__ONCE_LITE_IF(!safe && !wrmsr)) {
pr_warn("unchecked MSR access error: RDMSR from 0x%x at rIP: 0x%lx (%pS)\n",
- (unsigned int)regs->cx, regs->ip, (void *)regs->ip);
+ msr, regs->ip, (void *)regs->ip);
show_stack_regs(regs);
}
if (!wrmsr) {
/* Pretend that the read succeeded and returned 0. */
- regs->ax = 0;
- regs->dx = 0;
+ *regptr = 0;
+
+ if (!imm_insn)
+ regs->dx = 0;
}
+ out:
if (safe)
*pt_regs_nr(regs, reg) = -EIO;
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 08/17] x86/msr: Make wrmsrns() a first class citizen
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (6 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 07/17] x86/extable: Add support for " Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 09/17] x86/msr: Introduce sync_cpu_after_wrmsrns() Juergen Gross
` (9 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86, virtualization, llvm
Cc: Juergen Gross, Xin Li, H. Peter Anvin, Thomas Gleixner,
Ingo Molnar, Borislav Petkov, Dave Hansen, Ajay Kaher,
Alexey Makhalov, Broadcom internal kernel review list,
Nathan Chancellor, Nick Desaulniers, Bill Wendling, Justin Stitt
Today wrmsrns() is - apart from the potential use of the wrmsrns
instruction - equivalent to __wrmsrq(). Change that by supporting
MSR write trace entries and a safe variant.
wrmsrns() and wrmsrns_safe() will be the "normal" interfaces like
wrmsrq() and wrmsrq_safe(). They will call write_msrns[_safe]() and
conditionally create trace entries via do_trace_write_msr().
write_msrns[_safe]() are different between paravirt and non-paravirt
cases. For the paravirt case they will (for now) only use the wrmsr
paravirt functions, while for non-paravirt they call native_wrmsrns()
and native_wrmsrns_safe().
native_wrmsrns() is like wrmsrns() today, native_wrmsrns_safe() is just
the safe variant of it. The both rely on __wrmsrns(), which will use
the ALTERNATIVE*() macros for selecting WRMSR or WRMSRNS (with or
without an immediate operand specifying the MSR register) depending on
availability.
Switch the wrmsrns() call in fred_update_rsp0() to native_wrmsrns() in
order to avoid a change of functionality. The wrmsrns() call in
vmx_write_guest_host_msr() can be kept, as it has replaced a wrmsrq()
call, so eventually creating a trace entry is obviously fine here.
Originally-by: Xin Li (Intel) <xin@zytor.com>
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V2:
- new patch, partially taken from "[RFC PATCH v2 21/34] x86/msr: Utilize
the alternatives mechanism to write MSR" by Xin Li.
V4:
- don't modify __wrmsrq(), but create __wrmsrns().
---
arch/x86/include/asm/fred.h | 2 +-
arch/x86/include/asm/msr.h | 150 +++++++++++++++++++++++++++++---
arch/x86/include/asm/paravirt.h | 10 +++
3 files changed, 148 insertions(+), 14 deletions(-)
diff --git a/arch/x86/include/asm/fred.h b/arch/x86/include/asm/fred.h
index 18a2f811c358..0a6773b76968 100644
--- a/arch/x86/include/asm/fred.h
+++ b/arch/x86/include/asm/fred.h
@@ -101,7 +101,7 @@ static __always_inline void fred_update_rsp0(void)
unsigned long rsp0 = (unsigned long) task_stack_page(current) + THREAD_SIZE;
if (cpu_feature_enabled(X86_FEATURE_FRED) && (__this_cpu_read(fred_rsp0) != rsp0)) {
- wrmsrns(MSR_IA32_FRED_RSP0, rsp0);
+ native_wrmsrns(MSR_IA32_FRED_RSP0, rsp0);
__this_cpu_write(fred_rsp0, rsp0);
}
}
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index 266298b3d201..91d6f481732b 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -7,11 +7,11 @@
#ifndef __ASSEMBLER__
#include <asm/asm.h>
-#include <asm/errno.h>
#include <asm/cpumask.h>
#include <uapi/asm/msr.h>
#include <asm/shared/msr.h>
+#include <linux/errno.h>
#include <linux/types.h>
#include <linux/percpu.h>
@@ -56,6 +56,36 @@ static inline void do_trace_read_msr(u32 msr, u64 val, int failed) {}
static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
#endif
+/* The GNU Assembler (Gas) with Binutils 2.40 adds WRMSRNS support */
+#if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24000
+#define ASM_WRMSRNS "wrmsrns\n\t"
+#else
+#define ASM_WRMSRNS _ASM_BYTES(0x0f,0x01,0xc6)
+#endif
+
+/* The GNU Assembler (Gas) with Binutils 2.41 adds the .insn directive support */
+#if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24100
+#define ASM_WRMSRNS_IMM \
+ " .insn VEX.128.F3.M7.W0 0xf6 /0, %[val], %[msr]%{:u32}\n\t"
+#else
+/*
+ * Note, clang also doesn't support the .insn directive.
+ *
+ * The register operand is encoded as %rax because all uses of the immediate
+ * form MSR access instructions reference %rax as the register operand.
+ */
+#define ASM_WRMSRNS_IMM \
+ " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]"
+#endif
+
+#define PREPARE_RDX_FOR_WRMSR \
+ "mov %%rax, %%rdx\n\t" \
+ "shr $0x20, %%rdx\n\t"
+
+#define PREPARE_RCX_RDX_FOR_WRMSR \
+ "mov %[msr], %%ecx\n\t" \
+ PREPARE_RDX_FOR_WRMSR
+
/*
* __rdmsr() and __wrmsr() are the two primitives which are the bare minimum MSR
* accessors and should not have any tracing or other functionality piggybacking
@@ -83,6 +113,78 @@ static __always_inline void __wrmsrq(u32 msr, u64 val)
: : "c" (msr), "a" ((u32)val), "d" ((u32)(val >> 32)) : "memory");
}
+static __always_inline bool __wrmsrns_variable(u32 msr, u64 val, int type)
+{
+#ifdef CONFIG_X86_64
+ BUILD_BUG_ON(__builtin_constant_p(msr));
+#endif
+
+ /*
+ * WRMSR is 2 bytes. WRMSRNS is 3 bytes. Pad WRMSR with a redundant
+ * DS prefix to avoid a trailing NOP.
+ */
+ asm_inline volatile goto(
+ "1:\n"
+ ALTERNATIVE("ds wrmsr",
+ ASM_WRMSRNS,
+ X86_FEATURE_WRMSRNS)
+ _ASM_EXTABLE_TYPE(1b, %l[badmsr], %c[type])
+
+ :
+ : "c" (msr), "a" ((u32)val), "d" ((u32)(val >> 32)), [type] "i" (type)
+ : "memory"
+ : badmsr);
+
+ return false;
+
+badmsr:
+ return true;
+}
+
+#ifdef CONFIG_X86_64
+/*
+ * Non-serializing WRMSR or its immediate form, when available.
+ *
+ * Otherwise, it falls back to a serializing WRMSR.
+ */
+static __always_inline bool __wrmsrns_constant(u32 msr, u64 val, int type)
+{
+ BUILD_BUG_ON(!__builtin_constant_p(msr));
+
+ asm_inline volatile goto(
+ "1:\n"
+ ALTERNATIVE_2(PREPARE_RCX_RDX_FOR_WRMSR
+ "2: ds wrmsr",
+ PREPARE_RCX_RDX_FOR_WRMSR
+ ASM_WRMSRNS,
+ X86_FEATURE_WRMSRNS,
+ ASM_WRMSRNS_IMM,
+ X86_FEATURE_MSR_IMM)
+ _ASM_EXTABLE_TYPE(1b, %l[badmsr], %c[type]) /* For WRMSRNS immediate */
+ _ASM_EXTABLE_TYPE(2b, %l[badmsr], %c[type]) /* For WRMSR(NS) */
+
+ :
+ : [val] "a" (val), [msr] "i" (msr), [type] "i" (type)
+ : "memory", "ecx", "rdx"
+ : badmsr);
+
+ return false;
+
+badmsr:
+ return true;
+}
+#endif
+
+static __always_inline bool __wrmsrns(u32 msr, u64 val, int type)
+{
+#ifdef CONFIG_X86_64
+ if (__builtin_constant_p(msr))
+ return __wrmsrns_constant(msr, val, type);
+#endif
+
+ return __wrmsrns_variable(msr, val, type);
+}
+
static __always_inline u64 native_rdmsrq(u32 msr)
{
return __rdmsr(msr);
@@ -134,6 +236,16 @@ static inline int notrace native_write_msr_safe(u32 msr, u64 val)
return err;
}
+static __always_inline void native_wrmsrns(u32 msr, u64 val)
+{
+ __wrmsrns(msr, val, EX_TYPE_WRMSR);
+}
+
+static __always_inline int native_wrmsrns_safe(u32 msr, u64 val)
+{
+ return __wrmsrns(msr, val, EX_TYPE_WRMSR_SAFE) ? -EIO : 0;
+}
+
extern int rdmsr_safe_regs(u32 regs[8]);
extern int wrmsr_safe_regs(u32 regs[8]);
@@ -150,7 +262,6 @@ static inline u64 native_read_pmc(int counter)
#ifdef CONFIG_PARAVIRT_XXL
#include <asm/paravirt.h>
#else
-#include <linux/errno.h>
static __always_inline u64 read_msr(u32 msr)
{
return native_read_msr(msr);
@@ -171,6 +282,16 @@ static __always_inline int write_msr_safe(u32 msr, u64 val)
return native_write_msr_safe(msr, val);
}
+static __always_inline void write_msrns(u32 msr, u64 val)
+{
+ native_wrmsrns(msr, val);
+}
+
+static __always_inline int write_msrns_safe(u32 msr, u64 val)
+{
+ return native_wrmsrns_safe(msr, val);
+}
+
static __always_inline u64 rdpmc(int counter)
{
return native_read_pmc(counter);
@@ -223,19 +344,22 @@ static inline int wrmsrq_safe(u32 msr, u64 val)
return err;
}
-/* Instruction opcode for WRMSRNS supported in binutils >= 2.40 */
-#define ASM_WRMSRNS _ASM_BYTES(0x0f,0x01,0xc6)
-
-/* Non-serializing WRMSR, when available. Falls back to a serializing WRMSR. */
static __always_inline void wrmsrns(u32 msr, u64 val)
{
- /*
- * WRMSR is 2 bytes. WRMSRNS is 3 bytes. Pad WRMSR with a redundant
- * DS prefix to avoid a trailing NOP.
- */
- asm volatile("1: " ALTERNATIVE("ds wrmsr", ASM_WRMSRNS, X86_FEATURE_WRMSRNS)
- "2: " _ASM_EXTABLE_TYPE(1b, 2b, EX_TYPE_WRMSR)
- : : "c" (msr), "a" ((u32)val), "d" ((u32)(val >> 32)));
+ write_msrns(msr, val);
+
+ if (tracepoint_enabled(write_msr))
+ do_trace_write_msr(msr, val, 0);
+}
+
+static __always_inline int wrmsrns_safe(u32 msr, u64 val)
+{
+ int err = write_msrns_safe(msr, val);
+
+ if (tracepoint_enabled(write_msr))
+ do_trace_write_msr(msr, val, err);
+
+ return err;
}
struct msr __percpu *msrs_alloc(void);
diff --git a/arch/x86/include/asm/paravirt.h b/arch/x86/include/asm/paravirt.h
index a5a1fc4c88d1..b0c740316cf7 100644
--- a/arch/x86/include/asm/paravirt.h
+++ b/arch/x86/include/asm/paravirt.h
@@ -160,11 +160,21 @@ static inline void write_msr(u32 msr, u64 val)
paravirt_write_msr(msr, val);
}
+static __always_inline void write_msrns(u32 msr, u64 val)
+{
+ paravirt_write_msr(msr, val);
+}
+
static inline int write_msr_safe(u32 msr, u64 val)
{
return paravirt_write_msr_safe(msr, val);
}
+static __always_inline int write_msrns_safe(u32 msr, u64 val)
+{
+ return paravirt_write_msr_safe(msr, val);
+}
+
static __always_inline int read_msr_safe(u32 msr, u64 *p)
{
return paravirt_read_msr_safe(msr, p);
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 09/17] x86/msr: Introduce sync_cpu_after_wrmsrns()
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (7 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 08/17] x86/msr: Make wrmsrns() a first class citizen Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 10/17] x86/msr: Use the alternatives mechanism for RDMSR Juergen Gross
` (8 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin
In order to allow using wrmsrns() for multiple MSR register writes
introduce sync_cpu_after_wrmsrns() which will then do a CPU serialize
operation, if the hardware does support the WRMSRNS instruction. In
case the hardware doesn't support WRMSRNS, sync_cpu_after_wrmsrns()
will be a NOP.
Suggested-by: H. Peter Anvin <hpa@zytor.com>
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V4:
- new patch
---
arch/x86/include/asm/msr.h | 7 +++++++
1 file changed, 7 insertions(+)
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index 91d6f481732b..6ae8a49f9c07 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -14,6 +14,7 @@
#include <linux/errno.h>
#include <linux/types.h>
#include <linux/percpu.h>
+#include <asm/special_insns.h>
struct msr_info {
u32 msr_no;
@@ -362,6 +363,12 @@ static __always_inline int wrmsrns_safe(u32 msr, u64 val)
return err;
}
+static __always_inline void sync_cpu_after_wrmsrns(void)
+{
+ if (cpu_feature_enabled(X86_FEATURE_WRMSRNS))
+ serialize();
+}
+
struct msr __percpu *msrs_alloc(void);
void msrs_free(struct msr __percpu *msrs);
int msr_set_bit(u32 msr, u8 bit);
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 10/17] x86/msr: Use the alternatives mechanism for RDMSR
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (8 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 09/17] x86/msr: Introduce sync_cpu_after_wrmsrns() Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 11/17] x86/alternatives: Add ALTERNATIVE_4() Juergen Gross
` (7 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Xin Li (Intel)
When available use the immediate variant of RDMSR in __rdmsr().
For the safe/unsafe variants make __rdmsr() to be a common base
function instead of duplicating the ALTERNATIVE*() macros.
Modify native_rdmsr() and native_read_msr() to use native_rdmsrq().
The paravirt case will be handled later.
Originally-by: Xin Li (Intel) <xin@zytor.com>
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V2:
- new patch, partially taken from "[RFC PATCH v2 22/34] x86/msr: Utilize
the alternatives mechanism to read MSR" by Xin Li
---
arch/x86/include/asm/msr.h | 106 +++++++++++++++++++++++++++++--------
1 file changed, 84 insertions(+), 22 deletions(-)
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index 6ae8a49f9c07..fb1027dc7f64 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -66,6 +66,8 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
/* The GNU Assembler (Gas) with Binutils 2.41 adds the .insn directive support */
#if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24100
+#define ASM_RDMSR_IMM \
+ " .insn VEX.128.F2.M7.W0 0xf6 /0, %[msr]%{:u32}, %[val]\n\t"
#define ASM_WRMSRNS_IMM \
" .insn VEX.128.F3.M7.W0 0xf6 /0, %[val], %[msr]%{:u32}\n\t"
#else
@@ -75,10 +77,17 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
* The register operand is encoded as %rax because all uses of the immediate
* form MSR access instructions reference %rax as the register operand.
*/
+#define ASM_RDMSR_IMM \
+ " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
#define ASM_WRMSRNS_IMM \
" .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]"
#endif
+#define RDMSR_AND_SAVE_RESULT \
+ "rdmsr\n\t" \
+ "shl $0x20, %%rdx\n\t" \
+ "or %%rdx, %%rax\n\t"
+
#define PREPARE_RDX_FOR_WRMSR \
"mov %%rax, %%rdx\n\t" \
"shr $0x20, %%rdx\n\t"
@@ -94,16 +103,76 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
* think of extending them - you will be slapped with a stinking trout or a frozen
* shark will reach you, wherever you are! You've been warned.
*/
-static __always_inline u64 __rdmsr(u32 msr)
+static __always_inline bool __rdmsrq_variable(u32 msr, u64 *val, int type)
+ {
+#ifdef CONFIG_X86_64
+ BUILD_BUG_ON(__builtin_constant_p(msr));
+
+ asm_inline volatile goto(
+ "1:\n"
+ RDMSR_AND_SAVE_RESULT
+ _ASM_EXTABLE_TYPE(1b, %l[badmsr], %c[type]) /* For RDMSR */
+
+ : [val] "=a" (*val)
+ : "c" (msr), [type] "i" (type)
+ : "rdx"
+ : badmsr);
+#else
+ asm_inline volatile goto(
+ "1: rdmsr\n\t"
+ _ASM_EXTABLE_TYPE(1b, %l[badmsr], %c[type]) /* For RDMSR */
+
+ : "=A" (*val)
+ : "c" (msr), [type] "i" (type)
+ :
+ : badmsr);
+#endif
+
+ return false;
+
+badmsr:
+ *val = 0;
+
+ return true;
+}
+
+#ifdef CONFIG_X86_64
+static __always_inline bool __rdmsrq_constant(u32 msr, u64 *val, int type)
{
- EAX_EDX_DECLARE_ARGS(val, low, high);
+ BUILD_BUG_ON(!__builtin_constant_p(msr));
- asm volatile("1: rdmsr\n"
- "2:\n"
- _ASM_EXTABLE_TYPE(1b, 2b, EX_TYPE_RDMSR)
- : EAX_EDX_RET(val, low, high) : "c" (msr));
+ asm_inline volatile goto(
+ "1:\n"
+ ALTERNATIVE("mov %[msr], %%ecx\n\t"
+ "2:\n"
+ RDMSR_AND_SAVE_RESULT,
+ ASM_RDMSR_IMM,
+ X86_FEATURE_MSR_IMM)
+ _ASM_EXTABLE_TYPE(1b, %l[badmsr], %c[type]) /* For RDMSR immediate */
+ _ASM_EXTABLE_TYPE(2b, %l[badmsr], %c[type]) /* For RDMSR */
+
+ : [val] "=a" (*val)
+ : [msr] "i" (msr), [type] "i" (type)
+ : "ecx", "rdx"
+ : badmsr);
- return EAX_EDX_VAL(val, low, high);
+ return false;
+
+badmsr:
+ *val = 0;
+
+ return true;
+}
+#endif
+
+static __always_inline bool __rdmsr(u32 msr, u64 *val, int type)
+{
+#ifdef CONFIG_X86_64
+ if (__builtin_constant_p(msr))
+ return __rdmsrq_constant(msr, val, type);
+#endif
+
+ return __rdmsrq_variable(msr, val, type);
}
static __always_inline void __wrmsrq(u32 msr, u64 val)
@@ -188,7 +257,11 @@ static __always_inline bool __wrmsrns(u32 msr, u64 val, int type)
static __always_inline u64 native_rdmsrq(u32 msr)
{
- return __rdmsr(msr);
+ u64 val;
+
+ __rdmsr(msr, &val, EX_TYPE_RDMSR);
+
+ return val;
}
static __always_inline void native_wrmsrq(u32 msr, u64 val)
@@ -198,23 +271,12 @@ static __always_inline void native_wrmsrq(u32 msr, u64 val)
static inline u64 native_read_msr(u32 msr)
{
- return __rdmsr(msr);
+ return native_rdmsrq(msr);
}
-static inline int native_read_msr_safe(u32 msr, u64 *p)
+static inline int native_read_msr_safe(u32 msr, u64 *val)
{
- int err;
- EAX_EDX_DECLARE_ARGS(val, low, high);
-
- asm volatile("1: rdmsr ; xor %[err],%[err]\n"
- "2:\n\t"
- _ASM_EXTABLE_TYPE_REG(1b, 2b, EX_TYPE_RDMSR_SAFE, %[err])
- : [err] "=r" (err), EAX_EDX_RET(val, low, high)
- : "c" (msr));
-
- *p = EAX_EDX_VAL(val, low, high);
-
- return err;
+ return __rdmsr(msr, val, EX_TYPE_RDMSR_SAFE) ? -EIO : 0;
}
/* Can be uninlined because referenced by paravirt */
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 11/17] x86/alternatives: Add ALTERNATIVE_4()
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (9 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 10/17] x86/msr: Use the alternatives mechanism for RDMSR Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 12/17] x86/paravirt: Split off MSR related hooks into new header Juergen Gross
` (6 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin
For supporting WRMSR with CONFIG_PARAVIRT_XXL using direct instruction
replacement, ALTERNATIVE_4() is needed.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V3:
- new patch
---
arch/x86/include/asm/alternative.h | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/arch/x86/include/asm/alternative.h b/arch/x86/include/asm/alternative.h
index 08af86ef090a..4ec6e7653ac7 100644
--- a/arch/x86/include/asm/alternative.h
+++ b/arch/x86/include/asm/alternative.h
@@ -180,6 +180,12 @@ static __always_inline bool cpu_wants_rethunk_at(void *addr)
ALTERNATIVE(ALTERNATIVE_2(oldinstr, newinstr1, ft_flags1, newinstr2, ft_flags2), \
newinstr3, ft_flags3)
+#define ALTERNATIVE_4(oldinstr, newinstr1, ft_flags1, newinstr2, ft_flags2, \
+ newinstr3, ft_flags3, newinstr4, ft_flags4) \
+ ALTERNATIVE(ALTERNATIVE_3(oldinstr, newinstr1, ft_flags1, \
+ newinstr2, ft_flags2, newinstr3, ft_flags3),\
+ newinstr4, ft_flags4)
+
/*
* Alternative instructions for different CPU types or capabilities.
*
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 12/17] x86/paravirt: Split off MSR related hooks into new header
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (10 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 11/17] x86/alternatives: Add ALTERNATIVE_4() Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 13/17] x86/paravirt: Prepare support of MSR instruction interfaces Juergen Gross
` (5 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86, virtualization
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Ajay Kaher, Alexey Makhalov,
Broadcom internal kernel review list, Boris Ostrovsky,
Josh Poimboeuf, Peter Zijlstra, xen-devel
Move the WRMSR, RDMSR and RDPMC related parts of paravirt.h and
paravirt_types.h into a new header file paravirt-msr.h.
Switch all moved helper functions to __always_inline.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V3:
- new patch
V4:
- always use __always_inline
---
arch/x86/include/asm/msr.h | 2 +-
arch/x86/include/asm/paravirt-msr.h | 56 +++++++++++++++++++++++++++
arch/x86/include/asm/paravirt.h | 55 --------------------------
arch/x86/include/asm/paravirt_types.h | 13 -------
arch/x86/kernel/paravirt.c | 14 ++++---
arch/x86/xen/enlighten_pv.c | 11 +++---
tools/objtool/check.c | 1 +
7 files changed, 73 insertions(+), 79 deletions(-)
create mode 100644 arch/x86/include/asm/paravirt-msr.h
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index fb1027dc7f64..f7ad6d25bb65 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -323,7 +323,7 @@ static inline u64 native_read_pmc(int counter)
}
#ifdef CONFIG_PARAVIRT_XXL
-#include <asm/paravirt.h>
+#include <asm/paravirt-msr.h>
#else
static __always_inline u64 read_msr(u32 msr)
{
diff --git a/arch/x86/include/asm/paravirt-msr.h b/arch/x86/include/asm/paravirt-msr.h
new file mode 100644
index 000000000000..3e31648316a8
--- /dev/null
+++ b/arch/x86/include/asm/paravirt-msr.h
@@ -0,0 +1,56 @@
+/* SPDX-License-Identifier: GPL-2.0-only */
+#ifndef _ASM_X86_PARAVIRT_MSR_H
+#define _ASM_X86_PARAVIRT_MSR_H
+
+#include <asm/paravirt_types.h>
+
+struct pv_msr_ops {
+ /* Unsafe MSR operations. These will warn or panic on failure. */
+ u64 (*read_msr)(u32 msr);
+ void (*write_msr)(u32 msr, u64 val);
+
+ /* Safe MSR operations. Returns 0 or -EIO. */
+ int (*read_msr_safe)(u32 msr, u64 *val);
+ int (*write_msr_safe)(u32 msr, u64 val);
+
+ u64 (*read_pmc)(int counter);
+} __no_randomize_layout;
+
+extern struct pv_msr_ops pv_ops_msr;
+
+static __always_inline u64 read_msr(u32 msr)
+{
+ return PVOP_CALL1(u64, pv_ops_msr, read_msr, msr);
+}
+
+static __always_inline void write_msr(u32 msr, u64 val)
+{
+ PVOP_VCALL2(pv_ops_msr, write_msr, msr, val);
+}
+
+static __always_inline void write_msrns(u32 msr, u64 val)
+{
+ PVOP_VCALL2(pv_ops_msr, write_msr, msr, val);
+}
+
+static __always_inline int read_msr_safe(u32 msr, u64 *val)
+{
+ return PVOP_CALL2(int, pv_ops_msr, read_msr_safe, msr, val);
+}
+
+static __always_inline int write_msr_safe(u32 msr, u64 val)
+{
+ return PVOP_CALL2(int, pv_ops_msr, write_msr_safe, msr, val);
+}
+
+static __always_inline int write_msrns_safe(u32 msr, u64 val)
+{
+ return PVOP_CALL2(int, pv_ops_msr, write_msr_safe, msr, val);
+}
+
+static __always_inline u64 rdpmc(int counter)
+{
+ return PVOP_CALL1(u64, pv_ops_msr, read_pmc, counter);
+}
+
+#endif /* _ASM_X86_PARAVIRT_MSR_H */
diff --git a/arch/x86/include/asm/paravirt.h b/arch/x86/include/asm/paravirt.h
index b0c740316cf7..eb16d55f94d3 100644
--- a/arch/x86/include/asm/paravirt.h
+++ b/arch/x86/include/asm/paravirt.h
@@ -130,61 +130,6 @@ static inline void __write_cr4(unsigned long x)
PVOP_VCALL1(pv_ops, cpu.write_cr4, x);
}
-static inline u64 paravirt_read_msr(u32 msr)
-{
- return PVOP_CALL1(u64, pv_ops, cpu.read_msr, msr);
-}
-
-static inline void paravirt_write_msr(u32 msr, u64 val)
-{
- PVOP_VCALL2(pv_ops, cpu.write_msr, msr, val);
-}
-
-static inline int paravirt_read_msr_safe(u32 msr, u64 *val)
-{
- return PVOP_CALL2(int, pv_ops, cpu.read_msr_safe, msr, val);
-}
-
-static inline int paravirt_write_msr_safe(u32 msr, u64 val)
-{
- return PVOP_CALL2(int, pv_ops, cpu.write_msr_safe, msr, val);
-}
-
-static __always_inline u64 read_msr(u32 msr)
-{
- return paravirt_read_msr(msr);
-}
-
-static inline void write_msr(u32 msr, u64 val)
-{
- paravirt_write_msr(msr, val);
-}
-
-static __always_inline void write_msrns(u32 msr, u64 val)
-{
- paravirt_write_msr(msr, val);
-}
-
-static inline int write_msr_safe(u32 msr, u64 val)
-{
- return paravirt_write_msr_safe(msr, val);
-}
-
-static __always_inline int write_msrns_safe(u32 msr, u64 val)
-{
- return paravirt_write_msr_safe(msr, val);
-}
-
-static __always_inline int read_msr_safe(u32 msr, u64 *p)
-{
- return paravirt_read_msr_safe(msr, p);
-}
-
-static __always_inline u64 rdpmc(int counter)
-{
- return PVOP_CALL1(u64, pv_ops, cpu.read_pmc, counter);
-}
-
static inline void paravirt_alloc_ldt(struct desc_struct *ldt, unsigned entries)
{
PVOP_VCALL2(pv_ops, cpu.alloc_ldt, ldt, entries);
diff --git a/arch/x86/include/asm/paravirt_types.h b/arch/x86/include/asm/paravirt_types.h
index b4c4a23e77a1..2459163fa196 100644
--- a/arch/x86/include/asm/paravirt_types.h
+++ b/arch/x86/include/asm/paravirt_types.h
@@ -58,19 +58,6 @@ struct pv_cpu_ops {
void (*cpuid)(unsigned int *eax, unsigned int *ebx,
unsigned int *ecx, unsigned int *edx);
- /* Unsafe MSR operations. These will warn or panic on failure. */
- u64 (*read_msr)(u32 msr);
- void (*write_msr)(u32 msr, u64 val);
-
- /*
- * Safe MSR operations.
- * Returns 0 or -EIO.
- */
- int (*read_msr_safe)(u32 msr, u64 *val);
- int (*write_msr_safe)(u32 msr, u64 val);
-
- u64 (*read_pmc)(int counter);
-
void (*start_context_switch)(struct task_struct *prev);
void (*end_context_switch)(struct task_struct *next);
#endif
diff --git a/arch/x86/kernel/paravirt.c b/arch/x86/kernel/paravirt.c
index 00b59d774389..739dbfd8aadf 100644
--- a/arch/x86/kernel/paravirt.c
+++ b/arch/x86/kernel/paravirt.c
@@ -110,11 +110,6 @@ struct paravirt_patch_template pv_ops = {
.cpu.read_cr0 = native_read_cr0,
.cpu.write_cr0 = native_write_cr0,
.cpu.write_cr4 = native_write_cr4,
- .cpu.read_msr = native_read_msr,
- .cpu.write_msr = native_write_msr,
- .cpu.read_msr_safe = native_read_msr_safe,
- .cpu.write_msr_safe = native_write_msr_safe,
- .cpu.read_pmc = native_read_pmc,
.cpu.load_tr_desc = native_load_tr_desc,
.cpu.set_ldt = native_set_ldt,
.cpu.load_gdt = native_load_gdt,
@@ -212,6 +207,15 @@ struct paravirt_patch_template pv_ops = {
};
#ifdef CONFIG_PARAVIRT_XXL
+struct pv_msr_ops pv_ops_msr = {
+ .read_msr = native_read_msr,
+ .write_msr = native_write_msr,
+ .read_msr_safe = native_read_msr_safe,
+ .write_msr_safe = native_write_msr_safe,
+ .read_pmc = native_read_pmc,
+};
+EXPORT_SYMBOL(pv_ops_msr);
+
NOKPROBE_SYMBOL(native_load_idt);
#endif
diff --git a/arch/x86/xen/enlighten_pv.c b/arch/x86/xen/enlighten_pv.c
index 2c64b388f616..bf81e84ff261 100644
--- a/arch/x86/xen/enlighten_pv.c
+++ b/arch/x86/xen/enlighten_pv.c
@@ -1360,11 +1360,6 @@ asmlinkage __visible void __init xen_start_kernel(struct start_info *si)
pv_ops.cpu.read_cr0 = xen_read_cr0;
pv_ops.cpu.write_cr0 = xen_write_cr0;
pv_ops.cpu.write_cr4 = xen_write_cr4;
- pv_ops.cpu.read_msr = xen_read_msr;
- pv_ops.cpu.write_msr = xen_write_msr;
- pv_ops.cpu.read_msr_safe = xen_read_msr_safe;
- pv_ops.cpu.write_msr_safe = xen_write_msr_safe;
- pv_ops.cpu.read_pmc = xen_read_pmc;
pv_ops.cpu.load_tr_desc = paravirt_nop;
pv_ops.cpu.set_ldt = xen_set_ldt;
pv_ops.cpu.load_gdt = xen_load_gdt;
@@ -1385,6 +1380,12 @@ asmlinkage __visible void __init xen_start_kernel(struct start_info *si)
pv_ops.cpu.start_context_switch = xen_start_context_switch;
pv_ops.cpu.end_context_switch = xen_end_context_switch;
+ pv_ops_msr.read_msr = xen_read_msr;
+ pv_ops_msr.write_msr = xen_write_msr;
+ pv_ops_msr.read_msr_safe = xen_read_msr_safe;
+ pv_ops_msr.write_msr_safe = xen_write_msr_safe;
+ pv_ops_msr.read_pmc = xen_read_pmc;
+
xen_init_irq_ops();
/*
diff --git a/tools/objtool/check.c b/tools/objtool/check.c
index 464f6c9d9ff0..6c47d650b20e 100644
--- a/tools/objtool/check.c
+++ b/tools/objtool/check.c
@@ -529,6 +529,7 @@ static struct {
} pv_ops_tables[] = {
{ .name = "pv_ops", },
{ .name = "pv_ops_lock", },
+ { .name = "pv_ops_msr", },
{ .name = NULL, .idx_off = -1 }
};
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 13/17] x86/paravirt: Prepare support of MSR instruction interfaces
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (11 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 12/17] x86/paravirt: Split off MSR related hooks into new header Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 14/17] x86/paravirt: Switch MSR access pv_ops functions to " Juergen Gross
` (4 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86, virtualization
Cc: Juergen Gross, Ajay Kaher, Alexey Makhalov,
Broadcom internal kernel review list, Thomas Gleixner,
Ingo Molnar, Borislav Petkov, Dave Hansen, H. Peter Anvin
Make the paravirt callee-save infrastructure more generic by allowing
arbitrary register interfaces via prologue and epilogue helper macros.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V3:
- carved out from patch 5 of V1
---
arch/x86/include/asm/paravirt_types.h | 43 ++++++++++++++---------
arch/x86/include/asm/qspinlock_paravirt.h | 4 +--
2 files changed, 29 insertions(+), 18 deletions(-)
diff --git a/arch/x86/include/asm/paravirt_types.h b/arch/x86/include/asm/paravirt_types.h
index 2459163fa196..740ea819bbab 100644
--- a/arch/x86/include/asm/paravirt_types.h
+++ b/arch/x86/include/asm/paravirt_types.h
@@ -448,27 +448,38 @@ extern struct paravirt_patch_template pv_ops;
#define PV_SAVE_ALL_CALLER_REGS "pushl %ecx;"
#define PV_RESTORE_ALL_CALLER_REGS "popl %ecx;"
#else
+/* Save and restore caller-save registers, except %rax, %rcx and %rdx. */
+#define PV_SAVE_COMMON_CALLER_REGS \
+ "push %rsi;" \
+ "push %rdi;" \
+ "push %r8;" \
+ "push %r9;" \
+ "push %r10;" \
+ "push %r11;"
+
+#define PV_RESTORE_COMMON_CALLER_REGS \
+ "pop %r11;" \
+ "pop %r10;" \
+ "pop %r9;" \
+ "pop %r8;" \
+ "pop %rdi;" \
+ "pop %rsi;"
+
/* save and restore all caller-save registers, except return value */
#define PV_SAVE_ALL_CALLER_REGS \
"push %rcx;" \
"push %rdx;" \
- "push %rsi;" \
- "push %rdi;" \
- "push %r8;" \
- "push %r9;" \
- "push %r10;" \
- "push %r11;"
+ PV_SAVE_COMMON_CALLER_REGS
+
#define PV_RESTORE_ALL_CALLER_REGS \
- "pop %r11;" \
- "pop %r10;" \
- "pop %r9;" \
- "pop %r8;" \
- "pop %rdi;" \
- "pop %rsi;" \
+ PV_RESTORE_COMMON_CALLER_REGS \
"pop %rdx;" \
"pop %rcx;"
#endif
+#define PV_PROLOGUE_ALL(func) PV_SAVE_ALL_CALLER_REGS
+#define PV_EPILOGUE_ALL(func) PV_RESTORE_ALL_CALLER_REGS
+
/*
* Generate a thunk around a function which saves all caller-save
* registers except for the return value. This allows C functions to
@@ -482,7 +493,7 @@ extern struct paravirt_patch_template pv_ops;
* functions.
*/
#define PV_THUNK_NAME(func) "__raw_callee_save_" #func
-#define __PV_CALLEE_SAVE_REGS_THUNK(func, section) \
+#define __PV_CALLEE_SAVE_REGS_THUNK(func, section, helper) \
extern typeof(func) __raw_callee_save_##func; \
\
asm(".pushsection " section ", \"ax\";" \
@@ -492,16 +503,16 @@ extern struct paravirt_patch_template pv_ops;
PV_THUNK_NAME(func) ":" \
ASM_ENDBR \
FRAME_BEGIN \
- PV_SAVE_ALL_CALLER_REGS \
+ PV_PROLOGUE_##helper(func) \
"call " #func ";" \
- PV_RESTORE_ALL_CALLER_REGS \
+ PV_EPILOGUE_##helper(func) \
FRAME_END \
ASM_RET \
".size " PV_THUNK_NAME(func) ", .-" PV_THUNK_NAME(func) ";" \
".popsection")
#define PV_CALLEE_SAVE_REGS_THUNK(func) \
- __PV_CALLEE_SAVE_REGS_THUNK(func, ".text")
+ __PV_CALLEE_SAVE_REGS_THUNK(func, ".text", ALL)
/* Get a reference to a callee-save function */
#define PV_CALLEE_SAVE(func) \
diff --git a/arch/x86/include/asm/qspinlock_paravirt.h b/arch/x86/include/asm/qspinlock_paravirt.h
index 0a985784be9b..002b17f0735e 100644
--- a/arch/x86/include/asm/qspinlock_paravirt.h
+++ b/arch/x86/include/asm/qspinlock_paravirt.h
@@ -14,7 +14,7 @@ void __lockfunc __pv_queued_spin_unlock_slowpath(struct qspinlock *lock, u8 lock
*/
#ifdef CONFIG_64BIT
-__PV_CALLEE_SAVE_REGS_THUNK(__pv_queued_spin_unlock_slowpath, ".spinlock.text");
+__PV_CALLEE_SAVE_REGS_THUNK(__pv_queued_spin_unlock_slowpath, ".spinlock.text", ALL);
#define __pv_queued_spin_unlock __pv_queued_spin_unlock
/*
@@ -61,7 +61,7 @@ DEFINE_ASM_FUNC(__raw_callee_save___pv_queued_spin_unlock,
#else /* CONFIG_64BIT */
extern void __lockfunc __pv_queued_spin_unlock(struct qspinlock *lock);
-__PV_CALLEE_SAVE_REGS_THUNK(__pv_queued_spin_unlock, ".spinlock.text");
+__PV_CALLEE_SAVE_REGS_THUNK(__pv_queued_spin_unlock, ".spinlock.text", ALL);
#endif /* CONFIG_64BIT */
#endif
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 14/17] x86/paravirt: Switch MSR access pv_ops functions to instruction interfaces
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (12 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 13/17] x86/paravirt: Prepare support of MSR instruction interfaces Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 15/17] x86/msr: Reduce number of low level MSR access helpers Juergen Gross
` (3 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86, virtualization
Cc: Juergen Gross, Ajay Kaher, Alexey Makhalov,
Broadcom internal kernel review list, Thomas Gleixner,
Ingo Molnar, Borislav Petkov, Dave Hansen, H. Peter Anvin,
Boris Ostrovsky, xen-devel
In order to prepare for inlining RDMSR/WRMSR instructions via
alternatives directly when running not in a Xen PV guest, switch the
interfaces of the MSR related pvops callbacks to ones similar of the
related instructions.
In order to prepare for supporting the immediate variants of RDMSR/WRMSR
use a 64-bit interface instead of the 32-bit one of RDMSR/WRMSR.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V3:
- former patch 5 of V1 has been split
- use 64-bit interface (Xin Li)
---
arch/x86/include/asm/paravirt-msr.h | 64 ++++++++++++++++++++++++-----
arch/x86/kernel/paravirt.c | 36 ++++++++++++++--
arch/x86/xen/enlighten_pv.c | 45 +++++++++++++++-----
3 files changed, 120 insertions(+), 25 deletions(-)
diff --git a/arch/x86/include/asm/paravirt-msr.h b/arch/x86/include/asm/paravirt-msr.h
index 3e31648316a8..4b71a1cd780c 100644
--- a/arch/x86/include/asm/paravirt-msr.h
+++ b/arch/x86/include/asm/paravirt-msr.h
@@ -6,46 +6,90 @@
struct pv_msr_ops {
/* Unsafe MSR operations. These will warn or panic on failure. */
- u64 (*read_msr)(u32 msr);
- void (*write_msr)(u32 msr, u64 val);
+ struct paravirt_callee_save read_msr;
+ struct paravirt_callee_save write_msr;
/* Safe MSR operations. Returns 0 or -EIO. */
- int (*read_msr_safe)(u32 msr, u64 *val);
- int (*write_msr_safe)(u32 msr, u64 val);
+ struct paravirt_callee_save read_msr_safe;
+ struct paravirt_callee_save write_msr_safe;
u64 (*read_pmc)(int counter);
} __no_randomize_layout;
extern struct pv_msr_ops pv_ops_msr;
+#define PV_PROLOGUE_MSR(func) \
+ PV_SAVE_COMMON_CALLER_REGS \
+ PV_PROLOGUE_MSR_##func
+
+#define PV_EPILOGUE_MSR(func) PV_RESTORE_COMMON_CALLER_REGS
+
+#define PV_CALLEE_SAVE_REGS_MSR_THUNK(func) \
+ __PV_CALLEE_SAVE_REGS_THUNK(func, ".text", MSR)
+
static __always_inline u64 read_msr(u32 msr)
{
- return PVOP_CALL1(u64, pv_ops_msr, read_msr, msr);
+ u64 val;
+
+ asm volatile(PARAVIRT_CALL
+ : "=a" (val), ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, read_msr), "c" (msr)
+ : "rdx");
+
+ return val;
}
static __always_inline void write_msr(u32 msr, u64 val)
{
- PVOP_VCALL2(pv_ops_msr, write_msr, msr, val);
+ asm volatile(PARAVIRT_CALL
+ : ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, write_msr), "c" (msr), "a" (val)
+ : "memory", "rdx");
}
static __always_inline void write_msrns(u32 msr, u64 val)
{
- PVOP_VCALL2(pv_ops_msr, write_msr, msr, val);
+ asm volatile(PARAVIRT_CALL
+ : ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, write_msr), "c" (msr), "a" (val)
+ : "memory", "rdx");
}
static __always_inline int read_msr_safe(u32 msr, u64 *val)
{
- return PVOP_CALL2(int, pv_ops_msr, read_msr_safe, msr, val);
+ int err;
+
+ asm volatile(PARAVIRT_CALL
+ : [err] "=d" (err), "=a" (*val), ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, read_msr_safe), "c" (msr));
+
+ return err ? -EIO : 0;
}
static __always_inline int write_msr_safe(u32 msr, u64 val)
{
- return PVOP_CALL2(int, pv_ops_msr, write_msr_safe, msr, val);
+ int err;
+
+ asm volatile(PARAVIRT_CALL
+ : [err] "=a" (err), ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, write_msr_safe),
+ "c" (msr), "a" (val)
+ : "memory", "rdx");
+
+ return err ? -EIO : 0;
}
static __always_inline int write_msrns_safe(u32 msr, u64 val)
{
- return PVOP_CALL2(int, pv_ops_msr, write_msr_safe, msr, val);
+ int err;
+
+ asm volatile(PARAVIRT_CALL
+ : [err] "=a" (err), ASM_CALL_CONSTRAINT
+ : paravirt_ptr(pv_ops_msr, write_msr_safe),
+ "c" (msr), "a" (val)
+ : "memory", "rdx");
+
+ return err ? -EIO : 0;
}
static __always_inline u64 rdpmc(int counter)
diff --git a/arch/x86/kernel/paravirt.c b/arch/x86/kernel/paravirt.c
index 739dbfd8aadf..66c0d6b5423c 100644
--- a/arch/x86/kernel/paravirt.c
+++ b/arch/x86/kernel/paravirt.c
@@ -50,12 +50,40 @@ unsigned long pv_native_save_fl(void);
void pv_native_irq_disable(void);
void pv_native_irq_enable(void);
unsigned long pv_native_read_cr2(void);
+void pv_native_rdmsr(void);
+void pv_native_wrmsr(void);
+void pv_native_rdmsr_safe(void);
+void pv_native_wrmsr_safe(void);
DEFINE_ASM_FUNC(_paravirt_ident_64, "mov %rdi, %rax", .text);
DEFINE_ASM_FUNC(pv_native_save_fl, "pushf; pop %rax", .noinstr.text);
DEFINE_ASM_FUNC(pv_native_irq_disable, "cli", .noinstr.text);
DEFINE_ASM_FUNC(pv_native_irq_enable, "sti", .noinstr.text);
DEFINE_ASM_FUNC(pv_native_read_cr2, "mov %cr2, %rax", .noinstr.text);
+DEFINE_ASM_FUNC(pv_native_rdmsr,
+ "1: rdmsr\n"
+ "shl $32, %rdx; or %rdx, %rax\n"
+ "2:\n"
+ _ASM_EXTABLE_TYPE(1b, 2b, EX_TYPE_RDMSR), .noinstr.text);
+DEFINE_ASM_FUNC(pv_native_wrmsr,
+ "mov %rax, %rdx; shr $32, %rdx\n"
+ "1: wrmsr\n"
+ "2:\n"
+ _ASM_EXTABLE_TYPE(1b, 2b, EX_TYPE_WRMSR), .noinstr.text);
+DEFINE_ASM_FUNC(pv_native_rdmsr_safe,
+ "1: rdmsr\n"
+ "shl $32, %rdx; or %rdx, %rax\n"
+ "xor %edx, %edx\n"
+ "2:\n"
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, EX_TYPE_RDMSR_SAFE, %%edx),
+ .noinstr.text);
+DEFINE_ASM_FUNC(pv_native_wrmsr_safe,
+ "mov %rax, %rdx; shr $32, %rdx\n"
+ "1: wrmsr\n"
+ "xor %eax, %eax\n"
+ "2:\n"
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, EX_TYPE_WRMSR_SAFE, %%eax),
+ .noinstr.text);
#endif
static noinstr void pv_native_safe_halt(void)
@@ -208,10 +236,10 @@ struct paravirt_patch_template pv_ops = {
#ifdef CONFIG_PARAVIRT_XXL
struct pv_msr_ops pv_ops_msr = {
- .read_msr = native_read_msr,
- .write_msr = native_write_msr,
- .read_msr_safe = native_read_msr_safe,
- .write_msr_safe = native_write_msr_safe,
+ .read_msr = __PV_IS_CALLEE_SAVE(pv_native_rdmsr),
+ .write_msr = __PV_IS_CALLEE_SAVE(pv_native_wrmsr),
+ .read_msr_safe = __PV_IS_CALLEE_SAVE(pv_native_rdmsr_safe),
+ .write_msr_safe = __PV_IS_CALLEE_SAVE(pv_native_wrmsr_safe),
.read_pmc = native_read_pmc,
};
EXPORT_SYMBOL(pv_ops_msr);
diff --git a/arch/x86/xen/enlighten_pv.c b/arch/x86/xen/enlighten_pv.c
index bf81e84ff261..505a85c3869e 100644
--- a/arch/x86/xen/enlighten_pv.c
+++ b/arch/x86/xen/enlighten_pv.c
@@ -1148,15 +1148,32 @@ static void xen_do_write_msr(u32 msr, u64 val, int *err)
}
}
-static int xen_read_msr_safe(u32 msr, u64 *val)
+/*
+ * Prototypes for functions called via PV_CALLEE_SAVE_REGS_THUNK() in order
+ * to avoid warnings with "-Wmissing-prototypes".
+ */
+struct xen_rdmsr_safe_ret {
+ u64 val;
+ int err;
+};
+struct xen_rdmsr_safe_ret xen_read_msr_safe(u32 msr);
+int xen_write_msr_safe(u32 msr, u64 val);
+u64 xen_read_msr(u32 msr);
+void xen_write_msr(u32 msr, u64 val);
+#define PV_PROLOGUE_RDMSR "mov %ecx, %edi;"
+#define PV_PROLOGUE_WRMSR "mov %ecx, %edi; mov %rax, %rsi;"
+
+__visible struct xen_rdmsr_safe_ret xen_read_msr_safe(u32 msr)
{
- int err = 0;
+ struct xen_rdmsr_safe_ret ret = { 0, 0 };
- *val = xen_do_read_msr(msr, &err);
- return err;
+ ret.val = xen_do_read_msr(msr, &ret.err);
+ return ret;
}
+#define PV_PROLOGUE_MSR_xen_read_msr_safe PV_PROLOGUE_RDMSR
+PV_CALLEE_SAVE_REGS_MSR_THUNK(xen_read_msr_safe);
-static int xen_write_msr_safe(u32 msr, u64 val)
+__visible int xen_write_msr_safe(u32 msr, u64 val)
{
int err = 0;
@@ -1164,20 +1181,26 @@ static int xen_write_msr_safe(u32 msr, u64 val)
return err;
}
+#define PV_PROLOGUE_MSR_xen_write_msr_safe PV_PROLOGUE_WRMSR
+PV_CALLEE_SAVE_REGS_MSR_THUNK(xen_write_msr_safe);
-static u64 xen_read_msr(u32 msr)
+__visible u64 xen_read_msr(u32 msr)
{
int err = 0;
return xen_do_read_msr(msr, xen_msr_safe ? &err : NULL);
}
+#define PV_PROLOGUE_MSR_xen_read_msr PV_PROLOGUE_RDMSR
+PV_CALLEE_SAVE_REGS_MSR_THUNK(xen_read_msr);
-static void xen_write_msr(u32 msr, u64 val)
+__visible void xen_write_msr(u32 msr, u64 val)
{
int err;
xen_do_write_msr(msr, val, xen_msr_safe ? &err : NULL);
}
+#define PV_PROLOGUE_MSR_xen_write_msr PV_PROLOGUE_WRMSR
+PV_CALLEE_SAVE_REGS_MSR_THUNK(xen_write_msr);
/* This is called once we have the cpu_possible_mask */
void __init xen_setup_vcpu_info_placement(void)
@@ -1380,10 +1403,10 @@ asmlinkage __visible void __init xen_start_kernel(struct start_info *si)
pv_ops.cpu.start_context_switch = xen_start_context_switch;
pv_ops.cpu.end_context_switch = xen_end_context_switch;
- pv_ops_msr.read_msr = xen_read_msr;
- pv_ops_msr.write_msr = xen_write_msr;
- pv_ops_msr.read_msr_safe = xen_read_msr_safe;
- pv_ops_msr.write_msr_safe = xen_write_msr_safe;
+ pv_ops_msr.read_msr = PV_CALLEE_SAVE(xen_read_msr);
+ pv_ops_msr.write_msr = PV_CALLEE_SAVE(xen_write_msr);
+ pv_ops_msr.read_msr_safe = PV_CALLEE_SAVE(xen_read_msr_safe);
+ pv_ops_msr.write_msr_safe = PV_CALLEE_SAVE(xen_write_msr_safe);
pv_ops_msr.read_pmc = xen_read_pmc;
xen_init_irq_ops();
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 15/17] x86/msr: Reduce number of low level MSR access helpers
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (13 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 14/17] x86/paravirt: Switch MSR access pv_ops functions to " Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 16/17] x86/paravirt: Use alternatives for MSR access with paravirt Juergen Gross
` (2 subsequent siblings)
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin, Boris Ostrovsky, xen-devel
Some MSR access helpers are redundant now, so remove the no longer
needed ones.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
arch/x86/include/asm/msr.h | 15 ++-------------
arch/x86/xen/enlighten_pv.c | 4 ++--
2 files changed, 4 insertions(+), 15 deletions(-)
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index f7ad6d25bb65..dbc24550a504 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -269,22 +269,11 @@ static __always_inline void native_wrmsrq(u32 msr, u64 val)
__wrmsrq(msr, val);
}
-static inline u64 native_read_msr(u32 msr)
-{
- return native_rdmsrq(msr);
-}
-
static inline int native_read_msr_safe(u32 msr, u64 *val)
{
return __rdmsr(msr, val, EX_TYPE_RDMSR_SAFE) ? -EIO : 0;
}
-/* Can be uninlined because referenced by paravirt */
-static inline void notrace native_write_msr(u32 msr, u64 val)
-{
- native_wrmsrq(msr, val);
-}
-
/* Can be uninlined because referenced by paravirt */
static inline int notrace native_write_msr_safe(u32 msr, u64 val)
{
@@ -327,7 +316,7 @@ static inline u64 native_read_pmc(int counter)
#else
static __always_inline u64 read_msr(u32 msr)
{
- return native_read_msr(msr);
+ return native_rdmsrq(msr);
}
static __always_inline int read_msr_safe(u32 msr, u64 *p)
@@ -337,7 +326,7 @@ static __always_inline int read_msr_safe(u32 msr, u64 *p)
static __always_inline void write_msr(u32 msr, u64 val)
{
- native_write_msr(msr, val);
+ native_wrmsrq(msr, val);
}
static __always_inline int write_msr_safe(u32 msr, u64 val)
diff --git a/arch/x86/xen/enlighten_pv.c b/arch/x86/xen/enlighten_pv.c
index 505a85c3869e..bc572ca49a2c 100644
--- a/arch/x86/xen/enlighten_pv.c
+++ b/arch/x86/xen/enlighten_pv.c
@@ -1085,7 +1085,7 @@ static u64 xen_do_read_msr(u32 msr, int *err)
if (err)
*err = native_read_msr_safe(msr, &val);
else
- val = native_read_msr(msr);
+ val = native_rdmsrq(msr);
switch (msr) {
case MSR_IA32_APICBASE:
@@ -1144,7 +1144,7 @@ static void xen_do_write_msr(u32 msr, u64 val, int *err)
if (err)
*err = native_write_msr_safe(msr, val);
else
- native_write_msr(msr, val);
+ native_wrmsrq(msr, val);
}
}
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 16/17] x86/paravirt: Use alternatives for MSR access with paravirt
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (14 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 15/17] x86/msr: Reduce number of low level MSR access helpers Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-11 8:42 ` [PATCH v5 17/17] x86/msr: Make all MSR access functions __always_inline Juergen Gross
2026-09-23 19:36 ` [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Shreshth Srivastava
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86, virtualization, llvm
Cc: Juergen Gross, Ajay Kaher, Alexey Makhalov,
Broadcom internal kernel review list, Thomas Gleixner,
Ingo Molnar, Borislav Petkov, Dave Hansen, H. Peter Anvin,
Nathan Chancellor, Nick Desaulniers, Bill Wendling, Justin Stitt
When not running as Xen PV guest, patch in the optimal MSR instructions
via alternative and use direct calls otherwise.
This will especially have positive effects for performance when not
running as a Xen PV guest with paravirtualization enabled, as there
will be no call overhead for MSR access functions any longer.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V3:
- new patch
V4:
- fix build error with clang (kernel test robot)
---
arch/x86/include/asm/paravirt-msr.h | 136 ++++++++++++++++++++------
arch/x86/include/asm/paravirt_types.h | 1 +
2 files changed, 109 insertions(+), 28 deletions(-)
diff --git a/arch/x86/include/asm/paravirt-msr.h b/arch/x86/include/asm/paravirt-msr.h
index 4b71a1cd780c..ba3ee64446db 100644
--- a/arch/x86/include/asm/paravirt-msr.h
+++ b/arch/x86/include/asm/paravirt-msr.h
@@ -27,67 +27,147 @@ extern struct pv_msr_ops pv_ops_msr;
#define PV_CALLEE_SAVE_REGS_MSR_THUNK(func) \
__PV_CALLEE_SAVE_REGS_THUNK(func, ".text", MSR)
+#define ASM_CLRERR "xor %[err],%[err]\n"
+
+#define PV_RDMSR_VAR(__msr, __val, __type, __func, __err) \
+ asm volatile( \
+ "1:\n" \
+ ALTERNATIVE_2(PARAVIRT_CALL, \
+ RDMSR_AND_SAVE_RESULT ASM_CLRERR, X86_FEATURE_ALWAYS, \
+ ALT_CALL_INSTR, ALT_XEN_CALL) \
+ "2:\n" \
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, __type, %[err]) \
+ : [err] "=d" (__err), [val] "=a" (__val), \
+ ASM_CALL_CONSTRAINT \
+ : paravirt_ptr(pv_ops_msr, __func), "c" (__msr) \
+ : "cc")
+
+#define PV_RDMSR_CONST(__msr, __val, __type, __func, __err) \
+ asm volatile( \
+ "1:\n" \
+ ALTERNATIVE_3(PARAVIRT_CALL, \
+ RDMSR_AND_SAVE_RESULT ASM_CLRERR, X86_FEATURE_ALWAYS, \
+ ASM_RDMSR_IMM ASM_CLRERR, X86_FEATURE_MSR_IMM, \
+ ALT_CALL_INSTR, ALT_XEN_CALL) \
+ "2:\n" \
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, __type, %[err]) \
+ : [err] "=d" (__err), [val] "=a" (__val), \
+ ASM_CALL_CONSTRAINT \
+ : paravirt_ptr(pv_ops_msr, __func), \
+ "c" (__msr), [msr] "i" (__msr) \
+ : "cc")
+
+#define PV_WRMSR(__msr, __val, __type, __func, __err) \
+({ \
+ unsigned long rdx = rdx; \
+ asm volatile( \
+ "1:\n" \
+ ALTERNATIVE_2(PARAVIRT_CALL, \
+ "wrmsr;" ASM_CLRERR, X86_FEATURE_ALWAYS, \
+ ALT_CALL_INSTR, ALT_XEN_CALL) \
+ "2:\n" \
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, __type, %[err]) \
+ : [err] "=a" (__err), "=d" (rdx), ASM_CALL_CONSTRAINT \
+ : paravirt_ptr(pv_ops_msr, __func), \
+ "0" (__val), "1" ((__val) >> 32), "c" (__msr) \
+ : "memory", "cc"); \
+})
+
+#define PV_WRMSRNS_VAR(__msr, __val, __type, __func, __err) \
+({ \
+ unsigned long rdx = rdx; \
+ asm volatile( \
+ "1:\n" \
+ ALTERNATIVE_3(PARAVIRT_CALL, \
+ "wrmsr;" ASM_CLRERR, X86_FEATURE_ALWAYS, \
+ ASM_WRMSRNS ASM_CLRERR, X86_FEATURE_WRMSRNS, \
+ ALT_CALL_INSTR, ALT_XEN_CALL) \
+ "2:\n" \
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, __type, %[err]) \
+ : [err] "=a" (__err), "=d" (rdx), ASM_CALL_CONSTRAINT \
+ : paravirt_ptr(pv_ops_msr, __func), \
+ "0" (__val), "1" ((__val) >> 32), "c" (__msr) \
+ : "memory", "cc"); \
+})
+
+#define PV_WRMSRNS_CONST(__msr, __val, __type, __func, __err) \
+({ \
+ unsigned long rdx = rdx; \
+ asm volatile( \
+ "1:\n" \
+ ALTERNATIVE_4(PARAVIRT_CALL, \
+ "wrmsr;" ASM_CLRERR, X86_FEATURE_ALWAYS, \
+ ASM_WRMSRNS ASM_CLRERR, X86_FEATURE_WRMSRNS, \
+ ASM_WRMSRNS_IMM ASM_CLRERR, X86_FEATURE_MSR_IMM,\
+ ALT_CALL_INSTR, ALT_XEN_CALL) \
+ "2:\n" \
+ _ASM_EXTABLE_TYPE_REG(1b, 2b, __type, %[err]) \
+ : [err] "=a" (__err), "=d" (rdx), ASM_CALL_CONSTRAINT \
+ : paravirt_ptr(pv_ops_msr, __func), \
+ [val] "0" (__val), "1" ((__val) >> 32), \
+ "c" (__msr), [msr] "i" (__msr) \
+ : "memory", "cc"); \
+})
+
static __always_inline u64 read_msr(u32 msr)
{
u64 val;
+ u64 err;
- asm volatile(PARAVIRT_CALL
- : "=a" (val), ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, read_msr), "c" (msr)
- : "rdx");
+ if (__builtin_constant_p(msr))
+ PV_RDMSR_CONST(msr, val, EX_TYPE_RDMSR, read_msr, err);
+ else
+ PV_RDMSR_VAR(msr, val, EX_TYPE_RDMSR, read_msr, err);
return val;
}
static __always_inline void write_msr(u32 msr, u64 val)
{
- asm volatile(PARAVIRT_CALL
- : ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, write_msr), "c" (msr), "a" (val)
- : "memory", "rdx");
+ u64 err;
+
+ PV_WRMSR(msr, val, EX_TYPE_WRMSR, write_msr, err);
}
static __always_inline void write_msrns(u32 msr, u64 val)
{
- asm volatile(PARAVIRT_CALL
- : ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, write_msr), "c" (msr), "a" (val)
- : "memory", "rdx");
+ u64 err;
+
+ if (__builtin_constant_p(msr))
+ PV_WRMSRNS_CONST(msr, val, EX_TYPE_WRMSR, write_msr, err);
+ else
+ PV_WRMSRNS_VAR(msr, val, EX_TYPE_WRMSR, write_msr, err);
}
static __always_inline int read_msr_safe(u32 msr, u64 *val)
{
- int err;
+ u64 err;
- asm volatile(PARAVIRT_CALL
- : [err] "=d" (err), "=a" (*val), ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, read_msr_safe), "c" (msr));
+ if (__builtin_constant_p(msr))
+ PV_RDMSR_CONST(msr, *val, EX_TYPE_RDMSR_SAFE, read_msr_safe, err);
+ else
+ PV_RDMSR_VAR(msr, *val, EX_TYPE_RDMSR_SAFE, read_msr_safe, err);
return err ? -EIO : 0;
}
static __always_inline int write_msr_safe(u32 msr, u64 val)
{
- int err;
+ u64 err;
- asm volatile(PARAVIRT_CALL
- : [err] "=a" (err), ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, write_msr_safe),
- "c" (msr), "a" (val)
- : "memory", "rdx");
+ PV_WRMSR(msr, val, EX_TYPE_WRMSR_SAFE, write_msr_safe, err);
return err ? -EIO : 0;
}
static __always_inline int write_msrns_safe(u32 msr, u64 val)
{
- int err;
+ u64 err;
- asm volatile(PARAVIRT_CALL
- : [err] "=a" (err), ASM_CALL_CONSTRAINT
- : paravirt_ptr(pv_ops_msr, write_msr_safe),
- "c" (msr), "a" (val)
- : "memory", "rdx");
+ if (__builtin_constant_p(msr))
+ PV_WRMSRNS_CONST(msr, val, EX_TYPE_WRMSR_SAFE, write_msr_safe, err);
+ else
+ PV_WRMSRNS_VAR(msr, val, EX_TYPE_WRMSR_SAFE, write_msr_safe, err);
return err ? -EIO : 0;
}
diff --git a/arch/x86/include/asm/paravirt_types.h b/arch/x86/include/asm/paravirt_types.h
index 740ea819bbab..54f7c3d8fadf 100644
--- a/arch/x86/include/asm/paravirt_types.h
+++ b/arch/x86/include/asm/paravirt_types.h
@@ -442,6 +442,7 @@ extern struct paravirt_patch_template pv_ops;
#endif /* __ASSEMBLER__ */
#define ALT_NOT_XEN ALT_NOT(X86_FEATURE_XENPV)
+#define ALT_XEN_CALL ALT_DIRECT_CALL(X86_FEATURE_XENPV)
#ifdef CONFIG_X86_32
/* save and restore all caller-save registers, except return value */
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 17/17] x86/msr: Make all MSR access functions __always_inline
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (15 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 16/17] x86/paravirt: Use alternatives for MSR access with paravirt Juergen Gross
@ 2026-09-11 8:42 ` Juergen Gross
2026-09-23 19:36 ` [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Shreshth Srivastava
17 siblings, 0 replies; 23+ messages in thread
From: Juergen Gross @ 2026-09-11 8:42 UTC (permalink / raw)
To: linux-kernel, x86
Cc: Juergen Gross, Thomas Gleixner, Ingo Molnar, Borislav Petkov,
Dave Hansen, H. Peter Anvin
There are a few MSR access functions left which are not yet marked as
__always_inline. Do the conversion.
Remove a leftover comment no longer being true related to this.
Signed-off-by: Juergen Gross <jgross@suse.com>
---
V4:
- new patch
---
arch/x86/include/asm/msr.h | 13 ++++++-------
1 file changed, 6 insertions(+), 7 deletions(-)
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index dbc24550a504..eba325ecfe4c 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -269,13 +269,12 @@ static __always_inline void native_wrmsrq(u32 msr, u64 val)
__wrmsrq(msr, val);
}
-static inline int native_read_msr_safe(u32 msr, u64 *val)
+static __always_inline int native_read_msr_safe(u32 msr, u64 *val)
{
return __rdmsr(msr, val, EX_TYPE_RDMSR_SAFE) ? -EIO : 0;
}
-/* Can be uninlined because referenced by paravirt */
-static inline int notrace native_write_msr_safe(u32 msr, u64 val)
+static __always_inline int notrace native_write_msr_safe(u32 msr, u64 val)
{
int err;
@@ -301,7 +300,7 @@ static __always_inline int native_wrmsrns_safe(u32 msr, u64 val)
extern int rdmsr_safe_regs(u32 regs[8]);
extern int wrmsr_safe_regs(u32 regs[8]);
-static inline u64 native_read_pmc(int counter)
+static __always_inline u64 native_read_pmc(int counter)
{
EAX_EDX_DECLARE_ARGS(val, low, high);
@@ -363,7 +362,7 @@ static __always_inline u64 rdmsrq(u32 msr)
}
/* rdmsr with exception handling */
-static inline int rdmsrq_safe(u32 msr, u64 *p)
+static __always_inline int rdmsrq_safe(u32 msr, u64 *p)
{
int err;
@@ -375,7 +374,7 @@ static inline int rdmsrq_safe(u32 msr, u64 *p)
return err;
}
-static inline void wrmsrq(u32 msr, u64 val)
+static __always_inline void wrmsrq(u32 msr, u64 val)
{
write_msr(msr, val);
@@ -384,7 +383,7 @@ static inline void wrmsrq(u32 msr, u64 val)
}
/* wrmsr with exception handling */
-static inline int wrmsrq_safe(u32 msr, u64 val)
+static __always_inline int wrmsrq_safe(u32 msr, u64 val)
{
int err;
--
2.55.0
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions
2026-09-11 8:41 [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Juergen Gross
` (16 preceding siblings ...)
2026-09-11 8:42 ` [PATCH v5 17/17] x86/msr: Make all MSR access functions __always_inline Juergen Gross
@ 2026-09-23 19:36 ` Shreshth Srivastava
2026-09-23 20:15 ` Nick Desaulniers
2026-09-25 10:25 ` Jürgen Groß
17 siblings, 2 replies; 23+ messages in thread
From: Shreshth Srivastava @ 2026-09-23 19:36 UTC (permalink / raw)
To: Juergen Gross, linux-kernel, x86, linux-coco, kvm, linux-hyperv,
virtualization, llvm
Cc: tglx, mingo, bp, dave.hansen, hpa, xin, nathan, ndesaulniers,
jpoimboe, peterz, boris.ostrovsky, xen-devel
On 11.09.26 10:41, Juergen Gross wrote:
> When building a kernel with CONFIG_PARAVIRT_XXL the paravirt
> infrastructure will always use functions for reading or writing MSRs,
> even when running on bare metal.
Hi Juergen,
This doesn't build with CONFIG_PARAVIRT_XXL=y. 16/17 and 17/17 are where
it breaks, but the cause is the .byte fallbacks: ASM_WRMSRNS_IMM from
08/17 and ASM_RDMSR_IMM from 10/17 don't end in a separator, unlike the
.insn variants above them.
#define ASM_RDMSR_IMM \
" .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
That worked while they were only ever the last argument of an
ALTERNATIVE(), which appends its own newline. 16/17 concatenates them
with ASM_CLRERR:
ASM_RDMSR_IMM ASM_CLRERR, X86_FEATURE_MSR_IMM, \
so the .long operand runs into the xor. From
make arch/x86/kernel/cpu/common.s:
.byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long 266xor %rdx,%rdx
paravirt-msr.h:165: Error: junk at end of line, first unrecognized character is `x'
clang reports "error: unexpected token" in the same place. 71 objects
fail, the same 71 either way, no vmlinux.
msr.h chooses between the .insn form and the .byte fallback with:
#if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24100
Two kinds of toolchain end up on the .byte side of that test:
- GNU as older than 2.41. RHEL 9 and CentOS Stream 9 ship 2.35, and
Documentation/process/changes.rst sets the minimum at 2.30, so this
is a supported configuration rather than an old outlier.
- clang, any version. CONFIG_AS_IS_GNU is never set for clang, so the
&& short-circuits and the version comparison is never reached. Your
08/17 comment already notes that clang has no .insn support.
gcc with binutils 2.41 or newer takes the .insn path, where both macros
do end in a separator, and is unaffected.
Reproduced on v7.3-rc2 with your v3 00/13, v2 0/5 and v5 00/17 applied in
that order, x86_64 defconfig plus HYPERVISOR_GUEST, PARAVIRT, XEN and
XEN_PV. gcc 11.5.0 with GNU as 2.35.2, and clang 21.1.7.
Terminating both fallbacks fixes it, and both toolchains then build
vmlinux with no errors or warnings:
diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
index eba325ecfe4c..529c13553c63 100644
--- a/arch/x86/include/asm/msr.h
+++ b/arch/x86/include/asm/msr.h
@@ -78,9 +78,9 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
* form MSR access instructions reference %rax as the register operand.
*/
#define ASM_RDMSR_IMM \
- " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
+ " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]\n\t"
#define ASM_WRMSRNS_IMM \
- " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]"
+ " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]\n\t"
#endif
#define RDMSR_AND_SAVE_RESULT \
ASM_WRMSRNS needs no change, _ASM_BYTES() already emits a semicolon.
The WRMSRNS line belongs in 08/17 and the RDMSR line in 10/17.
Thanks,
Shreshth
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions
2026-09-23 19:36 ` [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Shreshth Srivastava
@ 2026-09-23 20:15 ` Nick Desaulniers
2026-09-25 10:25 ` Jürgen Groß
1 sibling, 0 replies; 23+ messages in thread
From: Nick Desaulniers @ 2026-09-23 20:15 UTC (permalink / raw)
To: Shreshth Srivastava, Juergen Gross
Cc: linux-kernel, x86, linux-coco, kvm, linux-hyperv, virtualization,
llvm, tglx, mingo, bp, dave.hansen, hpa, xin, nathan, jpoimboe,
peterz, boris.ostrovsky, xen-devel
On Wed, Sep 23, 2026 at 12:36 PM Shreshth Srivastava
<shreshth.srivastava@intel.com> wrote:
>
> On 11.09.26 10:41, Juergen Gross wrote:
> > When building a kernel with CONFIG_PARAVIRT_XXL the paravirt
> > infrastructure will always use functions for reading or writing MSRs,
> > even when running on bare metal.
>
> Hi Juergen,
>
> This doesn't build with CONFIG_PARAVIRT_XXL=y. 16/17 and 17/17 are where
> it breaks, but the cause is the .byte fallbacks: ASM_WRMSRNS_IMM from
> 08/17 and ASM_RDMSR_IMM from 10/17 don't end in a separator, unlike the
> .insn variants above them.
>
> #define ASM_RDMSR_IMM \
> " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
>
> That worked while they were only ever the last argument of an
> ALTERNATIVE(), which appends its own newline. 16/17 concatenates them
> with ASM_CLRERR:
>
> ASM_RDMSR_IMM ASM_CLRERR, X86_FEATURE_MSR_IMM, \
>
> so the .long operand runs into the xor. From
> make arch/x86/kernel/cpu/common.s:
>
> .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long 266xor %rdx,%rdx
>
> paravirt-msr.h:165: Error: junk at end of line, first unrecognized character is `x'
>
> clang reports "error: unexpected token" in the same place. 71 objects
> fail, the same 71 either way, no vmlinux.
>
> msr.h chooses between the .insn form and the .byte fallback with:
>
> #if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24100
>
> Two kinds of toolchain end up on the .byte side of that test:
>
> - GNU as older than 2.41. RHEL 9 and CentOS Stream 9 ship 2.35, and
> Documentation/process/changes.rst sets the minimum at 2.30, so this
> is a supported configuration rather than an old outlier.
> - clang, any version. CONFIG_AS_IS_GNU is never set for clang, so the
> && short-circuits and the version comparison is never reached. Your
> 08/17 comment already notes that clang has no .insn support.
Indeed, looks like we're missing support for .insn for x86.
Filed https://github.com/llvm/llvm-project/issues/225916.
(Please do file bugs against the toolchain when you encounter issues
like this, and cc someone from kernel development).
>
> gcc with binutils 2.41 or newer takes the .insn path, where both macros
> do end in a separator, and is unaffected.
>
> Reproduced on v7.3-rc2 with your v3 00/13, v2 0/5 and v5 00/17 applied in
> that order, x86_64 defconfig plus HYPERVISOR_GUEST, PARAVIRT, XEN and
> XEN_PV. gcc 11.5.0 with GNU as 2.35.2, and clang 21.1.7.
>
> Terminating both fallbacks fixes it, and both toolchains then build
> vmlinux with no errors or warnings:
>
> diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
> index eba325ecfe4c..529c13553c63 100644
> --- a/arch/x86/include/asm/msr.h
> +++ b/arch/x86/include/asm/msr.h
> @@ -78,9 +78,9 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
> * form MSR access instructions reference %rax as the register operand.
> */
> #define ASM_RDMSR_IMM \
> - " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
> + " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]\n\t"
> #define ASM_WRMSRNS_IMM \
> - " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]"
> + " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]\n\t"
> #endif
>
> #define RDMSR_AND_SAVE_RESULT \
>
> ASM_WRMSRNS needs no change, _ASM_BYTES() already emits a semicolon.
>
> The WRMSRNS line belongs in 08/17 and the RDMSR line in 10/17.
>
> Thanks,
> Shreshth
--
Thanks,
~Nick Desaulniers
^ permalink raw reply [flat|nested] 23+ messages in thread* Re: [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions
2026-09-23 19:36 ` [PATCH v5 00/17] x86/msr: Inline rdmsr/wrmsr instructions Shreshth Srivastava
2026-09-23 20:15 ` Nick Desaulniers
@ 2026-09-25 10:25 ` Jürgen Groß
1 sibling, 0 replies; 23+ messages in thread
From: Jürgen Groß @ 2026-09-25 10:25 UTC (permalink / raw)
To: Shreshth Srivastava, linux-kernel, x86, linux-coco, kvm,
linux-hyperv, virtualization, llvm
Cc: tglx, mingo, bp, dave.hansen, hpa, xin, nathan, ndesaulniers,
jpoimboe, peterz, boris.ostrovsky, xen-devel
[-- Attachment #1.1.1: Type: text/plain, Size: 3224 bytes --]
On 23.09.26 21:36, Shreshth Srivastava wrote:
> On 11.09.26 10:41, Juergen Gross wrote:
>> When building a kernel with CONFIG_PARAVIRT_XXL the paravirt
>> infrastructure will always use functions for reading or writing MSRs,
>> even when running on bare metal.
>
> Hi Juergen,
>
> This doesn't build with CONFIG_PARAVIRT_XXL=y. 16/17 and 17/17 are where
> it breaks, but the cause is the .byte fallbacks: ASM_WRMSRNS_IMM from
> 08/17 and ASM_RDMSR_IMM from 10/17 don't end in a separator, unlike the
> .insn variants above them.
>
> #define ASM_RDMSR_IMM \
> " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
>
> That worked while they were only ever the last argument of an
> ALTERNATIVE(), which appends its own newline. 16/17 concatenates them
> with ASM_CLRERR:
>
> ASM_RDMSR_IMM ASM_CLRERR, X86_FEATURE_MSR_IMM, \
>
> so the .long operand runs into the xor. From
> make arch/x86/kernel/cpu/common.s:
>
> .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long 266xor %rdx,%rdx
>
> paravirt-msr.h:165: Error: junk at end of line, first unrecognized character is `x'
>
> clang reports "error: unexpected token" in the same place. 71 objects
> fail, the same 71 either way, no vmlinux.
>
> msr.h chooses between the .insn form and the .byte fallback with:
>
> #if defined(CONFIG_AS_IS_GNU) && CONFIG_AS_VERSION >= 24100
>
> Two kinds of toolchain end up on the .byte side of that test:
>
> - GNU as older than 2.41. RHEL 9 and CentOS Stream 9 ship 2.35, and
> Documentation/process/changes.rst sets the minimum at 2.30, so this
> is a supported configuration rather than an old outlier.
> - clang, any version. CONFIG_AS_IS_GNU is never set for clang, so the
> && short-circuits and the version comparison is never reached. Your
> 08/17 comment already notes that clang has no .insn support.
>
> gcc with binutils 2.41 or newer takes the .insn path, where both macros
> do end in a separator, and is unaffected.
>
> Reproduced on v7.3-rc2 with your v3 00/13, v2 0/5 and v5 00/17 applied in
> that order, x86_64 defconfig plus HYPERVISOR_GUEST, PARAVIRT, XEN and
> XEN_PV. gcc 11.5.0 with GNU as 2.35.2, and clang 21.1.7.
>
> Terminating both fallbacks fixes it, and both toolchains then build
> vmlinux with no errors or warnings:
>
> diff --git a/arch/x86/include/asm/msr.h b/arch/x86/include/asm/msr.h
> index eba325ecfe4c..529c13553c63 100644
> --- a/arch/x86/include/asm/msr.h
> +++ b/arch/x86/include/asm/msr.h
> @@ -78,9 +78,9 @@ static inline void do_trace_rdpmc(u32 msr, u64 val, int failed) {}
> * form MSR access instructions reference %rax as the register operand.
> */
> #define ASM_RDMSR_IMM \
> - " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]"
> + " .byte 0xc4,0xe7,0x7b,0xf6,0xc0; .long %c[msr]\n\t"
> #define ASM_WRMSRNS_IMM \
> - " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]"
> + " .byte 0xc4,0xe7,0x7a,0xf6,0xc0; .long %c[msr]\n\t"
> #endif
>
> #define RDMSR_AND_SAVE_RESULT \
>
> ASM_WRMSRNS needs no change, _ASM_BYTES() already emits a semicolon.
>
> The WRMSRNS line belongs in 08/17 and the RDMSR line in 10/17.
Thanks, will be fixed in V6.
Juergen
[-- Attachment #1.1.2: OpenPGP public key --]
[-- Type: application/pgp-keys, Size: 3743 bytes --]
[-- Attachment #2: OpenPGP digital signature --]
[-- Type: application/pgp-signature, Size: 495 bytes --]
^ permalink raw reply [flat|nested] 23+ messages in thread