Assembly Hall of Shame
Mirrored from Hacker News — AI on Front Page for archival readability. Support the source by reading on the original site.
Instruction latency analysis usually focuses on performance
optimization—making code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of
single-instruction performance.
x86: fxrstor64
Strategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a
high-latency MMIO region in the PCIe fabric, then starve the fabric while the
load is in flight — a fleet of hammer cores pounds a different high-latency
MMIO register with tight 4-byte reads, saturating the PCIe root complex and
endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must
queue behind all that contending traffic.
Contender: AMD Ryzen 7 5800H
; CPU 0 — timed instruction movl $0xfcc68830, %rsi fxrstor64 %rsi ; CPUs 1..N — hammer loop against a different high-latency location movl 0xfcc68858, %eax
🏆 Score: 198,002,498,236 cycles
🏆 Time: 62 seconds
A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.
vmovdqu 0xfcc003b1, %ymm0
- Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored.
- Trapped/emulated/virtualized instructions may only time the trap, not the handler.
- Instructions must not be interruptible.
rep movs,pause, etc. are disqualified. - Times are normalized based on the CPU base clock frequency.
- All platforms must be in their factory stock configurations - no hardware modifications.
27. nop
Strategy: nop does nothing. It opens the leaderboard accordingly.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
nop
Score: 1 cycles
Time: 0 nanoseconds
26. nop16
Strategy: Regular nop was too short, but how do we make nothing take
longer? Try a lonnnnnng nop.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)
Score: 20 cycles
Time: 7 nanoseconds
25. rdtsc
Strategy: Just a reference instruction to get our bearings.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
rdtscScore: 49 cycles
Time: 18 nanoseconds
24. idiv
Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
xorq %rax, %rax ; rax = 0 (low 64 bits of dividend) movq $2, %rdx ; rdx = 2 (high 64 bits: full dividend = 2^65) movq $5, %rbx ; divisor → quotient = 2^65/5 ≈ 7.4×10^18 idivq %rbx
Score: 77 cycles
Time: 28 nanoseconds
23. enter
Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
enter $0, $31 ; 0 bytes allocated, nesting depth 31 (maximum)
Score: 112 cycles
Time: 41 nanoseconds
22. fldl
Strategy: Try a small denormal to trigger an FP microcode assist.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x0000000000000001, %rax movq %rax, -8(%rsp) fldl -8(%rsp)
Score: 133 cycles
Time: 49 nanoseconds
21. clflush
Strategy: Just ensure the cache line is dirty.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
clflush (%rax) ; rax -> dirty cache line resident in L3
Score: 165 cycles
Time: 60 nanoseconds
20. fsin
Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x7fffffffffffffff, %rax movq %rax, -8(%rsp) fldl -8(%rsp) fsin
Score: 257 cycles
Time: 94 nanoseconds
19. mfence
Strategy: Saturate all write-combining line-fill buffers with movnti
stores to distinct cache lines, forcing mfence to drain the full LFB write
path to the uncore before retiring.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movnti %r9, 0*64(%rdi) ; ×16 distinct cache lines — saturate the write-combining LFBs ; … movnti %r9, 15*64(%rdi) mfence ; must drain all pending LFB writes before retiring
Score: 326 cycles
Time: 120 nanoseconds
18. mov cr3
Strategy: Nothing for now, just check how long it takes to invalidate the TLB.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov %rax, %cr3
Score: 352 cycles
Time: 110 nanoseconds
17. fadd
Strategy: Hit x87 FP microcode assist path by using denormal source operand.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
fldl subnorm ; 1e-310: value < DBL_MIN, biased exponent = 0 faddl subnorm ; source is subnormal → FP microcode assist
Score: 677 cycles
Time: 249 nanoseconds
16. split lock
Strategy: Align lock-prefixed operand to straddle cache-line
boundary, forcing CPU to assert the external bus lock rather than using the fast
MESI cache-coherence path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
; split_ptr % 64 == 63 — dword spans bytes 63 (line N) and 64–66 (line N+1) lock xaddl %r9d, (%rdi)
Score: 865 cycles
Time: 319 nanoseconds
15. fdiv -
Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x3ff0000000000000, %rax ; 1.0 (normal dividend) movq %rax, -8(%rsp) fldl -8(%rsp) ; ST(0) = 1.0 movabsq $0x0000002000000000, %rax ; 6.79e-313 (subnormal divisor) movq %rax, -8(%rsp) fdivl -8(%rsp) ; ST(0) = 1.0 / subnormal → FP assist
Score: 883 cycles
Time: 325 nanoseconds
14. cpuid
Strategy: Use rakefield to find the highest latency CPUID leaves.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movl $6, %eax
cpuid
Score: 1248 cycles
Time: 460 nanoseconds
13. rdrand
Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
rdrand %rax
Score: 5,579 cycles
Time: 2.057 microseconds
12. wrmsr
Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movl $0x17b, %ecx ; MCG_CTL wrmsr
Score: 34,304 cycles
Time: 10.742 microseconds
11. out
Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0xf019, %dx outl %eax, %dx
Score: 49,857 cycles
Time: 15.580 microseconds
10. rdmsr
Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.
Contender: VIA Eden Processor 800MHz
movl $0x133, %ecx ; undocumented MSR rdmsr
Score: 161,602 cycles
Time: 202.004 microseconds
9. wbinvd
Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
wbinvdScore: 1,616,480 cycles
Time: 506.165 microseconds
Strategy: Target I/O port mapped to an ACPI PM block where an unaligned 4-byte read decodes into multiple non-posted loads from wherever this port goes.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0x0413, %dx inl %dx, %eax
Score: 12,524,415 cycles
Time: 3.921769 milliseconds
7. mov
Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, hit unkown GPU register.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movl 0xfcc003b0, %esi
Score: 443,937,696 cycles
Time: 139.010268 milliseconds
6. mov rax -
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 8-byte MMIO read to get two dword register accesses, which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movq 0xfcc003b0, %rax
Score: 887,716,864 cycles
Time: 277.971228 milliseconds
5. vmovdqu xmm -
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 16-byte MMIO read to get four dword register accesses, which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %xmm0
Score: 1,774,555,776 cycles
Time: 555.664133 milliseconds
4. vmovdqu ymm -
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte MMIO read to get eight dword register accesses, which still isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %ymm0
Score: 3,549,079,296 cycles
Time: 1.111345034 s
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte unaligned MMIO read to get nine dword register accesses, which is even less allowed than the aligned version, but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b1, %ymm0
Score: 4,453,212,256 cycles
Time: 1.394428818 seconds
2. fxrstor64 (baseline) -
Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, isolate region near 0's and offset state to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O transactions through slowest available memory aperture.
Contender: AMD Ryzen 7 5800H
movl $0xfcc68830, %rsi fxrstor64 %rsi
Score: 74,584,168,512 cycles
Time: 23.354502677 seconds
1. 🏆 fxrstor64 🏆
Strategy: Extend fxrstor64 (baseline) by starving the fabric while the load is in flight — a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
; CPU 0 — timed instruction movl $0xfcc68830, %rsi fxrstor64 %rsi ; CPUs 1..N — hammer loop against a different high-latency location movl 0xfcc68858, %eax
🏆 Score: 198,002,498,236 cycles
🏆 Time: 62 seconds
Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach
from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size ->
1,000,000,000,000 cycles
Contender: TODO
; XCR0 must enable AMX components (bits 17-18); state area ~8KB xrstor64 (%rsi) ; rsi -> MMIO region, same technique as fxrstor64
- T.B.D.
- T.B.D.
The assembly hall-of-shame is a research effort from Christopher Domas (@xoreaxeaxeax).

Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.