No, consumer Ryzen does not offer L3 as NUMA configuration.
No, consumer Ryzen does not offer L3 as NUMA configuration.
I just found this discussion elsewhere. It asks whether this feature applies to consumer Ryzen models or is exclusive to Epyc systems. Right now I don’t have a desktop AMD machine to verify, and my laptop doesn’t offer such settings either. It seems many users struggle with cache fragmentation, and this could improve how the OS and software manage data placement. Windows 10’s handling isn’t ideal, especially since manual affinity adjustments work well. I haven’t tried it on Windows 11 to compare the scheduler’s performance.
R9 5900x with MSI B550 Tomahawk. The only NUMA-related adjustment in my BIOS is "NUMA nodes per socket." It can be chosen as Auto (default) or set from NPS0 through NPS4. The note explains it defines how many NUMA nodes should be on each socket, with zero trying to mix the two together. I’m not sure if a CCD counts as a socket in this context or if they’re referring to physical slots the board doesn’t have.
It's a standard single socket consumer board. That’s why I considered a CCD might be seen as a "socket" in that context. You could check the number of NUMA domains in Windows or Linux to see if that matches your setup.
Thanks for the offer. I think the best way is to use prime95 built in benchmark as it also reports the CPU structure. Download latest version. Set min and max FFT size to 2048, uncheck "use hyper-threading" and reduce the test per time to 5s and let it run. Please do this with default and 2 setting. There is a small chance that 2048 might not work as optimal FFT sizes may vary with CPU, in which case set min and max to something like 2000-2100 and retry, should find something in a region. Example output for my 7920X is in spoiler below. This can be found in file results.bench.txt after running, so no need to extract anything from screen. While I'm most interested in the reported CPU organisation with that setting, the bench result is also separately interesting. Prime95 will pick number of "workers" appropriate to the CPU cores. This is usually 1, max, and something in between. e.g. for 12 core CPU, 1 worker = 12 cores per task. 2 workers = 6 cores per task, and so on. The FFT size of 2048 means each task has a data set of 16MB. Multiple tasks will take multiples of 16MB. If the total exceeds (usually L3) cache, you become ram impacted in performance. Based on my past testing I'd expect 2 workers to be optimal on 5900X, but it will be interesting to see how 1 worker compares. As Prime95 is aware of CPU structure, this setting shouldn't matter, but other similar software may not be and that's where it could help. Spoiler Intel® Core i9-7920X CPU @ 2.90GHz CPU speed: 3789.61 MHz, 12 hyperthreaded cores CPU features: Prefetchw, SSE, SSE2, SSE4, AVX, AVX2, FMA, AVX512F L1 cache size: 12x32 KB, L2 cache size: 12x1 MB, L3 cache size: 16896 KB L1 cache line size: 64 bytes, L2 cache line size: 64 bytes Machine topology as determined by hwloc library: Machine#0 (total=59624216KB, Backend=Windows, OSName=Windows, WindowsBuildEnvironment=MinGW, OSRelease=10, OSVersion=10.0.19045, Hostname=GARUDA, Architecture=x86_64, hwlocVersion=2.6.0, ProcessName=prime95.exe) Package (total=59624216KB, CPUVendor=GenuineIntel, CPUFamilyNumber=6, CPUModelNumber=85, CPUModel="Intel® Core i9-7920X CPU @ 2.90GHz", CPUStepping=4) L3 (size=16896KB, linesize=64, ways=11, Inclusive=0) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000003) PU#0 (cpuset: 0x00000001) PU#1 (cpuset: 0x00000002) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000000c) PU#2 (cpuset: 0x00000004) PU#3 (cpuset: 0x00000008) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000030) PU#4 (cpuset: 0x00000010) PU#5 (cpuset: 0x00000020) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000000c0) PU#6 (cpuset: 0x00000040) PU#7 (cpuset: 0x00000080) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000300) PU#8 (cpuset: 0x00000100) PU#9 (cpuset: 0x00000200) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000c00) PU#10 (cpuset: 0x00000400) PU#11 (cpuset: 0x00000800) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00003000) PU#12 (cpuset: 0x00001000) PU#13 (cpuset: 0x00002000) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000c000) PU#14 (cpuset: 0x00004000) PU#15 (cpuset: 0x00008000) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00030000) PU#16 (cpuset: 0x00010000) PU#17 (cpuset: 0x00020000) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000c0000) PU#18 (cpuset: 0x00040000) PU#19 (cpuset: 0x00080000) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00300000) PU#20 (cpuset: 0x00100000) PU#21 (cpuset: 0x00200000) L2 (size=1024KB, linesize=64, ways=16, Inclusive=0) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00c00000) PU#22 (cpuset: 0x00400000) PU#23 (cpuset: 0x00800000) Prime95 64-bit version 30.8, RdtscTiming=1 Timings for 2048K FFT length (12 cores, 1 worker): 0.72 ms. Throughput: 1393.16 iter/sec. Timings for 2048K FFT length (12 cores, 3 workers): 2.65, 2.64, 2.45 ms. Throughput: 1164.47 iter/sec. Timings for 2048K FFT length (12 cores, 12 workers): 12.67, 12.16, 12.05, 12.38, 12.15, 12.33, 12.32, 12.23, 12.92, 12.12, 12.05, 12.67 ms. Throughput: 973.10 iter/sec.
Ok, I hope I did it correctly Used the pre-selected throughput benchmark, changed min/max FFT to 2048, disabled HT, time 5s, left everything else at its default. My BIOS seems to have a weird bug. When I change the NUMA option it tells me nothing has changed when I go to save. When I use the search and change the option, it seems to detect the changes I've made. So that's what I went with. In task manager, Performance > CPU > Graph > Right click > Change graph to. The "NUMA nodes" options is grayed out, which supposedly means I have a single NUMA node. Doesn't seem any setting influences this. Doesn't seem to change too much about the p95 results. Here's the results: Nodes per socket: Auto Spoiler [Sun Jan 8 12:57:46 2023] Compare your results to other computers at http://www.mersenne.org/report_benchmarks AMD Ryzen 9 5900X 12-Core Processor CPU speed: 4625.11 MHz, 12 hyperthreaded cores CPU features: 3DNow! Prefetch, SSE, SSE2, SSE4, AVX, AVX2, FMA L1 cache size: 12x32 KB, L2 cache size: 12x512 KB, L3 cache size: 2x32 MB L1 cache line size: 64 bytes, L2 cache line size: 64 bytes Machine topology as determined by hwloc library: Machine#0 (total=30874688KB, Backend=Windows, OSName=Windows, WindowsBuildEnvironment=MinGW, OSRelease=10, OSVersion=10.0.19044, Hostname=DESKTOP-7CRB6SJ, Architecture=x86_64, hwlocVersion=2.6.0, ProcessName=prime95.exe) Package (total=30874688KB, CPUVendor=AuthenticAMD, CPUFamilyNumber=25, CPUModelNumber=33, CPUModel="AMD Ryzen 9 5900X 12-Core Processor ", CPUStepping=0) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000003) PU#0 (cpuset: 0x00000001) PU#1 (cpuset: 0x00000002) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000000c) PU#2 (cpuset: 0x00000004) PU#3 (cpuset: 0x00000008) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000030) PU#4 (cpuset: 0x00000010) PU#5 (cpuset: 0x00000020) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000000c0) PU#6 (cpuset: 0x00000040) PU#7 (cpuset: 0x00000080) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000300) PU#8 (cpuset: 0x00000100) PU#9 (cpuset: 0x00000200) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000c00) PU#10 (cpuset: 0x00000400) PU#11 (cpuset: 0x00000800) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00003000) PU#12 (cpuset: 0x00001000) PU#13 (cpuset: 0x00002000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000c000) PU#14 (cpuset: 0x00004000) PU#15 (cpuset: 0x00008000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00030000) PU#16 (cpuset: 0x00010000) PU#17 (cpuset: 0x00020000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000c0000) PU#18 (cpuset: 0x00040000) PU#19 (cpuset: 0x00080000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00300000) PU#20 (cpuset: 0x00100000) PU#21 (cpuset: 0x00200000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00c00000) PU#22 (cpuset: 0x00400000) PU#23 (cpuset: 0x00800000) Prime95 64-bit version 30.8, RdtscTiming=1 Timings for 2048K FFT length (12 cores, 1 worker): 0.84 ms. Throughput: 1195.16 iter/sec. Timings for 2048K FFT length (12 cores, 2 workers): 1.12, 1.12 ms. Throughput: 1784.17 iter/sec. Timings for 2048K FFT length (12 cores, 12 workers): 22.44, 29.97, 28.79, 30.39, 29.75, 29.05, 28.09, 28.71, 26.97, 28.46, 28.65, 30.43 ms. Throughput: 424.02 iter/sec. Nodes per socket: 2 Spoiler [Sun Jan 8 13:01:19 2023] Compare your results to other computers at http://www.mersenne.org/report_benchmarks AMD Ryzen 9 5900X 12-Core Processor CPU speed: 4625.16 MHz, 12 hyperthreaded cores CPU features: 3DNow! Prefetch, SSE, SSE2, SSE4, AVX, AVX2, FMA L1 cache size: 12x32 KB, L2 cache size: 12x512 KB, L3 cache size: 2x32 MB L1 cache line size: 64 bytes, L2 cache line size: 64 bytes Machine topology as determined by hwloc library: Machine#0 (total=30667528KB, Backend=Windows, OSName=Windows, WindowsBuildEnvironment=MinGW, OSRelease=10, OSVersion=10.0.19044, Hostname=DESKTOP-7CRB6SJ, Architecture=x86_64, hwlocVersion=2.6.0, ProcessName=prime95.exe) Package (total=30667528KB, CPUVendor=AuthenticAMD, CPUFamilyNumber=25, CPUModelNumber=33, CPUModel="AMD Ryzen 9 5900X 12-Core Processor ", CPUStepping=0) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000003) PU#0 (cpuset: 0x00000001) PU#1 (cpuset: 0x00000002) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000000c) PU#2 (cpuset: 0x00000004) PU#3 (cpuset: 0x00000008) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000030) PU#4 (cpuset: 0x00000010) PU#5 (cpuset: 0x00000020) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000000c0) PU#6 (cpuset: 0x00000040) PU#7 (cpuset: 0x00000080) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000300) PU#8 (cpuset: 0x00000100) PU#9 (cpuset: 0x00000200) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000c00) PU#10 (cpuset: 0x00000400) PU#11 (cpuset: 0x00000800) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00003000) PU#12 (cpuset: 0x00001000) PU#13 (cpuset: 0x00002000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000c000) PU#14 (cpuset: 0x00004000) PU#15 (cpuset: 0x00008000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00030000) PU#16 (cpuset: 0x00010000) PU#17 (cpuset: 0x00020000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000c0000) PU#18 (cpuset: 0x00040000) PU#19 (cpuset: 0x00080000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00300000) PU#20 (cpuset: 0x00100000) PU#21 (cpuset: 0x00200000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00c00000) PU#22 (cpuset: 0x00400000) PU#23 (cpuset: 0x00800000) Prime95 64-bit version 30.8, RdtscTiming=1 Timings for 2048K FFT length (12 cores, 1 worker): 0.84 ms. Throughput: 1192.56 iter/sec. Timings for 2048K FFT length (12 cores, 2 workers): 1.14, 1.14 ms. Throughput: 1752.38 iter/sec. Timings for 2048K FFT length (12 cores, 12 workers): 28.97, 30.00, 29.40, 28.19, 24.89, 29.36, 11.22, 31.42, 31.15, 30.50, 30.86, 30.66 ms. Throughput: 462.41 iter/sec. Nodes per socket: 0 Spoiler [Sun Jan 8 13:11:45 2023] Compare your results to other computers at http://www.mersenne.org/report_benchmarks AMD Ryzen 9 5900X 12-Core Processor CPU speed: 4625.16 MHz, 12 hyperthreaded cores CPU features: 3DNow! Prefetch, SSE, SSE2, SSE4, AVX, AVX2, FMA L1 cache size: 12x32 KB, L2 cache size: 12x512 KB, L3 cache size: 2x32 MB L1 cache line size: 64 bytes, L2 cache line size: 64 bytes Machine topology as determined by hwloc library: Machine#0 (total=30862288KB, Backend=Windows, OSName=Windows, WindowsBuildEnvironment=MinGW, OSRelease=10, OSVersion=10.0.19044, Hostname=DESKTOP-7CRB6SJ, Architecture=x86_64, hwlocVersion=2.6.0, ProcessName=prime95.exe) Package (total=30862288KB, CPUVendor=AuthenticAMD, CPUFamilyNumber=25, CPUModelNumber=33, CPUModel="AMD Ryzen 9 5900X 12-Core Processor ", CPUStepping=0) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000003) PU#0 (cpuset: 0x00000001) PU#1 (cpuset: 0x00000002) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000000c) PU#2 (cpuset: 0x00000004) PU#3 (cpuset: 0x00000008) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000030) PU#4 (cpuset: 0x00000010) PU#5 (cpuset: 0x00000020) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000000c0) PU#6 (cpuset: 0x00000040) PU#7 (cpuset: 0x00000080) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000300) PU#8 (cpuset: 0x00000100) PU#9 (cpuset: 0x00000200) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000c00) PU#10 (cpuset: 0x00000400) PU#11 (cpuset: 0x00000800) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00003000) PU#12 (cpuset: 0x00001000) PU#13 (cpuset: 0x00002000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000c000) PU#14 (cpuset: 0x00004000) PU#15 (cpuset: 0x00008000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00030000) PU#16 (cpuset: 0x00010000) PU#17 (cpuset: 0x00020000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000c0000) PU#18 (cpuset: 0x00040000) PU#19 (cpuset: 0x00080000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00300000) PU#20 (cpuset: 0x00100000) PU#21 (cpuset: 0x00200000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00c00000) PU#22 (cpuset: 0x00400000) PU#23 (cpuset: 0x00800000) Prime95 64-bit version 30.8, RdtscTiming=1 Timings for 2048K FFT length (12 cores, 1 worker): 0.81 ms. Throughput: 1237.28 iter/sec. Timings for 2048K FFT length (12 cores, 2 workers): 1.16, 1.14 ms. Throughput: 1739.45 iter/sec. Timings for 2048K FFT length (12 cores, 12 workers): 29.01, 28.69, 28.42, 27.52, 29.03, 17.13, 7.82, 30.66, 30.17, 29.97, 30.89, 30.78 ms. Throughput: 525.47 iter/sec. Nodes per socket: 4 Spoiler [Sun Jan 8 13:13:44 2023] Compare your results to other computers at http://www.mersenne.org/report_benchmarks AMD Ryzen 9 5900X 12-Core Processor CPU speed: 4600.09 MHz, 12 hyperthreaded cores CPU features: 3DNow! Prefetch, SSE, SSE2, SSE4, AVX, AVX2, FMA L1 cache size: 12x32 KB, L2 cache size: 12x512 KB, L3 cache size: 2x32 MB L1 cache line size: 64 bytes, L2 cache line size: 64 bytes Machine topology as determined by hwloc library: Machine#0 (total=30863048KB, Backend=Windows, OSName=Windows, WindowsBuildEnvironment=MinGW, OSRelease=10, OSVersion=10.0.19044, Hostname=DESKTOP-7CRB6SJ, Architecture=x86_64, hwlocVersion=2.6.0, ProcessName=prime95.exe) Package (total=30863048KB, CPUVendor=AuthenticAMD, CPUFamilyNumber=25, CPUModelNumber=33, CPUModel="AMD Ryzen 9 5900X 12-Core Processor ", CPUStepping=0) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000003) PU#0 (cpuset: 0x00000001) PU#1 (cpuset: 0x00000002) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000000c) PU#2 (cpuset: 0x00000004) PU#3 (cpuset: 0x00000008) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000030) PU#4 (cpuset: 0x00000010) PU#5 (cpuset: 0x00000020) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000000c0) PU#6 (cpuset: 0x00000040) PU#7 (cpuset: 0x00000080) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000300) PU#8 (cpuset: 0x00000100) PU#9 (cpuset: 0x00000200) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00000c00) PU#10 (cpuset: 0x00000400) PU#11 (cpuset: 0x00000800) L3 (size=32768KB, linesize=64, ways=16, Inclusive=0) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00003000) PU#12 (cpuset: 0x00001000) PU#13 (cpuset: 0x00002000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x0000c000) PU#14 (cpuset: 0x00004000) PU#15 (cpuset: 0x00008000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00030000) PU#16 (cpuset: 0x00010000) PU#17 (cpuset: 0x00020000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x000c0000) PU#18 (cpuset: 0x00040000) PU#19 (cpuset: 0x00080000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00300000) PU#20 (cpuset: 0x00100000) PU#21 (cpuset: 0x00200000) L2 (size=512KB, linesize=64, ways=8, Inclusive=1) L1d (size=32KB, linesize=64, ways=8, Inclusive=0) Core (cpuset: 0x00c00000) PU#22 (cpuset: 0x00400000) PU#23 (cpuset: 0x00800000) Prime95 64-bit version 30.8, RdtscTiming=1 Timings for 2048K FFT length (12 cores, 1 worker): 0.85 ms. Throughput: 1170.19 iter/sec. Timings for 2048K FFT length (12 cores, 2 workers): 1.17, 1.11 ms. Throughput: 1751.82 iter/sec. Timings for 2048K FFT length (12 cores, 12 workers): 30.08, 7.67, 30.01, 30.11, 29.75, 29.83, 28.64, 11.99, 29.91, 28.45, 29.73, 29.37 ms. Throughput: 551.91 iter/sec.
Thanks for sharing the findings. In my view, everything appears identical. I’m curious if this is similar to an older setup with two NUMA nodes. My earlier belief that two workers would offer optimal performance still holds. I wasn’t expecting a different outcome, but I did wonder about the four-worker configuration—maybe the software isn’t detecting it or it’s not visible. I’d like to experiment myself, but I’d need a compatible system. Still, considering Zen 4, we might hold off for now. Appreciate your help.
I believe enabling "ACPI SRAT L3 Cache as NUMA Domain" is necessary for having two NUMA nodes on consumer Ryzen systems—it's a distinct setting compared to "NUMA nodes per socket." The options are clearly visible in the screenshots I provided, though I can't test it because I only have one CCX.