F5F Stay Refreshed Hardware Desktop Explain Memory BW and AMD 7950X memory controller in simpler terms.

Explain Memory BW and AMD 7950X memory controller in simpler terms.

Explain Memory BW and AMD 7950X memory controller in simpler terms.

B
baconandfries
Member
215
07-19-2023, 04:08 AM
#1
In my professional work, I operate a 3D simulation tool that demands significant memory and processing power to generate gigabytes of information. Memory bandwidth plays a crucial role in achieving quicker solutions, which can span several hours. Since large datasets often exceed the capacity of faster cache memory—something common in gaming—I see this clearly. The best evidence comes from AMD X3D versions with enhanced caches and modest bandwidth, achieving better performance than standard gaming setups. Essentially, everything hinges on bandwidth. It’s similar to how a book works: it’s swift when nearby, slower if stored elsewhere, and even more so when you have to fetch it from a distant location. After searching for a reliable benchmark linking solve time with memory speed, I found some useful data. PassMark offers RAM tests, though they’re less effective for multi-core tasks. The two examples I analyzed were an AMD 7950X with dual 48GB sticks and an AMD Threadripper 3955WX with eight 16GB sticks. Their performance metrics were impressive—around 83.6 GB/s and 79.3 GB/s respectively. Running the actual solver, I observed 2117 seconds for the 7950X versus 2180 seconds for the 3955WX. This aligns closely with PassMark results, though the Aida64 test suggested the 3955WX might edge ahead at higher speeds. Adjusting the RAM speed on the 7950X to 4800MHz reduced its performance by about 13%, confirming the sensitivity of bandwidth. Further testing showed that lowering the memory speed to 60K on the 7950X cut performance by roughly 14%. Inside Performance Test, a more detailed memory assessment revealed that threads (cores) impact results significantly when HT/SMT is disabled. This highlights a potential issue with AMD’s memory controller—its bandwidth drops after 10C. It appears that only 8C of memory per bank is effectively used, and losing one bank drastically reduces speed. This seems to be a notable limitation. Intel’s performance tends to improve linearly with core count on certain architectures, but the 9700K reached only 37 GB/s on DDR4 while 5000-series RAM handled 50-54 GB/s. Today’s DDR5 shows mixed results: some systems max out at 84 Kb/s, whereas others push 90-96 K. This inconsistency makes accurate benchmarking difficult. I also experimented with SMT settings—disabling it improved solve times by about 2 seconds for long runs. Overall, the data suggests that memory bandwidth and cache configuration are critical factors, especially for multi-threaded applications. I’m still exploring options for RAM with lower latency (<10ns) to further enhance performance. While a 7800MHz 36th core at CL36 is appealing, it’s pricier, and DDR5 options are promising but not yet widely available. I’m hoping to find more detailed PassMark charts for the 7900X, 13900K, or 14900K with SMT off, along with precise latency figures, to validate these observations.
B
baconandfries
07-19-2023, 04:08 AM #1

In my professional work, I operate a 3D simulation tool that demands significant memory and processing power to generate gigabytes of information. Memory bandwidth plays a crucial role in achieving quicker solutions, which can span several hours. Since large datasets often exceed the capacity of faster cache memory—something common in gaming—I see this clearly. The best evidence comes from AMD X3D versions with enhanced caches and modest bandwidth, achieving better performance than standard gaming setups. Essentially, everything hinges on bandwidth. It’s similar to how a book works: it’s swift when nearby, slower if stored elsewhere, and even more so when you have to fetch it from a distant location. After searching for a reliable benchmark linking solve time with memory speed, I found some useful data. PassMark offers RAM tests, though they’re less effective for multi-core tasks. The two examples I analyzed were an AMD 7950X with dual 48GB sticks and an AMD Threadripper 3955WX with eight 16GB sticks. Their performance metrics were impressive—around 83.6 GB/s and 79.3 GB/s respectively. Running the actual solver, I observed 2117 seconds for the 7950X versus 2180 seconds for the 3955WX. This aligns closely with PassMark results, though the Aida64 test suggested the 3955WX might edge ahead at higher speeds. Adjusting the RAM speed on the 7950X to 4800MHz reduced its performance by about 13%, confirming the sensitivity of bandwidth. Further testing showed that lowering the memory speed to 60K on the 7950X cut performance by roughly 14%. Inside Performance Test, a more detailed memory assessment revealed that threads (cores) impact results significantly when HT/SMT is disabled. This highlights a potential issue with AMD’s memory controller—its bandwidth drops after 10C. It appears that only 8C of memory per bank is effectively used, and losing one bank drastically reduces speed. This seems to be a notable limitation. Intel’s performance tends to improve linearly with core count on certain architectures, but the 9700K reached only 37 GB/s on DDR4 while 5000-series RAM handled 50-54 GB/s. Today’s DDR5 shows mixed results: some systems max out at 84 Kb/s, whereas others push 90-96 K. This inconsistency makes accurate benchmarking difficult. I also experimented with SMT settings—disabling it improved solve times by about 2 seconds for long runs. Overall, the data suggests that memory bandwidth and cache configuration are critical factors, especially for multi-threaded applications. I’m still exploring options for RAM with lower latency (<10ns) to further enhance performance. While a 7800MHz 36th core at CL36 is appealing, it’s pricier, and DDR5 options are promising but not yet widely available. I’m hoping to find more detailed PassMark charts for the 7900X, 13900K, or 14900K with SMT off, along with precise latency figures, to validate these observations.

Z
Zen_Browncoat
Junior Member
8
07-19-2023, 06:57 AM
#2
A 13700K could work. Your kits support DDR5 8000 CL34, and most configurations below that should be fine, though you’ll need to check your specific settings chart—those advanced features might require a paid tool.
Z
Zen_Browncoat
07-19-2023, 06:57 AM #2

A 13700K could work. Your kits support DDR5 8000 CL34, and most configurations below that should be fine, though you’ll need to check your specific settings chart—those advanced features might require a paid tool.

M
Maxhos_
Junior Member
30
07-19-2023, 11:19 AM
#3
Here are the results you asked for. I completed four distinct tests: 6400 CL32, 7200 CL36, and 8000 CL34. The 8000 CL34 was tested with both stock CPU configurations and without HT/E cores enabled. The 6400 CL32 used only XMP settings, while the 7200 CL36 followed exact settings from a G.Skill kit (details I’ll discuss later). The 8000 CL34 involved fully customized sub-timings. All tests used a 2x24GB RAM package, though the 2x16GB would likely perform slightly better in theory.

- 6400 CL32: Result included
- 7200 CL36: Result included
- 8000 CL34 (E cores/HT off): Result included
- 8000 CL34 (E cores/HT on): Result included

The process was done with a 2x24GB kit, though the 2x16GB should generally run faster.
M
Maxhos_
07-19-2023, 11:19 AM #3

Here are the results you asked for. I completed four distinct tests: 6400 CL32, 7200 CL36, and 8000 CL34. The 8000 CL34 was tested with both stock CPU configurations and without HT/E cores enabled. The 6400 CL32 used only XMP settings, while the 7200 CL36 followed exact settings from a G.Skill kit (details I’ll discuss later). The 8000 CL34 involved fully customized sub-timings. All tests used a 2x24GB RAM package, though the 2x16GB would likely perform slightly better in theory.

- 6400 CL32: Result included
- 7200 CL36: Result included
- 8000 CL34 (E cores/HT off): Result included
- 8000 CL34 (E cores/HT on): Result included

The process was done with a 2x24GB kit, though the 2x16GB should generally run faster.

S
SkeleRooMC
Junior Member
20
07-19-2023, 01:32 PM
#4
Here’s a revised version of your message:

Thank you for running some calculations for me. Under Advanced there’s a memory test that offers a closer look at threaded outcomes. It’s only available until the 30-day review period ends, but it remains accessible otherwise. Turning off hyper-threading makes it easier to interpret the data. I’m planning to purchase the software since my gaming rig has surpassed the 30-day mark, though the latest build performed well. It seems you’re using a stock 6400 unit with an OC configuration. I own a Corsair 96GB kit at 6400MHz latency: 32-40-40-84. I had to lower the clock speed slightly when running two GPUs, which caused some problems. Another Corsair model with 6000MHz settings (30-36-36-76) helped boost performance to around 84500MB/s. I also managed to reduce voltage from 1.3 or 1.35 down to 1.1 without any issues. I’ll check the tRFC:tREFI on the box tomorrow to see more details. From what I observe in your results, it looks like adjusting the ratio for 6400MHz to match the 7200 settings could increase bandwidth for that configuration. I’ll need to experiment a bit further. Thanks again for your guidance—it’s useful to hear what others are doing and learn from them.
S
SkeleRooMC
07-19-2023, 01:32 PM #4

Here’s a revised version of your message:

Thank you for running some calculations for me. Under Advanced there’s a memory test that offers a closer look at threaded outcomes. It’s only available until the 30-day review period ends, but it remains accessible otherwise. Turning off hyper-threading makes it easier to interpret the data. I’m planning to purchase the software since my gaming rig has surpassed the 30-day mark, though the latest build performed well. It seems you’re using a stock 6400 unit with an OC configuration. I own a Corsair 96GB kit at 6400MHz latency: 32-40-40-84. I had to lower the clock speed slightly when running two GPUs, which caused some problems. Another Corsair model with 6000MHz settings (30-36-36-76) helped boost performance to around 84500MB/s. I also managed to reduce voltage from 1.3 or 1.35 down to 1.1 without any issues. I’ll check the tRFC:tREFI on the box tomorrow to see more details. From what I observe in your results, it looks like adjusting the ratio for 6400MHz to match the 7200 settings could increase bandwidth for that configuration. I’ll need to experiment a bit further. Thanks again for your guidance—it’s useful to hear what others are doing and learn from them.

M
MaddiBlake
Member
241
07-19-2023, 09:50 PM
#5
It's accurate that the two kits I own are 2x16GB 6000 CL30-38-38 Flare X5 and a 2x24GB 6400 CL32 Ripjaws S5. The 2x16GB model performs better overall, offering higher maximum frequency and suitable timing options. However, the 2x24GB kit is nearly comparable in both aspects, which is why I prefer using it for extra storage. I can elaborate further if needed, but the main point is that memory kits often don't differ much, particularly with DDR5 technology. If a kit uses the same memory ICs as the top performers, it will simply achieve those speeds—so a 6000 CL30 should handle 8000 CL38 without issues, provided the board and processor support it. That said, the label doesn’t reflect this nuance. Manufacturers usually highlight main timings, while secondary and tertiary settings matter more for real performance. To check actual timing, install the kit in a system and use tools like Aida64 or Thaiphoon Burner to view the SPDE data (note: Windows 11 may complicate this). Alternatively, enable XMP and use software such as MemTweakIt (for Intel) or ZenTimings (for AMD) to see current settings. This adjustment can cut latency noticeably. I also reprocessed the data to include graphs, which show a substantial improvement—around 9.5% more bandwidth. Of course, this setting is quite sensitive to temperature; if you live in a hot climate or have poor airflow, stress testing is wise. Raising it too high can be risky, but generally it’s safe. The 6400 CL32 version was tested, though I didn’t go through it extensively due to potential instability at such speeds. This should give you a clear idea of what to expect.
M
MaddiBlake
07-19-2023, 09:50 PM #5

It's accurate that the two kits I own are 2x16GB 6000 CL30-38-38 Flare X5 and a 2x24GB 6400 CL32 Ripjaws S5. The 2x16GB model performs better overall, offering higher maximum frequency and suitable timing options. However, the 2x24GB kit is nearly comparable in both aspects, which is why I prefer using it for extra storage. I can elaborate further if needed, but the main point is that memory kits often don't differ much, particularly with DDR5 technology. If a kit uses the same memory ICs as the top performers, it will simply achieve those speeds—so a 6000 CL30 should handle 8000 CL38 without issues, provided the board and processor support it. That said, the label doesn’t reflect this nuance. Manufacturers usually highlight main timings, while secondary and tertiary settings matter more for real performance. To check actual timing, install the kit in a system and use tools like Aida64 or Thaiphoon Burner to view the SPDE data (note: Windows 11 may complicate this). Alternatively, enable XMP and use software such as MemTweakIt (for Intel) or ZenTimings (for AMD) to see current settings. This adjustment can cut latency noticeably. I also reprocessed the data to include graphs, which show a substantial improvement—around 9.5% more bandwidth. Of course, this setting is quite sensitive to temperature; if you live in a hot climate or have poor airflow, stress testing is wise. Raising it too high can be risky, but generally it’s safe. The 6400 CL32 version was tested, though I didn’t go through it extensively due to potential instability at such speeds. This should give you a clear idea of what to expect.

X
220
07-20-2023, 12:05 AM
#6
I tested the 7950X today and achieved a maximum of 85,400 MB/s for memory bandwidth. The results are similar regardless of clock speed.
X
xXStrikeBackXx
07-20-2023, 12:05 AM #6

I tested the 7950X today and achieved a maximum of 85,400 MB/s for memory bandwidth. The results are similar regardless of clock speed.

K
Kamikaze_007
Senior Member
625
07-27-2023, 08:10 PM
#7
I was pleased to notice this. It really aligns with how AMD’s memory controller functions. The process is somewhat intricate, so I might overlook certain points or explain them less clearly—feel free to ask for more details. There are three primary clock speeds in the AMD memory setup: the FCLK, the MEMCLK, and the UCLK. The FCLK (infinity fabric) dictates how cores communicate with the memory controller. The MEMCLK refers to the memory clock rate. The UCLK indicates the speed at which the controller operates. On AM5 systems, the FCLK is relatively low—around 2GHz with a 32-byte bus—limiting throughput to roughly 64GB/s. Of course, this applies per chip, so on Ryzen 9 boards you can theoretically reach up to 128GB/s to the CPU, assuming ideal scheduling and no internal chatter. Still, in real-world use, this remains a constraint. Then there are the MEMCLK and UCLK aspects, which seem to define most of the constraints. They should ideally match a 1:1 or 2:1 ratio (since I’m more familiar with Intel naming). Often, higher-end kits cap performance not by hardware but by controller speed—having the controller run at half the memory’s rate can boost overall speeds, though it introduces noticeable delays and reduces bandwidth. This creates a trade-off where top-tier memory often struggles to match both extremes. By default, most systems use Gear 1 until DDR5 6000 is reached, then switch to Gear 2. If you manually adjust to Gear 1 (around 6200-6400 MHz), most chips still function well, though rare models push it higher. Once you hit Gear 2, performance penalties become more obvious after DDR5 7600+, bringing in long boot times and strict voltage requirements. These speeds are mainly viable with single-chip memory configurations. On Intel, the situation is similar but more nuanced—multiple controller modes exist, and while DDR5 offers Gear 2 and Gear 4 options, the latter is mostly relevant for maximum memory rates on LN2 boards. The FCLK in Intel is handled differently, often through ring or uncore clocks, which run faster but aren’t a major bandwidth bottleneck. Overall, it’s not ideal to rely on Gear 2 unless you’re targeting specific memory specs and conditions.
K
Kamikaze_007
07-27-2023, 08:10 PM #7

I was pleased to notice this. It really aligns with how AMD’s memory controller functions. The process is somewhat intricate, so I might overlook certain points or explain them less clearly—feel free to ask for more details. There are three primary clock speeds in the AMD memory setup: the FCLK, the MEMCLK, and the UCLK. The FCLK (infinity fabric) dictates how cores communicate with the memory controller. The MEMCLK refers to the memory clock rate. The UCLK indicates the speed at which the controller operates. On AM5 systems, the FCLK is relatively low—around 2GHz with a 32-byte bus—limiting throughput to roughly 64GB/s. Of course, this applies per chip, so on Ryzen 9 boards you can theoretically reach up to 128GB/s to the CPU, assuming ideal scheduling and no internal chatter. Still, in real-world use, this remains a constraint. Then there are the MEMCLK and UCLK aspects, which seem to define most of the constraints. They should ideally match a 1:1 or 2:1 ratio (since I’m more familiar with Intel naming). Often, higher-end kits cap performance not by hardware but by controller speed—having the controller run at half the memory’s rate can boost overall speeds, though it introduces noticeable delays and reduces bandwidth. This creates a trade-off where top-tier memory often struggles to match both extremes. By default, most systems use Gear 1 until DDR5 6000 is reached, then switch to Gear 2. If you manually adjust to Gear 1 (around 6200-6400 MHz), most chips still function well, though rare models push it higher. Once you hit Gear 2, performance penalties become more obvious after DDR5 7600+, bringing in long boot times and strict voltage requirements. These speeds are mainly viable with single-chip memory configurations. On Intel, the situation is similar but more nuanced—multiple controller modes exist, and while DDR5 offers Gear 2 and Gear 4 options, the latter is mostly relevant for maximum memory rates on LN2 boards. The FCLK in Intel is handled differently, often through ring or uncore clocks, which run faster but aren’t a major bandwidth bottleneck. Overall, it’s not ideal to rely on Gear 2 unless you’re targeting specific memory specs and conditions.

T
TMGC_Oderic
Member
78
07-27-2023, 10:47 PM
#8
From the latest Passmark data I analyzed, several intriguing patterns emerged. My 7950X reached a peak of 85,400 MB/s, while a 7900X managed around 87,500 MB/s. This suggests that the 12C chip may be more efficient, possibly due to memory constraints rather than actual performance differences. Notably, the 7950X3D reported values near 97.9 and 98.6K, hinting at the significant impact of X3D’s large cache on multi-core threaded memory bandwidth. The 7900X3D achieved a high of 95.6K. These results raise questions about whether increased cache size correlates with program-specific memory usage or if it boosts multi-core CPU performance by up to 16.7% (98/84). If I acquire additional CPUs, I’ll share more details. Another point worth noting is Apple’s M2 Ultra model offers just 242K, whereas two M2 Pros linked together function as a single unit at 121K each—compared to the regular M2’s 63K MB/s. This could reflect Passmark’s methodology or an artificial effect in testing.
T
TMGC_Oderic
07-27-2023, 10:47 PM #8

From the latest Passmark data I analyzed, several intriguing patterns emerged. My 7950X reached a peak of 85,400 MB/s, while a 7900X managed around 87,500 MB/s. This suggests that the 12C chip may be more efficient, possibly due to memory constraints rather than actual performance differences. Notably, the 7950X3D reported values near 97.9 and 98.6K, hinting at the significant impact of X3D’s large cache on multi-core threaded memory bandwidth. The 7900X3D achieved a high of 95.6K. These results raise questions about whether increased cache size correlates with program-specific memory usage or if it boosts multi-core CPU performance by up to 16.7% (98/84). If I acquire additional CPUs, I’ll share more details. Another point worth noting is Apple’s M2 Ultra model offers just 242K, whereas two M2 Pros linked together function as a single unit at 121K each—compared to the regular M2’s 63K MB/s. This could reflect Passmark’s methodology or an artificial effect in testing.