Highlighting a 3950x format and emphasizing it strongly.
Highlighting a 3950x format and emphasizing it strongly.
So I have a bright and shinny 3950x. So far this thing has proved to be the beast it was advertised to be. But there's always that one fly that gets in the ointment. The whole rig (which is not at all intended for gaming) looks like this: 3950x Gigabyte X570 Aorus Pro 32GB Corsair memory (don't remember the model, but I chose it off the QVL for the MB and to clear the cooler.) MSI GTX-1660 Super (no, I'm really not gaming with this) 1 TB Intel MVME (because I could) Seasonic GX-750 gold PSU Noctua NH-D15 Windows 10 Pro All the latest drivers have been installed and I flashed the board to the most recent UEFI. It has handled Cinebench and Aida64 testing with no problem (both were run for hours, I know not complete tests, but I only built it yesterday.) Temps are reasonable (hovering around 61 under load, with occasional spikes to around 80 for a second or two, can't really explain why, but they're there) and I don't see any errors reported. Being a bit old school, I then fired up Prime95. And right out of the gate got errors on small FFTs. There were consistent on the same "core" numbers (18 and 19, so I assume them to be virtual as the physicals seem to be 0 through 15.) Long story short, I could "walk" the errors around by messing with the number of works I used and how many "cores" I told each to exercise. Everything in the UEFI was set stock, I hadn't even enabled XMP yet. (Though when I did that it made the errors worse.) I didn't disable Turbo (or whatever AMD is calling it these days.) So in a sense the chip was trying to OC itself when it detected load. I consider that "normal" behavior and should have been included in the test. I started playing with XMP as the memory voltage seemed low. But that got me looking at other voltages and ultimately lead me to start thinking about vdroop. (That was a long and twisted path that I won't bore everyone with.) Ultimately if I did find that if I set LLC to "Low", the system stopped throwing errors and Prime95 ran for slightly over 8 hours before I stopped it. I have another small FFT run going right now and it's behaving similarly. What I'm now faced with is what do with this mess. I do think the chip is beast and even in my short time I've come to really like it, but I also want something that's long term stable and doesn't have monsters lurking inside it just waiting for the right (and inopportune) time to come leaping out and reeking havoc. While I now know how to keep said monsters locked up where they don't show themselves, I don't really like that I had to tweak something in the UEFI to get it to be stable. I've never had to do that with any other chip and I've never seen a chip that didn't pass basic (albeit strenuous) tests. I cannot decide: If there is some problem with how the "Auto"/"Normal"/"Standard" LLC setting is implemented in the Gigabyte UEFI (all three seem to be the same.) If there is some system power supply issue that is showing up on power hungry chips like the 3950x and the Threadripper series (there are several folks in other forums reporting similar issues with Prime95 and these chips) If this has something to do with the fact that the memory seems to be under-volted. If there is a problem with my specific chip. If there is a problem with some other component (motherboard, memory, or PSU.) All are going to get tested as best I can. Or if I should be glad that I know how to keep this controlled and be happy with what I have. I did also read a description posted by a guy in another forum who had an issue similar to mine. He decided to RMA his chip, and has ended up, two RMAs later, with one that behaves worse than either of it's predecessors. That is train I totally don't want to get on. I figured that I'd see what folks here thought about this. Thank you in advance for any input you might have, and sorry for the long post.
I wouldn't stress small FFT even with stock configurations. If you're not encountering issues beyond p95, I'd consider you fine. The voltage should be at least 1.1V and RAM around 1.35V or 1.4V. The 3950x isn't overly power-hungry unless you overclock it, and even then it won't perform well, so a 750W PSU is sufficient. The biggest concern seems to be voltage droop, but if it's just p95 errors I wouldn't be too concerned—even someone who stopped using p95 would still be cautious.
You're observing a voltage drop on the +12V rail during brief FFT measurements. Turbo behavior is typical and not an issue for this chip, but several BIOS versions might increase power demands, pushing it into an overdrive state. Turning off PBO or manually adjusting power settings could help. Implementing an LLC regulator is a suitable solution, as it mitigates high-current voltage sag while maintaining efficiency during light loads. Excessive LLC usage can be detrimental, especially for boot performance, so ensure it's not pushed to its limits.
I'm confident I reset the CMOS during the UEFI flash, though I don't recall doing it on purpose. It's odd I haven't remembered this step, and I might try again just to feel more complete. I didn't use a voltmeter at the time, but I can check tools like HWMonitor or HWiNFO later if needed.
It's a solid story overall, though it misses some visual appeal. Prime95 isn't completely transparent. A few LLC adjustments seem reasonable; low starting values often lead to clock throttling and errors. If you're aiming for performance, consider OCCT AVX2. Benchmark your memory with Linpack.
@ShrimpBrine, your point isn't entirely far from the truth. The design leans heavily into tempered glass with minimal RGB options—most of what remains is pre-integrated into the GPU and motherboard, and much of that is disabled. For a short while I imagined creating something impressive enough to impress Vegas, but reality set in. It will remain unobtrusive, elegantly placed, almost like the monolith from 2001 A Space Odyssey. OCCT is on my agenda for testing. I’m a bit worried about heat management because of the case layout (there’s ample space around and behind it) and I have four 120 fans inside, plus two 140s on the NH-D15, which are directed straight at the exhaust on the back, along with the GPU fans and the chipset’s little issue. I just need to ensure it won’t quietly fail in that corner. Is OCCT’s documentation solid? I’ve downloaded it but haven’t checked yet. I’m unfamiliar with Linpack; I’ll look it up online. Previously I intended to use Memtest, but I’d rather stick with G.Skill memory—though I couldn’t find any that would fit the cooler without issues. It’s a personal choice, as Ford and Chevy both work well, but people tend to stick with one brand. The real frustration is finding Samsung-based RAM with 16GB sticks that don’t clutter the cooler (which adds about an extra centimeter). G.Skill offers several options, but they’d likely shift the front fan out for a RGB strip I’d never actually see. Anything low-profile or bare, available on QVL with Samsung chips, and without RGB would be ideal. Only then would I be confident it won’t get damaged internally before I commit.
I gave it a try. OCCT seemed more fun to watch than P95. After an hour running all 32 cores at full speed (4.19 GHz) without any issues, I might let it go longer. But I couldn’t figure out what OCCT was doing to the system and ended up stopping it for the night. Temperatures were lower than with P95, and the CPU stayed at turbo throughout. The problem likely came from the transition down from turbo when the CPU reached its limit.
Update on 6/16/20. By the end of the day, I realized this device couldn't handle math without an LLC set to "Medium." ("Low" only provided some stability, I needed "Medium" for real reliability.) It makes no sense that any CPU would fail at stock settings when running software. My chip was returned via RMA. It should take about two to three weeks before it comes back. It's frustrating, but I know it's the correct course. I'll aim for a better sample next time.