Occasional freezes and system failures after testing various solutions.
Occasional freezes and system failures after testing various solutions.
Hey all, I have had an issue for over a year now where my computer will randomly crash out of nowhere, sometimes I get the BSOD reading "WHEA_UNCORRECTABLE_ERROR", other times the screens go black, and then it restarts. I am at my wits end with it, because I have tried so many things and so many troubleshooting steps to see if they would fix it but nothing has worked so far. Things I have tried: -Updated Drivers (everything from GPU to LAN driver) -Updated BIOS -Updated Windows -Restored Windows -Removed ALL overclocking from my CPU, GPU and RAM, even tried underclocking and undervolting my CPU -Tried REBAR on and off -Removed all plugged in peripheral devices -shrank my set up to the bare bones, 1 monitor, my boot drive, mouse and keyboard -Replaced my old GPU (which ran at 90 Degrees c permanently, I thought it was that, it wasn't) -Disabled everything at startup -Ran Memtest overnight and it found no errors in my RAM -Stress tested my CPU with cinebench overnight and it didn't crash -Stress tested my new GPU overnight with Furmark and it didn't crash -Cranked all cooling to 100% and even used 2 desk fans to confirm it wasn't a thermal issue -Uninstalling all antivirus software, including Malwarebytes. (re-installed malwarebytes afterwards) -Tried turning off all manner of windows BS, from auto updates, to firewall settings, to power settings, nothing changed. Reset them to where they were afterwards -Disabled Gamemode and any Nvidia optimisation, didn't fix it -All thermals are fine, my CPU sits at 54 when i'm using my pc normally, and around 60-70 when using it for production -my new GPU just sits at 55 solid and even under load doesn't move There are probably some things I have missed, but those are what I can remember off the top of my head. There is also a weird thing it does that could be related but I am not sure, once, while the computer is on, at random the entire system will stutter for a few seconds, mouse, audio, screens, everything. It's been doing that for well over a year though, and the random crashes only started in the last year. System specs: Cpu: Ryzen 7 5800x OLD GPU (incase it matters): MSI Ventus 2x 3060TI New GPU: Gigabyte 4070TI RAM: 24GB Corsair vengeance RGB pro. CMW16GX4M2D3600C18 (one of my sticks died a few months back, so down from 32 to 24) (Also one of the sticks doesn't have the same serial, but I know the RAM is fine due to Memtest and running this config for 3 years) PSU: Corsair RM750 Motherboard: ASUS ROG B550 F Gaming 7 drives, but Crystal disk mark says they are all at 100% health Any and all help would be greatly appreciated, as I use this PC for music production and it's hindering my work. I am attaching pictures of my system through HWmonitor, as well as a picture of the last log before the last crash. EDIT: Forgot to mention, I am running Windows 10
Visit C:\Windows\Minidump and verify the presence of any minidump files. If found, return to the Windows directory and transfer the entire Minidump folder to the Downloads folder (use your desktop if OneDrive isn't available). Compress the copied folder and attach it to a message. Please adhere strictly to the provided directions since Windows typically resists such actions in this area. WHEA indicates a CPU or PCIe device problem, and Microsoft has also linked NVMe errors to WHEA. When the crash originates from storage, it usually suggests the drive was offline during the failure, potentially blocking dump file generation if the crashed drive lacks the page file. Therefore, NVMe is likely the culprit. If no dumps exist, we can modify a registry entry to capture the necessary crash details displayed on the BSOD screen. Follow the steps carefully as Windows generally dislikes these modifications. WHEA means hardware trouble with the CPU or a PCIe interface. Microsoft has also noted NVMe-related issues in this context. If you encounter a BSOD, check if it persists on the screen (without dumps) – this step may not be needed. If it restarts normally after a short time, skip that part and proceed to the next guide, turning off automatic reboots. To manually restart, press the power button. To show extra details on the BSOD, edit the registry at HKEY_LOCAL_MACHINE\System\CurrentControlSet\Control\CrashControl, add a DWORD named "DisplayParameters" with value 1, and save changes. Reboot to apply. Afterward, photograph this configuration. The General section only lists detected components; the detailed error data resides in the Details tab’s RawData field. Please paste or save this information because the RawData is extensive and I won’t want to transcribe it manually. Also, note that SMART metrics have been reduced for NVMe SSDs, which is disappointing. The percentage value reflects wear, not current drive health—it mainly shows how much of the warranty write capacity remains. A drive at full capacity can still be in poor condition compared to one at 70%.
Hey there! Glad you got the files ready. I’ve already compressed the minidump files—two were sent to a friend who couldn’t crack them, which is why I’ve zipped two more. I’ll also run a Reg edit to highlight extra BSOD details, since that’s really helpful even beyond this issue. For checking drive health, what tools do you suggest? Thanks! Minidump.rar
When files aren't available, focus on NVMe systems where reliable tools are limited. For SATA, CDI remains functional, though some devices reuse the same SMART status as NVMe. Regarding dump files, those not compressed initially reveal hardware issues with the CPU. Three indicate problems during L1 cache reads, two show bus errors between cores and cache, and one reports a timeout. After reviewing other dumps, the pattern points to a defective CPU. Consider disabling overclocking, undervolting, and checking XMP profiles and PBO settings in BIOS. Watch CPU temperatures to avoid overheating. The 5000 series may face voltage-related read errors rather than timeouts. Updating the BIOS is advisable since flashback protection can help recover from crashes during updates. Other rare causes include socket issues or insufficient PSU voltage. Generally, testing on another system or swapping CPUs provides the best insight.
It's good to know things are in order. I've already disabled all overclocking and updated the BIOS to the latest version. I'll double-check everything again, since I've had to reset multiple times recently. Probably needs a new CPU to test for a few weeks—if it still crashes, I can return it.