Problem with unstable system - I'm struggling
Problem with unstable system - I'm struggling
Hello everyone
I'm hoping someone can assist me with this issue. I've been experiencing problems with an unstable system for more than a year. There are frequent BSODs, random crashes, and the computer just stops responding completely.
The situation began in February of last year. I put my PC to sleep and left for a short time to shop, then came back home and tried to restart it. When I turned on the monitor, no signal appeared, and even changing the HDMI port didn't help. A few days later, using the original HDMI cable worked properly, which was unusual.
Since then, the problem has worsened. The crashes are linked to nvlddmkm.sys and often trigger a DPC_WATCHDOG_VIOLATION (133). I've reinstalled Windows, updated the display drivers multiple times, and used memtest to check my RAM—all came back normal. I also updated the BIOS, which seemed to smooth things out for a week before things deteriorated again.
Today I had another unexpected reboot, accompanied by a new error from WHEA Logger.
A critical hardware failure has occurred.
Reported by component: Processor Core
Error Source: Machine Check Exception
Error Type: Cache Hierarchy Error
Processor APIC ID: 0
The details section provides more information.
My system specifications are:
Ryzen 5 3600
Gigabyte x570i Pro Wifi
MSI RTX 2080 Super
Corsair SF750 Platinum PSU
Kingston HyperX Fury 2x8gb RAM
Please let me know if you can offer any help.
It could be useful to try sfc/DISM; the latest version should have addressed that issue.
Do you have any other branded RAM available to test, even just one stick?
Are temperatures being tracked?
What are you observing in Event Viewer?
I have also tested the sfc/DISM multiple times, but only minor changes occurred. There are no other RAM issues. I’m keeping an eye on temperatures—CPU hits 70°C during games, GPU reaches 80°C. It might be worth noting that another forum suggested running FurMark yesterday, which I did; it only ran for a couple of minutes and the GPU hit 77°C before the whole system froze completely. I recognize signs of an impending crash: fans slow down, screen flickers, then restarts and slows again. The event viewer shows errors like nvlddmkm and kernel power issues.
This is my second Ryzen 5 3600; the first one failed in 2021 along with my motherboard, but they were still under warranty. In 2020, as a new PC user and making some mistakes, I ran it with the CPU (likely the cause) and this GPU too, which may now be causing problems again.
Consider using a standard fan and positioning it so the air flows straight into the housing with the side panel or glass removed. Try some tests and let us know the results.
Lots of updates - none of them positive.
I reset the bios and changed the settings, no change.
I used Furmark to try a stress test, this caused a crash.
Random reboot which spat out this error in event viewer regarding WHEA Logger.
A fatal hardware error has occurred.
"Reported by component: Processor Core
Error Source: Machine Check Exception
Error Type: Cache Hierarchy Error
Processor APIC ID: 0
The details view of this entry contains further information."
I stress tested the CPU using prime 95. This gave no errors.
I have booted the system with a clean boot, at first I had no crashes but then they returned.
Got a 133 DPC WATCHDOG VIOLATION again. Then randomly lost signal and locked up again. Had to hold power to shut down as usual. Rebooted and the screen refused to show anything but the system was fully responsive. Logged in and heard case fans respond to login as well as number lock light respond and LEDs react and set themselves as per software. Shut down the system, switched back on a picture came up fine. During this GPU fans were still running.
Then had a new error code - 0x144: BUGCODE_USB3_DRIVER.
Then had new errors in event viewer and reliability monitor.
Event viewer: Dwminit - The Desktop Window Manager process has exited. (Process exit code: 0x0000042b, Restart count: 1, Primary display device ID: NVIDIA GeForce RTX 2080 SUPER)
About 30 seconds before that crash was Application Hang error. Saying
The program explorer.exe version 10.0.26100.3323 stopped interacting with Windows and was closed. To see if more information about the problem is available, check the problem history in the Security and Maintenance control panel.
In reliability history, a new Live Kernal Dump.
VIDEO_MINIPORT_BLACK_SCREEN_LIVEDUMP (1b8)
This time caused by dxgkrnl.sys.
I thought it would be worth mentioning that both my RAM and GPU are the oldest original components still in my system. These are now over 5 years old and my system is used everyday. My CPU and motherboard died mid 2021 and were replaced. I'm not sure what the lifespan on these components are but perhaps they are wearing out?
Another thing that has just popped into my head. I can't remember if it was 2022 or 2023 but one day I switched my system on and it was near the RAM sticks as soon as I pushed the power button there was a spark and the system immediately shut down. However I pushed the power button and the system was perfectly fine.
Attempted to open a game today, noticed a big slow down in the mouse performance and the system totally locked up after a couple of minutes. Had to hold the power button as per usual. Checked event viewer and reliability monitor but no errors showed this time.
I tried using a fan and plugging the system straight into a power socket and saw no change.