F5F Stay Refreshed Hardware Desktop Crash caused by SSD failure (or motherboard/PSU problem)

Crash caused by SSD failure (or motherboard/PSU problem)

Crash caused by SSD failure (or motherboard/PSU problem)

_
_GiovanniPvP_
Member
58
05-29-2023, 10:29 AM
#1
Hello everyone, this is my debut post here, so I’m aiming for clarity while still offering detailed insight. This problem appears fairly complex but seems close to a resolution. My system ran smoothly with the following configuration:

- CPU: AMD Ryzen 7 7800X3D @ 4.2 GHz, 8-core processor
- Cooler: Noctua NH-U12S chromax.black (55 CFM)
- Motherboard: MSI B650 GAMING PLUS WIFI ATX (AM5)
- RAM: Corsair Vengeance RGB 32 GB (2 x 16 GB) DDR5-6000 CL36
- Storage: Western Digital Green 240 GB 2.5" SSD (for Windows), Samsung 970 Evo Plus 2 TB M.2-2280 NVMe SSD
- Drive: Seagate Barracuda Compute 2 TB 3.5" HDD
- Video: XFX Speedster MERC 310 Black Edition RX 7900 XTX 24 GB
- Case: Corsair 4000D Airflow ATX ($84.97)
- PSU: NZXT C850 (2022, 850W, 80+ Gold)
- Fan: Corsair iCUE SP120 RGB ELITE 47.7 CFM (3-pack)
- Power: AVR UPS (also bought for safety)

Initially, everything worked perfectly. Then I ordered an NVMe SSD for my wife’s PC, but it didn’t arrive on time. I used the one that came, which was the Samsung 990 EVO 1 TB. After installing it and doing a clean Windows install, I also updated the BIOS for faster boot times.

For a few weeks, I faced random crashes—especially when switching from heavy gaming to idle. These crashes often triggered a BSOD with errors like UNEXPECTED STORE EXCEPTION or CRITICAL PROCESS DIED. I tried everything: checking settings, updating drivers, resetting the system, using event viewer, and even replacing the SSD.

Eventually, after several attempts, I got an error indicating KERNAL_DATA_INPAGE_ERROR. This strongly suggested a storage issue. I tried plugging in a different SSD, which worked. It made me think of two main possibilities:

1. The motherboard couldn’t handle the power from the new SSD, causing instability.
2. The PSU was supplying too much power, leading to shorts or overheating.

I also considered other less likely reasons like wiring issues or thermal problems. However, since I switched to a different SSD and it functioned, it’s more likely one of those two factors.

Now I’m deciding whether to buy a new PSU, a new motherboard, or both. Given the age of my SSD (about 5 years), replacing it seems reasonable. The wiring in my home might have contributed initially, but it appears resolved now.

I hope this helps clarify what happened and how I resolved it. Thanks for your support!
_
_GiovanniPvP_
05-29-2023, 10:29 AM #1

Hello everyone, this is my debut post here, so I’m aiming for clarity while still offering detailed insight. This problem appears fairly complex but seems close to a resolution. My system ran smoothly with the following configuration:

- CPU: AMD Ryzen 7 7800X3D @ 4.2 GHz, 8-core processor
- Cooler: Noctua NH-U12S chromax.black (55 CFM)
- Motherboard: MSI B650 GAMING PLUS WIFI ATX (AM5)
- RAM: Corsair Vengeance RGB 32 GB (2 x 16 GB) DDR5-6000 CL36
- Storage: Western Digital Green 240 GB 2.5" SSD (for Windows), Samsung 970 Evo Plus 2 TB M.2-2280 NVMe SSD
- Drive: Seagate Barracuda Compute 2 TB 3.5" HDD
- Video: XFX Speedster MERC 310 Black Edition RX 7900 XTX 24 GB
- Case: Corsair 4000D Airflow ATX ($84.97)
- PSU: NZXT C850 (2022, 850W, 80+ Gold)
- Fan: Corsair iCUE SP120 RGB ELITE 47.7 CFM (3-pack)
- Power: AVR UPS (also bought for safety)

Initially, everything worked perfectly. Then I ordered an NVMe SSD for my wife’s PC, but it didn’t arrive on time. I used the one that came, which was the Samsung 990 EVO 1 TB. After installing it and doing a clean Windows install, I also updated the BIOS for faster boot times.

For a few weeks, I faced random crashes—especially when switching from heavy gaming to idle. These crashes often triggered a BSOD with errors like UNEXPECTED STORE EXCEPTION or CRITICAL PROCESS DIED. I tried everything: checking settings, updating drivers, resetting the system, using event viewer, and even replacing the SSD.

Eventually, after several attempts, I got an error indicating KERNAL_DATA_INPAGE_ERROR. This strongly suggested a storage issue. I tried plugging in a different SSD, which worked. It made me think of two main possibilities:

1. The motherboard couldn’t handle the power from the new SSD, causing instability.
2. The PSU was supplying too much power, leading to shorts or overheating.

I also considered other less likely reasons like wiring issues or thermal problems. However, since I switched to a different SSD and it functioned, it’s more likely one of those two factors.

Now I’m deciding whether to buy a new PSU, a new motherboard, or both. Given the age of my SSD (about 5 years), replacing it seems reasonable. The wiring in my home might have contributed initially, but it appears resolved now.

I hope this helps clarify what happened and how I resolved it. Thanks for your support!

P
Plan2031
Junior Member
3
05-29-2023, 11:58 PM
#2
Kernel_data_inpage issues may stem from defective RAM or a malfunctioning storage unit. Download CrystalDiskInfo to assess the condition of each drive. Additionally, perform a memory test to evaluate your RAM status.
P
Plan2031
05-29-2023, 11:58 PM #2

Kernel_data_inpage issues may stem from defective RAM or a malfunctioning storage unit. Download CrystalDiskInfo to assess the condition of each drive. Additionally, perform a memory test to evaluate your RAM status.

M
Miteus_St
Member
56
05-30-2023, 08:37 PM
#3
I ran a RAM test on the system and everything worked smoothly. The Samsung Magician also reported no issues with either SSD.
M
Miteus_St
05-30-2023, 08:37 PM #3

I ran a RAM test on the system and everything worked smoothly. The Samsung Magician also reported no issues with either SSD.

A
AeroJirachi
Junior Member
18
05-31-2023, 04:58 AM
#4
As @BillBill mentioned, check Crystaldiskinfo and memtest86. They're widely recognized and effective tools. I'm not claiming Windows testing or Samsung tricks are bad, but using Crystal and memtest would be beneficial.
A
AeroJirachi
05-31-2023, 04:58 AM #4

As @BillBill mentioned, check Crystaldiskinfo and memtest86. They're widely recognized and effective tools. I'm not claiming Windows testing or Samsung tricks are bad, but using Crystal and memtest would be beneficial.

S
Sven_Weetj
Member
220
05-31-2023, 06:17 AM
#5
It’s interesting how the Ram issue appears only on the NVME SSD while the SATA SSD remains unaffected. Testing multiple SSDs helps clarify whether the problem is hardware-specific or tied to the interface type.
S
Sven_Weetj
05-31-2023, 06:17 AM #5

It’s interesting how the Ram issue appears only on the NVME SSD while the SATA SSD remains unaffected. Testing multiple SSDs helps clarify whether the problem is hardware-specific or tied to the interface type.

D
Draker59
Member
126
06-02-2023, 03:44 AM
#6
I’m experiencing a familiar sense of déjà vu with this discussion. Were you on Discord or Reddit when this problem arose? I’ll respond here without hesitation. It’s frustrating to seem biased, but checking SMART status on NVMe SSDs doesn’t really help much. SMART is essentially the SSD’s self-test data that CrystalDiskInfo interprets. With NVMe drives, this feature has been significantly reduced or removed. From my experience, hundreds of faulty NVMe SSDs have passed any SMART checks, and most have stripped away useful health metrics. The main part of CDI’s top information isn’t reliable for HDDs or SATA SSDs either, since manufacturers control status updates and many are unreliable. Based on the BSOD logs, storage issues seem most likely. The M.2 port or motherboard could also be involved, as these components are more prone to failure. Since you’ve already replaced the SSD, the next possible culprit is probably the motherboard.
D
Draker59
06-02-2023, 03:44 AM #6

I’m experiencing a familiar sense of déjà vu with this discussion. Were you on Discord or Reddit when this problem arose? I’ll respond here without hesitation. It’s frustrating to seem biased, but checking SMART status on NVMe SSDs doesn’t really help much. SMART is essentially the SSD’s self-test data that CrystalDiskInfo interprets. With NVMe drives, this feature has been significantly reduced or removed. From my experience, hundreds of faulty NVMe SSDs have passed any SMART checks, and most have stripped away useful health metrics. The main part of CDI’s top information isn’t reliable for HDDs or SATA SSDs either, since manufacturers control status updates and many are unreliable. Based on the BSOD logs, storage issues seem most likely. The M.2 port or motherboard could also be involved, as these components are more prone to failure. Since you’ve already replaced the SSD, the next possible culprit is probably the motherboard.

D
Dat_Asian_
Member
146
06-02-2023, 07:02 AM
#7
I've posted in a few different places without much success but I swear this is a newly written post written specifically for the forum! I appreciate the advice. I'm assuming its most likely just an issue with that specific port as I had never used it before and I just got unlucky. Thank you!
D
Dat_Asian_
06-02-2023, 07:02 AM #7

I've posted in a few different places without much success but I swear this is a newly written post written specifically for the forum! I appreciate the advice. I'm assuming its most likely just an issue with that specific port as I had never used it before and I just got unlucky. Thank you!

K
KingJaydxn
Member
240
06-19-2023, 02:49 AM
#8
On theory, what are your NVMe controller temperatures? I also have a 4 TB Samsung 990 Pro on my X670E ProArt as my OS drive. The controller on the Samsung 990 Pro NVMe runs hot, especially if you don’t have the heatsinked version and depend only on case airflow or the motherboard’s heatsink cover—these might not be very effective. On your motherboard, the NVMe heatsinks aren’t optimized for maximum cooling. I’ve installed the 4000D, which is a solid case, but it doesn’t provide the best airflow, which could affect performance. For instance, in my setup, because my GPU blocks much vertical airflow (Fractal Torrent) and the first NVMe slot is positioned above it, I had to swap NVMe drives between slots. I also added a small finned heatsink on the 990 Pro, which helped maintain lower temperatures. The 990 Pro still idles around 50-55°C—often exceeding 65°C under similar workloads, nearing Samsung’s stated upper limit of 70°C. A 2 TB Corsair MP600 Pro NH in the other slot stays between 44-48°C. Use tools like HWiNFO64 or a similar SMART monitor to check NVMe temps. If the controller is excessively hot, it could be a contributing factor, even though Samsung claims it should handle up to 70°C. Other potential issues include unstable overclocking, aggressive RAM timings, memory running in an EXPO profile unsupported by the board, a failing PSU, or overly aggressive undervolting/overclocking. Are your system’s clock speeds and voltages set to stock levels?
K
KingJaydxn
06-19-2023, 02:49 AM #8

On theory, what are your NVMe controller temperatures? I also have a 4 TB Samsung 990 Pro on my X670E ProArt as my OS drive. The controller on the Samsung 990 Pro NVMe runs hot, especially if you don’t have the heatsinked version and depend only on case airflow or the motherboard’s heatsink cover—these might not be very effective. On your motherboard, the NVMe heatsinks aren’t optimized for maximum cooling. I’ve installed the 4000D, which is a solid case, but it doesn’t provide the best airflow, which could affect performance. For instance, in my setup, because my GPU blocks much vertical airflow (Fractal Torrent) and the first NVMe slot is positioned above it, I had to swap NVMe drives between slots. I also added a small finned heatsink on the 990 Pro, which helped maintain lower temperatures. The 990 Pro still idles around 50-55°C—often exceeding 65°C under similar workloads, nearing Samsung’s stated upper limit of 70°C. A 2 TB Corsair MP600 Pro NH in the other slot stays between 44-48°C. Use tools like HWiNFO64 or a similar SMART monitor to check NVMe temps. If the controller is excessively hot, it could be a contributing factor, even though Samsung claims it should handle up to 70°C. Other potential issues include unstable overclocking, aggressive RAM timings, memory running in an EXPO profile unsupported by the board, a failing PSU, or overly aggressive undervolting/overclocking. Are your system’s clock speeds and voltages set to stock levels?

F
finnster20
Member
161
06-21-2023, 12:44 AM
#9
I located it on Reddit's TechSupport forum. It looked like the decision was between a PSU and a motherboard, so I expressed doubt about it being the power supply unit.
F
finnster20
06-21-2023, 12:44 AM #9

I located it on Reddit's TechSupport forum. It looked like the decision was between a PSU and a motherboard, so I expressed doubt about it being the power supply unit.

S
SedentarySauS
Senior Member
411
06-21-2023, 01:24 AM
#10
Thanks for the thorough breakdown. Everything is operating at normal clock speeds and voltages on my PC, aside from EXPO. I noticed some elevated temperatures on the controller temperature probe, while the other reading was technically within acceptable limits. The 970 with the Mobo heatshield appears to be functioning properly.
S
SedentarySauS
06-21-2023, 01:24 AM #10

Thanks for the thorough breakdown. Everything is operating at normal clock speeds and voltages on my PC, aside from EXPO. I noticed some elevated temperatures on the controller temperature probe, while the other reading was technically within acceptable limits. The 970 with the Mobo heatshield appears to be functioning properly.