Logs of random NVMe errors
Logs of random NVMe errors
Hi there, here are your PC details: CPU i9-14900K, board MSI PRO Z790-A, storage Kingston Fury Renegade 1TB PCIe 4 SSD, running Ubuntu 24.10 on Linux version 6.11.0-21-generic. When reviewing NVMe error logs via `nvme-cli` on Linux, you notice several issues: numerous errors appear, such as a massive count of 73484 entries with an invalid field message. The smart log also reports a critical warning about temperature at 123 °F (324 K) and a low endurance score. Everything else seems normal in the smart checks. Are you worried about the drive? It appears to be brand new, but the logs suggest potential problems that might need attention.
I also own a Kingston Fury Renegade 2TB, but it's running at 0tb. In the SMART section under "Number of Error Information Log Entries," a tiny count increases each time it boots. I've searched for details, but everyone with the same drive seems to report this number rising slightly with every boot. Likely, just this model logs errors or warnings, while my more affordable NV2 doesn't record anything.
It seems odd here. Could this be a software problem? I reached out to Kingston and will check their response.
I recently installed a new experimental drive for Proxmox and noticed the same pattern—growth of 2 every 4-8 seconds. It’s reassuring to see similar updates appear only after startup, though I’m a bit concerned. My OCD makes me cautious, but I think it’s probably safe. Unfortunately, Kingston doesn’t offer an update option for Linux, so I won’t be able to fix this myself by replacing the drive. I’m planning to ignore it and see if it resolves on its own. Would anyone have heard from Kingston about this issue? I’d appreciate their perspective.
I've been investigating more closely — it seems related to my monitoring tools (telegraf, lm-sensors, glances for home assistant). By removing certain sensors from the Kingston drive, I significantly reduced errors, except for occasional pings from lm-sensors every few seconds. I'm still encountering a few issues each minute, likely tied to my glances setup. I'm still working through the problem there...
I found this discussion interesting. The same warnings appear in Promox for a 2tb Fury Renegade when I check `nvme error-log /dev/nvme0` on a brand-new drive. Also, during `sensors` I notice these issues which are typical for this device: nvme-pci-0200 Adapter: PCI adapter Composite: +27.9°C (low = -20.1°C, high = +83.8°C) (crit = +88.8°C) ERROR: Can't get value of subfeature temp3_min: I/O error ERROR: Can't get value of subfeature temp3_max: I/O error Sensor 2: +63.9°C (low = +0.0°C, high = +0.0°C). When using `smartctl -a /dev/nvme0` the logs keep increasing rapidly—after a few days it reaches about 1,372 entries. But with Glances running in a Docker container, the count rises by two every couple of seconds. Closing the Glances page slows the increase noticeably. @robertoleonardo you mention several monitoring tools might be contributing. You’ve managed to reduce errors by omitting certain sensors not present on the drive—can you clarify what that involved? I was almost ready to return the drive to Amazon, but now it seems safe to keep it if I can lower the count further.
Updated the configuration to exclude the temp2 sensor from lm-sensors. Created a tailored config file so it doesn’t keep polling sensor2. Noted the drive identifier and followed steps to stop Glances from scanning those sensors. Adjusted settings via nano and verified no further error logs appeared. Looking forward to addressing Glances polling later.