this post was submitted on 01 Sep 2026
16 points (94.4% liked)

Linux

18935 readers
110 users here now

Welcome to c/linux!

Welcome to our thriving Linux community! Whether you're a seasoned Linux enthusiast or just starting your journey, we're excited to have you here. Explore, learn, and collaborate with like-minded individuals who share a passion for open-source software and the endless possibilities it offers. Together, let's dive into the world of Linux and embrace the power of freedom, customization, and innovation. Enjoy your stay and feel free to join the vibrant discussions that await you!

Rules:

  1. Stay on topic: Posts and discussions should be related to Linux, open source software, and related technologies.

  2. Be respectful: Treat fellow community members with respect and courtesy.

  3. Quality over quantity: Share informative and thought-provoking content.

  4. No spam or self-promotion: Avoid excessive self-promotion or spamming.

  5. No NSFW adult content

  6. Follow general lemmy guidelines.

founded 3 years ago
MODERATORS
16
SSD dying? (sopuli.xyz)
submitted 9 hours ago* (last edited 9 hours ago) by Svinhufvud@sopuli.xyz to c/linux@lemmy.world
 

This is a second time now in the last couple of months that my openSUSE Tumbleweed server has become inaccessible over the night with the terminal printing over and over:

BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45080, flush 8, corrupt 0, gen 0
systemd-journald[404389]: Failed to send WATCHDOG=1 notification message: Transport endpoint is not connected
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45081, flush 8, corrupt 0, gen 0
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45102, flush 8, corrupt 0, gen 0
systemd-journald[404389]: Failed to send WATCHDOG=1 notification message: Transport endpoint is not connected
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45103, flush 8, corrupt 0, gen 0
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45106, flush 8, corrupt 0, gen 0
...
systemd-journald[404389]: Failed to send WATCHDOG=1 notification message: Transport endpoint is not connected
...
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45128, flush 8, corrupt 0, gen 0
BTRFS error (device dm-0 state EA): bdev /dev/mapper/cr_root errs: wr 549, rd 45138, flush 8, corrupt 0, gen 0

Both times a restart seemingly fixed everything, as in the computer booted up normally and my services started up and seemed to work just fine.

I ran

btrfs scrub status /

UUID:             redacted
Scrub started:    Tue Sep  1 09:26:40 2026
Status:           finished
Duration:         0:02:36
Total to scrub:   495.58GiB
Rate:             2.96GiB/s
Error summary:    no errors found

and

smartctl -a /dev/nvme0

smartctl 7.5 2025-04-30 r5714 [x86_64-linux-7.2.2-1-default] (SUSE RPM)
Copyright (C) 2002-25, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Number:                       Samsung SSD 990 PRO 4TB
Serial Number:                      redacted
Firmware Version:                   4B2QJXD7
PCI Vendor/Subsystem ID:            redacted
IEEE OUI Identifier:                redacted
Total NVM Capacity:                 4,000,787,030,016 [4.00 TB]
Unallocated NVM Capacity:           0
Controller ID:                      1
NVMe Version:                       2.0
Number of Namespaces:               1
Namespace 1 Size/Capacity:          4,000,787,030,016 [4.00 TB]
Namespace 1 Utilization:            733,311,033,344 [733 GB]
Namespace 1 Formatted LBA Size:     512
Namespace 1 IEEE EUI-64:            redacted
Local Time is:                      Tue Sep  1 10:01:48 2026 EEST
Firmware Updates (0x16):            3 Slots, no Reset required
Optional Admin Commands (0x0017):   Security Format Frmw_DL Self_Test
Optional NVM Commands (0x0055):     Comp DS_Mngmt Sav/Sel_Feat Timestmp
Log Page Attributes (0x2f):         S/H_per_NS Cmd_Eff_Lg Ext_Get_Lg Telmtry_Lg Log0_FISE_MI
Maximum Data Transfer Size:         512 Pages
Warning  Comp. Temp. Threshold:     82 Celsius
Critical Comp. Temp. Threshold:     85 Celsius

Supported Power States
St Op     Max   Active     Idle   RL RT WL WT  Ent_Lat  Ex_Lat
 0 +     9.39W       -        -    0  0  0  0        0       0
 1 +     9.39W       -        -    1  1  1  1        0       0
 2 +     9.39W       -        -    2  2  2  2        0       0
 3 -   0.0400W       -        -    3  3  3  3     4200    2700
 4 -   0.0050W       -        -    4  4  4  4      500   21800

Supported LBA Sizes (NSID 0x1)
Id Fmt  Data  Metadt  Rel_Perf
 0 +     512       0         0

=== START OF SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

SMART/Health Information (NVMe Log 0x02, NSID 0xffffffff)
Critical Warning:                   0x00
Temperature:                        45 Celsius
Available Spare:                    100%
Available Spare Threshold:          10%
Percentage Used:                    2%
Data Units Read:                    36,634,591 [18.7 TB]
Data Units Written:                 88,390,637 [45.2 TB]
Host Read Commands:                 173,217,468
Host Write Commands:                2,664,966,303
Controller Busy Time:               19,920
Power Cycles:                       30
Power On Hours:                     9,824
Unsafe Shutdowns:                   15
Media and Data Integrity Errors:    0
Error Information Log Entries:      0
Warning  Comp. Temperature Time:    0
Critical Comp. Temperature Time:    0
Temperature Sensor 1:               45 Celsius
Temperature Sensor 2:               47 Celsius

Error Information (NVMe Log 0x01, 16 of 64 entries)
No Errors Logged

Self-test Log (NVMe Log 0x06, NSID 0xffffffff)
Self-test status: No self-test in progress
Num  Test_Description  Status                       Power_on_Hours  Failing_LBA  NSID Seg SCT Code
 0   Extended          Completed without error                9824            -     -   -   -    -
 1   Short             Completed without error                9824            -     -   -   -    -

which dont seem to show any errors. The SSD writes are quite high though. The SSD in question is under two years old, and acts both as a boot drive and a data drive for my data.

So is my SSD cooked? What troubleshooting steps should I take? Or is this a warranty case? (The data is backed up alredy.)

Thanks in advance!

you are viewing a single comment's thread
view the rest of the comments
[–] Svinhufvud@sopuli.xyz 1 points 8 hours ago (1 children)

m.2 nvme, and the only nvme drive in the system. I have two hard drives with SATA, but I doubt that's the issue.

[–] frongt@lemmy.zip 1 points 4 hours ago (1 children)

I'd just reseat the drive. Easy to do, and might fix it.

If it really is dying, though, I'd just keep using it until you can't stand it, then replace it, since you already have is m a backup.

[–] Svinhufvud@sopuli.xyz 1 points 4 hours ago

Yeah but it is still covered by the warranty, so if it is indeed dying, I'd rather RMA it now than than later.

But I will try the reseating.