An NVMe Power-State Bug Kept Crashing My Ceph Cluster
Visit link →
The same class of NVMe drive failed three times in 48 hours on my three-node Thunderbolt cluster: a real NVMe firmware bug, a smartctl red herring, a fix I already had half of by accident, and a ten-hour alerting blind spot I didn't know I had.
2026/w33/an-nvme-power-state-bug-kept-crashing-my-ceph-cluster