not a resource per se, but i recommend getting a pair of machinist's indicator holders. they make the best helping hands i've ever used when retrofitted with alligator clips, because once tightened they have zero springback so they permit holding parts far more precisely than the usual blue plastic affair. i have a pair of these (https://a.co/d/0dyGqDEh) clockwise tools ones and they're fine, but there are definitely cheaper ones out there that'll work.
Location: New York City, USA
Remote: Yes
Willing to relocate: No
Technologies: Python, C & C++, Rust; kernel module development; Verilog, Verilator, Vivado, Cadence Virtuoso; analog hardware
Résumé/CV: https://khz.ac/cv.pdf
Email: work@khz.ac
i'm nausicaä van west, and i'm a software, firmware, & hardware engineer; i've a background in solar power electronics, infosec (including classified work), circuit design, and a little bit of mechanical engineering & robotics. recent-ish grad (2024) but with more work experience than you would think. if you have electrical engineering work, send me an email!
iirc i address this in a footnote --- absence of SMART errors has very little predictive power, presence has slightly more but it depends which errors they are and how they accumulate
BTW transfer errors aren't indicating the drive errors. SATA is a quite nimble protocol so CRC errors do indicate the problem between the drive and the controller.
the reason RAID is not backup, aside from snapshotting, is colocated drives (same faulty machine, same flammable location, same malicious thunderhead) have a far higher chance of failing at the same time due to evnironmental effects --- hence the decision to have half the drives stored remotely. having a third copy, stored cold (blu-rays are a good idea) helps, but only if its failure is sufficiently independent af the others to make a real dent in the probabilities (e.g., a safe-deposit box or a friend's house, not where any of the drives are!)
oh it's definitely a bit absurd, but i felt i needed a workout.... hence this post. multiple sites, hardware diversity, staggered upgrades, cold storage, are all really good ways to improve this, as they pull failure events closer to independence