Anecdotally, RAM failures are more common than they used to be in previous decades, I assume due to increased capacities and smaller process nodes. Looser manufacturing standards may also play a part.
The component category I've never seen fail is CPU.
I've worked in and around datacenters for a few decades now and I've seen pretty much everything fail, but power supplies in particular do so at a higher rate than I expected. I've had one literally start smoking in my home office recently too. Luckily, it didn't damage any components, but it was pretty dramatic and abrupt -- just powered on my desktop and got nothing but a fog machine coming out of the rear fan until I yanked the power cord a few seconds later.
My evga 1000w has been going strong for 11 years now. Cheaper power supplies fail often but if you buy quality and oversized they’ll outlast all your other components as they don’t become outdated.
It’s the caps that fail. Electrolytic caps have the highest failure rate of any passive component. My team tries to avoid them in our designs as much as possible. But you’re right, caps in a good design can easily last 20 years, making a power supply the longest lasting component of most builds.
My go to source for reviews was always jonnyguru.com but he closed up shop. Not sure who’s actually doing in-depth performance and component reviews these days.
Power supplies - regardless of quality - die pretty regularly at scale. Only disks fail more often. They are the #2 component we replace in the datacenter.
Are data-center power supplies comparable to higher-end desktop power supplies? And how much do you oversize them compared to the load they will actually handle? E.g. for a desktop buying even double the power supply capacity you need won't really affect the overall PC price that much but I somehow doubt data center operators are willing to not optimize away that cost.
They typically are made by the same OEMs and I would say at par or better quality in general. I’ve had (percentage wise) more consumer grade PSUs fail on me over the past 20 years, and I don’t buy cheap stuff. Although the sample size of failure is low single digits vs thousands on the datacenter end.
I’d say on average DC PSUs are vastly overspeced for most machines. Most machines we work with are not sitting at peak load 24x7, and might be drawing half or so of their max rated power draw. Plus every server is redundant - so the modules themselves are operating at around 20-40% of full rated capacity most of their life. The only time they ever spin up past 50% is during power maintenance or a power failure of some sort that brings one of the redundant power sides down.
They do probably operate in slightly hotter conditions on average, but in much more stabilized temperature environments. Plus of course 24x7x365. But that goes for all components.
I’d say other than fans, failure rates in the PC world vs datacenter world are relatively comparable it terms of what fails most. Dust and temp swings are the large variables the average PC needs to contend with vs the datacenter.
Storage of course is the one that stands out the most - everything else is kind of a rounding error. Here though, datacenter work really does beat on the hardware more than most consumer gear so that’s expected. For every 100 disk swaps we do, I figure we do about 2 or 3 power supply modules. And maybe 1 fan swap.
They just rebrand PSUs, not all of them are bomber. But they have 10yr warranties on PSUs and are good about honoring them. I wish they still sold gpus. I would only buy evga for every component if I could specifically for the painless warranty. Had to warranty an asus mobo and it was like pulling teeth.
Yeah I also had an EVGA 1080 Ti die on me, and got a better GPU as upgrade. Their warranty was worth the premium price for sure.
Overall I have made very good experiences with long hardware warranties. Soldered RAM on my ThinkPad died once and I got an entire new mainboard after almost 3 years (it died 2 months before warranty expired.) New one still works today but the hardware is just too old.
Funny, just had an 8 year old 1600W Corsair Titanium efficiency power supply go. Generally very high quality, GaN transistors, all japanese caps, etc etc. It was still under warranty, and Corsair refunded the original price, stand-up company. Too bad 1600W power supplies have apparently inflated ~50% since then, but still.
> The component category I've never seen fail is CPU.
The CPU can fail if the heatsink falls off! That is something that happens (I've had one come off in my hand that was being held on by nothing but hopes and dreams, fortunately not while the CPU was powered on). More likely in a desktop tower form factor.
Even that is not guaranteed to break a modern CPU though as they all monitor their thermals and down-throttle (even below way specified minimum speeds) as needed to protect themselves. Not that I'd recommend relying on that but still, I have had cooling fail without apparent damage.
How do you know they weren't faulty? Blip flips are usually silent unless you have ECC memory or are actively testing for faulty cells.
Somewhat recently, Firefox implemented memory testing in their crash reporter and found 10% of reports were due to bit-flips[0].
If it weren't for reading that HN thread, I would never have being prompted to check and find the faulty memory in my laptop. It would have gone undetected and continued to corrupt whatever was stored in the affected cells. I did run memtest86 not long after I bought the laptop, so sometime between running that first memtest86 and the 12 months after, it developed the fault.
Non-ECC memory sometimes fails and there's a good chance you'll never know.
Most devices never test their memory and don’t have ecc, so you generally never know. Random instability can be fixed by restarting and your program not being in an area with bad bits anymore.
I recently had part of a 32 gig DIMM go bad on a 128 gig system. One of the larger VMs was spontaneously rebooting for no apparent reason. I was able to find the addresses with MemTest86 and map out 512 megs of the affected DIMM with the Linux kernel memmap option.
Don't send it back. Since they can't easily replace it with new hardware, they will probably refund you the price of purchase. Which is normally a very good deal, but obviously not in these circumstances. I've seen it happen with GPUs as well.
I've seen every single component of a system fail, the least I've seen fail has been the CPU itself (I've only seen one fail out of thousands).
Memory seems to fail pretty commonly overall, less common than harddisks but more common than motherboards.
The issue is that I'm used to using ECC ram, which fails loud when it's actually bad.. consumer memory won't tell you unless you can't boot.. and everyone disables the startup memory testing too.. (in fact, I think it's disabled by default for the last 10 years because people want to boot quickly).
I was so sure that a CPU wouldn't suddenly fail that I ended up replacing every other part in my computer before getting a new CPU and realizing that was the issue.
I had a 2012 mac book pro that burned out 2 different crucial memory sticks. Would have bricked the machine if it was the soldered-on type, but it was the last model apple sold that have swappable memory. Crucial had a lifetime warranty so they replaced them for free.
I used to game on it and I think the design just couldn't handle the cpu and dgpu being active for long periods, it would get extremely hot in the area close to where the memory was.
Granted, I've been building my own computer for 30 years and I've only seen it happen ONCE.
But I've seen plenty of other failures:
- Two hard drives -- One got dropped on the floor, so it was no surprise, the other was showing degraded performance and SMART showed some scary numbers and I was able to replace it before I lost data.
- One GPU, replaced under warranty.
- One AIO water cooler, replaced under warranty. Somehow, the water vanished from the loop. I'm guessing a microscopic leak that evaporated as fast as it leaked.
- Two power supplies -- One randomly exploded. Just started making popping noises and shooting sparks out the back. At autopsy, I determined that the fan failed (it was hard to spin manually) and it overheated. The other was just being overloaded. I had just gotten a new GPU and the PSU wasn't big enough for it.
- Several case fans.
Never seen a CPU or motherboard failure, even when my AIO water cooler was failing and my CPU was constantly at thermal limits. Never had an SSD/NVMe failure.
Anecdata: a coworker found a single bit flip corruption in a file he wrote in his computer. He ran memtest86 on his memory, and which found one of the two sticks was defective. The computer otherwise ran perfectly fine for years; and we would have never noticed if the file format didn't have built-in checksums.
> I've been using computers for more than 40 years now and the one component I have never seen fail is RAM.
I think over the decades I got one RAM disk fail. And another one didn't fail but, although I bought a pack of four sticks (2x x2), was detected (by memtest, when building the rig), as having another model name than the three others: even the shop who sold me the RAM was confused (don't know how that happened).
The usual component that I've seen fail the most would be the PSU? (before I started buying quality ones). One NVMe drive (an "ADATA") died on me even though it was nearly new.
PC assembled with quality parts usually last a very long time.
I also have been using PC's for a long time, thirty years. The only time I have seen ram fail during use was within the last 5 years. My 10th gen platform and 32 gigs of ram DDR4 and one of the sticks failed, this was before the ramapocalypse of course, the sticks were replaced under warranty. But I have had systems in my shop with failed ram. it happens, but its rare, unless its cheap ram, less rare.
I have had a RAM stick start frequently failing ECC checks after a couple years use. I also had the VRAM on a GPU fail, but you could blame that one on overheating from the nearby compute core. I agree is rarer though.
Oh well, it's fairly common.Maybe not in a singular/several rarely used (and not monitored) desktop computers, but.... it's common once you have quite a few under your care.
on older PC bios would perform a memory test - you may remember those - and that's where it would mark a bunch of pages dead/unusuable and it was not uncommon at all.
I've seen some fail. Might have had to do with the fact that the computers were literally flooded (river came up faster than the sandbaggers could keep up) while running.
I've seen computers brought back to life after flooding by lying in rice for a week, but they weren't turned on when the water hit (and not all survived).
I've had a few RAM sticks fail, but the ECC usually caught it first. Also had one case that somehow the stick slightly unseated itself; dunno how but it was out of service for months until someone opened it up and reseated the stick.
...granted though I've also seen a backpane on a server fail. I didn't think that was even possible until that point.
All the other bits in a computer I've seen fail.