>Impact Statement: Starting at 18:44 UTC on 26 Dec 2024, you have been identified as a customer who was impacted by a power incident in South Central US and may experience a degraded experience.
>Current Status: There was a power incident in the South Central US AZ03 which affected multiple services. We have applied mitigation and are actively validating recovery to the impacted services. Further updates will be provided in 60 minutes, or sooner as events warrant.
The times are the same for OpenAI - first notice from 11:00 PST (19:00 UTC)
I don't think this stuff is the work of the devil personally by a long shot.
We don't know what exactly caused the power issue and they might not have had a root cause at the time either. Let's assume that their power redundancy equipment failed, say, due to insufficient maintenance. This is not an active action, it's a passive one (they didn't do their maintenance duties properly and now it blew). So there is nothing to say for point #1 and #2.
There's also the part where they say that the customers they identified as impacted may be experiencing a service degradation. This may sound pedantic, but I think it is not an entirely unreasonable phrasing. Maybe my business isn't actively relying on the resources I have deployed in that datacenter. How would they know (#3)? Should I clean those resources up? Possibly. Depends on my access patterns and other considerations.
It reads like face (and ass) saving legal esque language. But there's a reason face and ass saving legalese sounds like it does.
It’s the same accountability shirking language as when layoffs “have affected you” instead of “I mismanaged this business and as a result I’m firing you”.
It’s always the same abstract invisible hand that just keeps affecting everyone! Scott Alexander’s Moloch perhaps :)
Power outage seems really odd. Don't datacenters usually have multiple redundant power supplies + on-site backup power generation? Maybe power "incident" mean something else?
The switching equipment can fail. Had this happen at a DC where the switching equipment arc flashed when going from mains to diesel generators. The switching equipment detected the arc and then locked out until someone onsite could inspect the equipment and override it. The rack UPSes only lasted like 5 minutes and then everything went dark.
this happened to a dupont fabros facility in northern va in I wanna say ~2012?
derecho storms hammered the area and killed power. external power lines in failed, and the ATS hung or died when switching to the N+1 diesel generators.
since it never got switched to diesel, the UPS systems kept things going for the standard interval (e.g. ~3-5 minutes) and then ran out of power, and then everything went down. AWS died and IIRC it took a lot of stuff with it, most notably reddit, etc.
I had equipment at a colo facility -- that had a (licensed, bonded, not fly-by-night) electrical tech accidentally drop a tool into the main bus connecting mains, generators, and batteries.
They are lucky they were able to walk away, but the facility was dark till someone could get in there and give the power equipment the green light.
Back in the 00's there was a power outage in downtown Vancouver. It caused Peer1 to fall back to their generators... that weren't tested for ages. They struggled for 5 minutes, gave up and bursted into flames, resulting in the colo not being able to go back online even when the main power was restored.
That was an epic mess. Especially considering they positioned themselves as the most technically sophisticated colo in the region. So, yeah, it happens.
Things happen, e.g. the redundant system also failing, the system that should handle the failover failing, a short circuit that causes enough chaos that the redundant supply shuts down rather than feeding power into a potential fault, ...
Tl;Dr fire in one data center hall was put out with water, water leaked into other hall's power generator and battery area. Turns out loads of water and power generation equipment don't mix well, and servers don't like sitting in puddles of water.
>Impact Statement: Starting at 18:44 UTC on 26 Dec 2024, you have been identified as a customer who was impacted by a power incident in South Central US and may experience a degraded experience.
>Current Status: There was a power incident in the South Central US AZ03 which affected multiple services. We have applied mitigation and are actively validating recovery to the impacted services. Further updates will be provided in 60 minutes, or sooner as events warrant.
The times are the same for OpenAI - first notice from 11:00 PST (19:00 UTC)