When one update breaks everything

Two years ago, a faulty CrowdStrike update caused widespread disruption across Windows systems worldwide.

Airports, businesses and critical services suddenly found themselves dealing with machines that could no longer operate normally.

What made the incident interesting wasn’t just the scale of the outage. It showed how dependent modern organisations have become on software that sits quietly underneath everything else. One small component can affect hundreds of other systems without most users even knowing it exists.

We often talk about redundancy, backups and disaster recovery as if they’re secondary technical tasks. Then something fails, and those “boring” precautions suddenly become the most important part of the entire infrastructure.

Technology is becoming more interconnected every year, which gives us incredible capabilities but also creates increasingly complex dependencies.

I don’t think good IT is about building systems that never fail. That’s unrealistic. Good IT is about making sure that when something does fail, the organisation can recover quickly, understand what happened and keep operating. Resilience is probably one of the least exciting parts of technology, but also one of the most important.

Alessio

Technology consultant, software builder and problem solver, sharing practical thoughts on tech, operations and digital products.