Back

[Analysis]

CrowdStrike Falcon outage: a faulty content update deployed to every host

In July 2024 a faulty configuration update, not an attacker, crashed 8.5 million Windows machines. It's the clearest lesson we know in testing the edge case, treating configuration as code, and never shipping to everyone at once.

A large screen hanging over the B gates at Washington Dulles airport, showing a Windows crash screen above the signs for the toilets and gates.
A display over the B gates at Washington Dulles International Airport on the morning of 19 July 2024. Crash screens like this one, in airports, shops and stations, were the most visible sign of the outage. Cropped from the original. Photo by reivax, CC BY-SA 2.0.

At 04:09 UTC on 19 July 2024, CrowdStrike released a routine content configuration update for its Falcon sensor on Windows. It was reverted 78 minutes later, at 05:27. In between, it reached every online Windows host running an affected version of the sensor, and crashed them. Microsoft estimated that 8.5 million devices were affected. Airlines, broadcasters, banks and GP surgeries were among those disrupted, and many machines had to be recovered by hand.

The NCSC was quick to assess that it was not a security incident or malicious cyber activity. It told affected organisations to follow the vendor’s guidance, and warned of a rise in phishing emails referencing the outage, as criminals tried to take advantage of the confusion.

Nobody attacked anyone. But this is still a security story, and one of the best we know about why software assurance matters.

What went wrong

CrowdStrike’s root cause analysis is worth reading in full. In short:

[All at once, or a ring at a time]

19 July 2024

  1. Update
  2. Every Windows host with the sensor online between 04:09 and 05:27 UTC

Staged rollout

  1. Update
  2. Internal
  3. Canary
  4. Early
  5. Everyone else

Each ring is a checkpoint: if something breaks, stop before the next one.

CrowdStrike’s fixes included automated tests for every template type, validating the number of input fields, bounds checking in the interpreter, a deployment ring process so content passes successive stages before full rollout, more customer control over when updates arrive, and independent review of its code and release processes.

What this means for anyone who ships software

You don’t need to build endpoint security products for these to apply:

The UK’s Software Security Code of Practice asks vendors to have a clear process for testing software and software updates before distribution. After July 2024, that line reads rather differently.

And if you’re the customer

Ask your critical suppliers how they roll out updates, and whether you can control the timing. For most software, taking updates promptly is still the right call: the NCSC was clear, even in the middle of the outage, that installing security updates is still essential. But “promptly” and “all at once, with no way back” aren’t the same thing.


Sources

Back