[Analysis]
CrowdStrike Falcon outage: a faulty content update deployed to every host
In July 2024 a faulty configuration update, not an attacker, crashed 8.5 million Windows machines. It's the clearest lesson we know in testing the edge case, treating configuration as code, and never shipping to everyone at once.

At 04:09 UTC on 19 July 2024, CrowdStrike released a routine content configuration update for its Falcon sensor on Windows. It was reverted 78 minutes later, at 05:27. In between, it reached every online Windows host running an affected version of the sensor, and crashed them. Microsoft estimated that 8.5 million devices were affected. Airlines, broadcasters, banks and GP surgeries were among those disrupted, and many machines had to be recovered by hand.
The NCSC was quick to assess that it was not a security incident or malicious cyber activity. It told affected organisations to follow the vendor’s guidance, and warned of a rise in phishing emails referencing the outage, as criminals tried to take advantage of the confusion.
Nobody attacked anyone. But this is still a security story, and one of the best we know about why software assurance matters.
What went wrong
CrowdStrike’s root cause analysis is worth reading in full. In short:
- A sensor capability added in February 2024 defined 21 input fields for a particular kind of content. The code that fed it supplied only 20.
- Earlier content updates for that capability had worked, partly because tests used a wildcard for the 21st field, so the mismatch never showed.
- The July update was the first to use the 21st field for real. A logic error in the content validator let it through, and the sensor read beyond the end of its input and crashed.
- The update was deployed to all eligible hosts at once, so the first place it failed was everywhere.
[All at once, or a ring at a time]
19 July 2024
- Update
- Every Windows host with the sensor online between 04:09 and 05:27 UTC
Staged rollout
- Update
- Internal
- Canary
- Early
- Everyone else
Each ring is a checkpoint: if something breaks, stop before the next one.
CrowdStrike’s fixes included automated tests for every template type, validating the number of input fields, bounds checking in the interpreter, a deployment ring process so content passes successive stages before full rollout, more customer control over when updates arrive, and independent review of its code and release processes.
What this means for anyone who ships software
You don’t need to build endpoint security products for these to apply:
- Configuration is code. Feature flags, rules files, templates and infrastructure definitions can take a system down as surely as a code change. They belong in version control, in review and under test.
- Test the edge case you’re relying on. A wildcard in a test says “anything goes here”. That’s rarely what production does. Where code indexes into an array or parses input, the test that matters is the one with the unexpected length.
- Validators need tests too. A check that’s meant to catch bad input, and quietly doesn’t, is worse than no check at all, because everyone downstream trusts it.
- Roll out in rings. Send a change to yourself first, then a small group, then more, and watch for trouble at each stage. It costs a little speed and turns a global outage into a contained one.
- Plan to recover from “good” software failing. Many organisations found that recovering machines by hand meant finding BitLocker recovery keys that were stored on the very systems that were down. Know where yours are.
The UK’s Software Security Code of Practice asks vendors to have a clear process for testing software and software updates before distribution. After July 2024, that line reads rather differently.
And if you’re the customer
Ask your critical suppliers how they roll out updates, and whether you can control the timing. For most software, taking updates promptly is still the right call: the NCSC was clear, even in the middle of the outage, that installing security updates is still essential. But “promptly” and “all at once, with no way back” aren’t the same thing.
Sources
- CrowdStrike, Preliminary post-incident review (July 2024)
- CrowdStrike, Executive summary: root cause analysis, Channel File 291 (August 2024)
- Microsoft, Helping our customers through the CrowdStrike outage (July 2024)
- NCSC, Statement on major IT outage and phishing threat (July 2024)
- DSIT and NCSC, Software Security Code of Practice (May 2025)