Rarely does an IT outage make national headlines, and even more rarely does that outage bring good news. The CrowdStrike error, which impacted millions, brought critical services like Microsoft 365, Azure, and countless others to a grinding halt. The culprit? A seemingly innocuous software update that was anything but routine.
While there will be no shortage of second-guessing, we know that this could happen even to the most well-run organizations. This underscores the importance of robust IT procedures, backup, disaster recovery plans, and maintaining control over your environment to avoid being at the mercy of another company. If there were ever a time to address adherence to protocols, it would be now when this is fresh in everyone’s mind.
One Small Programmer Error One Large Global Impact
The outage’s root cause stemmed from a human error, a bug in the code written by a CrowdStrike developer. You can find a full breakdown of the technical specifics here from CrowdStrike.
Essentially, a mistake happened while updating CrowdStrike’s software, which was intended to improve communication between Windows programs. The update asked for 21 pieces of information, but the code only provided 20. This error wasn’t noticed during testing because the tests allowed a “wildcard” to fill in the missing piece, so everything seemed fine at first.
Later, when the system tried to find the missing piece of info, it looked in the wrong place, causing the system to crash for safety reasons, which is a very simple explanation of a very complex issue.
Mitigation Efforts and the Road to Recovery
While the cause may seem like a simple mistake, the impact was far-reaching. Thankfully, both CrowdStrike and Microsoft responded swiftly:
- CrowdStrike: Acknowledged the issue and released a public statement along with a workaround solution.
- Microsoft: Communicated with CrowdStrike and external developers to expedite a solution. They also provided technical guidance and support to help customers recover safely.
The fix involved a solution from CrowdStrike that addressed the null pointer issue and prevented further crashes. Additionally, Microsoft posted instructions on the Windows Message Center to guide users on how to remedy the situation on their Windows endpoints.
A Global Impact
The outage wasn’t localized; it affected users across the globe. Critical business operations, healthcare services, airlines, stock exchanges, and countless individuals across various countries were impacted.
The level of companies affected was staggering. From airlines to healthcare, the ripple effect was profound, emphasizing the need for robust resilient infrastructure and business continuity plans.
– Dan Frasco, VP of Cyber Security for Amplix
While the exact number of affected users and locations remains unclear, reports suggest the outage spanned continents, causing significant disruptions.
Learning from the Outage: Ensuring Infrastructure Resilience
The CrowdStrike-Microsoft outage serves as a reminder of the domino effect a seemingly minor software bug can have. It highlights the importance of rigorous code reviews and collaboration’s crucial role in mitigating widespread disruptions.
Best Practices for Deployment and Testing
- Rigorous Code Reviews: Ensuring thorough code reviews and testing can prevent such errors from slipping through. Dan Frasco pointed out, “Proper validation and testing are non-negotiable. N-minus-1 testing strategies can help in identifying issues before they hit production environments.”
- Staggered Deployments: Instead of rolling out updates across all systems simultaneously, a phased deployment can help isolate and mitigate issues without widespread impact. “At Amplix, we always recommend deploying to a subset of systems first, monitoring the impact, and then proceeding with a full rollout,” Frasco advised.
- Rollback Procedures: Having a robust rollback plan in place can minimize downtime. If an update causes issues, the ability to quickly revert to the previous stable version is crucial.
- Disaster Recovery and Business Continuity Plans: Every organization needs to have a tested comprehensive disaster recovery plan and business continuity plans. This includes regular backups, redundant and resilient systems, and a clear action plan for restoring services in case of an outage.
The Importance of Having Your Own Network Control
One of the key takeaways from this incident is the risk of dependency on third-party services. While leveraging cloud services and external solutions can offer significant advantages, it’s crucial to maintain control over your network.
It’s not just about responding to outages; it’s about being prepared and having a plan. Ensure that your network and infrastructure are truly yours, not at the mercy of another company.
– Dan Frasco, VP of Cyber Security for Amplix
Strategies to Maintain Control
- Internal Testing and Validation: Even if a third-party vendor assures the quality of their updates, organizations should conduct their own testing. This ensures compatibility with their specific environment and prevents unexpected issues.
- Configurable Update Management: Dan Frasco emphasized the importance of having control over update schedules. “Organizations should have the flexibility to delay updates and conduct thorough testing. Automatic updates can be risky if not properly managed.”
- Comprehensive Monitoring and Alert Systems: Implementing robust monitoring tools can help detect issues early and initiate immediate responses.
- Backup and Redundancy: Regular backups and redundant systems can ensure business continuity. In the event of an outage, these measures can facilitate a swift recovery.
Business Is Better With Partners You Can Trust
The CrowdStrike outage was a good reminder to all of us not just in the IT world, but the entire business community of the importance of creating IT protocol and following that protocol. The old carpenter adage of “measure twice, cut once” applies, as well as having a routine backup of all your critical components in place to spin up can mean the difference between being out for minutes versus hours, days, and even weeks.
This is what we do at Amplix, we are not just another IT vendor, we are a partner who helps companies create strategies and implement those strategies to leverage IT as a competitive advantage. Contact us today to see how we can help your company achieve its goals.