Welcome to Roya News, stay informed with the most important news at your fingertips.

CrowdStrike protection.

1
Image 1 from gallery

How a faulty CrowdStrike update crippled global IT systems

Published :  
22/7/2024 15:28|
|
Editor Name:  
Mohammad_Alakaileh

By: Muhammed Akayleh

In a dramatic turn of events that mirrored some of the worst cyber disasters in history, a flawed update from cybersecurity giant CrowdStrike sent computers worldwide into a catastrophic reboot loop. This digital calamity, unfolding over the past 12 hours, disrupted vital sectors including air travel, healthcare, banking, and more, underscoring the fragility of our interconnected digital infrastructure.

The chaos began with a widespread outage on Microsoft’s Azure cloud platform on Thursday night. By Friday morning, the situation worsened as CrowdStrike released a defective update for its Falcon monitoring product, an advanced antivirus platform that operates with deep system access on various endpoints like laptops, servers, and routers. This combination of events created a perfect storm, bringing down IT systems globally. According to Wired, a Microsoft spokesperson clarified that the Azure outage and the CrowdStrike issue were unrelated.

CrowdStrike CEO George Kurtz identified the cause as a defect in the code released for Windows, which did not affect Mac and Linux systems. The problem stemmed from a configuration file update aimed at improving Falcon’s inspection of "named pipes" in Windows, a feature that facilitates data transfer between processes on the same machine or across a network. This update, intended to enhance detection of new hacking methods, inadvertently triggered a logic error causing operating systems to crash.

The update’s faulty configuration file, which ended in the .sys extension commonly used by kernel drivers, altered the functionality of the driver, leading to widespread system crashes. Despite the .sys extension, CrowdStrike clarified that the problematic file was not a kernel driver, though it interacted with one in a way that caused the crash. Kurtz added in a statement, "The issue has been identified, isolated, and a fix has been deployed," and he apologized for the disruption, noting that it may take some time for things to return to normal.

The impact was immediate and extensive. Hospitals in the UK, Israel, and Germany faced disruptions, with some canceling appointments due to communication breakdowns. Emergency services in the US reported issues with 911 lines. Television stations like Sky News experienced interruptions in live broadcasts. Airports worldwide, including those in India and the US, saw significant delays and grounded flights, with some resorting to handwritten boarding passes.

Cybersecurity authorities globally, including the UK's National Cyber Security Center and Australian officials, quickly ruled out malicious cyber attacks, attributing the issue to human error in the update process. The consensus was that the fault lay in the update mechanism itself rather than an external attack. "The NCSC assesses that these have not been caused by malicious cyber attacks," stated Felicity Oswald, CEO of the UK’s National Cyber Security Center.

Security experts highlighted the inherent risks of deep-system access required by advanced antivirus software. Matthieu Suiche, head of detection engineering at Magnet Forensics, likened running malicious code detection at the kernel level to "open-heart surgery," emphasizing the potential for such software to cause system-wide failures if mishandled.

Costin Raiu, a former Kaspersky expert, noted that driver updates undergo rigorous testing due to their critical role in system stability. However, the configuration file in question appeared to have bypassed this level of scrutiny, leading to the massive outage. Raiu stressed the need for careful oversight and testing of all updates to prevent such incidents. "One simple driver can bring down everything. Which is what we saw here," Raiu remarked.

CrowdStrike's immediate response involved isolating the issue and deploying a fix. However, the nature of the problem meant that many affected systems required manual intervention to recover. The company provided a workaround, instructing users to boot Windows machines in safe mode, delete a specific file, and reboot. This process was time-consuming, as millions of machines globally needed individual attention. "The fixes we’ve seen so far mean that you have to physically go to every machine, which will take days," said Mikko Hyppönen, the chief research officer at cybersecurity company WithSecure.

According to Wired, the CrowdStrike incident serves as a stark reminder of the vulnerabilities inherent in our digital infrastructure. Ciaran Martin, a professor at the University of Oxford and former head of the UK’s National Cyber Security Center, emphasized the lesson this event teaches about the fragility of core internet systems and the need for robust safeguards against both human error and malicious exploitation.

Jake Williams, vice president of research and development at cybersecurity consultancy Hunter Strategy, pointed out the potential for significant changes in how updates are managed. He argued that the CrowdStrike debacle might lead to demands for more controlled and supervised update processes, reducing the risk of such widespread failures in the future. "For better or worse, CrowdStrike has just shown why pushing updates without IT intervention is unsustainable," Williams stated.

The CrowdStrike update catastrophe highlights the delicate balance between security and stability in our digital world. As the global community recovers from this unprecedented outage, the lessons learned will hopefully lead to more resilient systems and safer update practices, ensuring that the tools designed to protect us do not become the cause of our downfall.