One of the biggest challenges that organizations face in a cyber incident is determining whether the data they plan to restore is valid or already compromised.
Hopefully you’ve got a good start on a path to data resilience which includes creating immutable copies, known as Safeguarded Copies on an IBM FlashSystem. The next question I get asked is about data validation!
“How do I test that Safeguarded Copy to make sure it’s usable for recovery?”
That’s the data validation question if you ask me. The next hurdle.
That usually sparks a debate, do we do that testing reactively or should we be proactive? Personally, I think that’s a journey, just like getting in the habit of taking Safeguarded Copies.
Starting out, most teams I see, build a process and automation around being able to reactively test their primary storage immutable copies ’IF’ they need to. This allows them to at least plan out the process and get a feel for how things would work in testing. That’s great!!!!
Then, data validation, that often starts out as human inspection, which is something I think we can all identify with. And that part of the process doesn’t really ever go away, somewhere along the line, I think we all expect a human will want to have a test or a ‘look’ at what is about to be restored.
Next, the question becomes, how do I do this testing
a) proactively or in an automated fashion where the human doesn’t have to be involved,
b) how can I use tooling to validate the data for me and report the outcome to the humans and ideally the SOC.
As an aside, this is also where we see a lot of organizations start building a cross-team integration between the Storage teams and the Security teams. It goes back to that fundamental idea that the two can work hand in hand; one on cyber resiliency and one on cyber security.
Now, to the testing question; one of the advantages of what we can do with IBM FlashSystem, is the integration with IBM Defender Sentinel. Wait, what is IBM Defender Sentinel?
IBM Storage Defender Sentinel is workload-specific software, powered by IndexEngines, that detects, diagnoses and identifies the sources of ransomware attacks and provides automated recovery orchestration for Oracle, SAP HANA, Linux, VMware and the Epic healthcare system.
It does just what is says on the ‘tin’ as the expression goes. It lets me automate the creation, lifespan, validation and tracking of the IBM Safeguarded Copies for my FlashSystem. Now this is where things get really interesting, so using Sentinel I can set up a schedule that will not only take the Safeguarded Copies for me, but it will also automate the testing of the copies and alert me if it finds a corruption, but that’s where things get different.
IBM Storage Sentinel is an AI powered tool that leverages the read performance of the IBM FlashCore modules to be able to test the Safeguarded Copy right on the FlashSystem, which means I don’t have to move the data. It already includes the automation to take application consistent copies and orchestration to scan the copies and alert on results. The AI component of it is two-fold!
First, it’s trained against malware in a ‘dirty room’ that lets the AI training identify how malware behaves so it’s able to identify, based on 200+ different points, how the data is supposed to look and what a potentially corrupted workload might look like based on how the malware creates the corruption.
In principal, what does that mean; there are several different types of ransomware / malware some of it maintains original metadata, but then corrupts the data itself, other will do partial or intermittent encryption, some will corrupt data but the changes will maintain the same size so it doesn't change the size or data patterns.
This means different types of corruption attacks behave in unique ways, so knowing how data is corrupted and how a malware attack behaves is an advantage in detection.
Now, the other part of the AI assist is IBM Defender Sentinel is trained against specific workload types so it knows what the construct of the data is that its inspecting across the 200+ points and can see if there is an issue. I think this really interesting from an intermittent corruption type of impact, where only parts of the data are impacted.
What does it look like? For me the real impact is the information I can get if the data has been corrupted. I’m able to look in one place that will give me the files that are impacted, how, where, the mix, as well as an understanding that impact. Sort of like this:
#community-stories1