Skip to content
All writing

Why Not, AI on Your Protected Data?

A customer asked me a very important question last week that I have been pondering upon ever since.

We discussed about the AI, how it can help them in what they wanted to do, what was getting in the way and out of nowhere I was asked

"Why would I even do AI on protected data?"

I kept thinking on it, and the more I thought about it, the more I realised the question had it backwards.

Why not?

Here's what I've noticed talking to engineers and founders building AI products right now. The conversation almost always ends up in the same place. What's blocking them is that their data is everywhere and nowhere at the same time - scattered across on-prem systems, cloud storage, SaaS tools, file servers and how can we firstly connect, consolidate this data?

Everyone agrees on this, the single bottleneck for AI in enterprise use is data access.

So when someone asks why you'd run AI on protected data, I think the implicit assumption is that backup data is somehow second-class. A copy of a copy and not the real thing.

But that assumption deserves a harder look. Here's why.

Your protected data estate is probably the most complete view of your organisation that exists anywhere. It has depth i.e, months or years of history. It has breadth which spans sources that your production systems don't talk to each other about. Silos is already consolidated and already governed.

Are you starting from scratch to connect your data? Why would we? You can run AI on the data which is already consolidated, governed, which saves you cost getting data ready for your data(Lets speak about it in different article)

##What about latency? I understand, intuitively it feels like querying a backup platform should be slower than hitting a live production database. It maybe marginally slower, but ms (milliseconds) more? We are not talking about minutes, or even seconds at Cohesity.

Here's again everyone agrees, in an AI workload, data retrieval is not where the time goes. Then where?

When an AI answers a question, finding the data takes a fraction of a second, while the LLM's thinking and writing takes seconds or minutes. Worrying about how fast the data is found is like stressing over the 10 seconds it takes to grab coffee beans from the cupboard, when the coffee itself takes five minutes to brew.

Which one takes time here? Retrieval or brewing?

and last but not the least

##What about data protection? This is even more critical.

If your AI is built on top of unprotected data infrastructure, how would you recover the data if you haven't even protected?

With Cohesity, Your data is already protected and You can restore to a point in time.

Recovering your data from an agent deleting the data?

I genuinely haven't seen another company approach this the way Cohesity does. Not because others aren't trying but because most people are still treating data protection and AI data infrastructure as completely separate problems. The insight that your backup estate is your AI data estate hasn't landed broadly yet.

I think it will.

So back to the customer's question.

"Why would I do AI on protected data?"

Because it's not protected data or AI data. It's just data, and what does AI need? Data. And you already have it, consolidated, governed, and recoverable, sitting in a platform that can now let you actually use it.

The question isn't whether this makes sense. The question is why you'd build something parallel when what you need already exists.

More writingGet in touch