Your AI Problem Might Actually Be a Data Problem

Picture of Chris Boshoff
Chris Boshoff
Chris Boshoff is the Managing Director of GoRebel Artificial Intelligence, with more than 25 years of experience in Microsoft technologies, cloud architecture and enterprise IT. He specialises in AI strategy, Azure, Entra ID, data engineering and modern cloud solutions, with a strong focus on aligning technology with real business outcomes. His background spans development, technical leadership and CTO-level roles, giving him a practical perspective on building secure, scalable and intelligent solutions.

Businesses often blame disappointing AI results on the model. But fragmented systems, unreliable information and unclear access controls are frequently the real constraint. Before investing in better AI, it’s worth looking closely at the data underneath it.

The proof of concept looked excellent.

We’d given the AI a carefully chosen set of documents, clean information and a controlled list of questions to answer. It performed beautifully. Everyone in the room was convinced.

Then we connected it to the real organisation and things changed.

Suddenly there were six versions of the same policy. Product names differed between systems. Information was missing. Permissions were inconsistent. Half the documents were years out of date, and nobody could say with confidence which source was the authoritative one.

The model hadn’t got worse between the demo and the deployment. It had simply met the business’s real data.

I’ve watched this play out enough times to say it plainly: when an AI initiative disappoints, the instinct is to blame the model. Often, the model isn’t the main problem.

AI doesn’t remove the weaknesses in your business data. It exposes them and sometimes amplifies them.

AI makes existing data problems visible

Most organisations have lived with fragmented data for years. What’s easy to miss is how well people have learned to compensate for it.

Your team knows which spreadsheet is actually current. They know which folder to check, who to phone when the system is wrong, which field in the ERP can’t be trusted, which SOP is obsolete, and which of two conflicting records is the real one.

None of that knowledge is written down. It lives in people’s heads, and it quietly holds the whole thing together.

AI has none of that institutional intuition, not unless you deliberately build it in. Point a model at your systems and it will work with the information it can access, conflicts and all.

That’s why AI projects are so good at surfacing weaknesses a business has spent years learning to tolerate.

Humans are remarkably good at working around bad information. AI is much less forgiving.

Having lots of data is not the same as having AI-ready data

Here’s a distinction I find myself explaining in almost every early conversation: the fact that data exists tells you very little about whether AI can use it effectively.

A company can hold terabytes of information and still have very little an AI system can work with reliably.

“We have the data” and “our data is usable, accessible, trustworthy and governed” are two completely different statements – and the gap between them is where many projects get into trouble.

In plain terms, data is AI-ready when:

  • the right information exists
  • it can be accessed
  • it’s accurate enough to rely on
  • the organisation knows which source to trust when there’s a conflict
  • permissions are clear
  • sensitive information is protected
  • the systems can surface the information when it’s needed

Miss any one of those, and the model inherits the gap.

Five data problems that quietly derail AI

In practice, the same handful of issues come up again and again.

1. The data is scattered everywhere

Across ERP, CRM, SharePoint, Google Drive, document stores, email, spreadsheets and legacy systems that predate half the team.

AI does not automatically turn all of that into a reliable source of truth. Someone still has to decide how the information connects, which sources matter and how conflicts should be handled.

2. The organisation doesn’t know what to trust

Multiple versions of the same document. Old records sitting alongside new ones. Duplicates. Conflicting entries with no clear indication of which source wins.

An AI system can retrieve what is there. That does not necessarily mean what is there is correct.

3. The data was designed for systems, not for AI

The information technically exists, but it’s often stored in structures built for a transactional system to process, not for a model to retrieve with the right context attached.

Making data usable by AI often means thinking differently about how that information is exposed, connected and interpreted.

4. Access controls are unclear

This one matters more than many leaders expect.

Just because an AI system can retrieve a piece of information doesn’t mean every user asking should be allowed to see it.

Getting this wrong doesn’t just risk a bad answer, it can create a serious security or compliance problem.

5. Nobody really owns the data

Who’s responsible for its accuracy? For maintaining it, setting permissions, removing what’s out of date and defining the authoritative source?

When the answer is “no one in particular”, quality only ever moves in one direction.

What this looks like in the real world

Two examples make the pattern concrete.

Take a manufacturer.

The information needed to answer a genuinely useful operational question may be spread across the ERP, MES, WMS, maintenance systems, quality records, machine data and a library of SOPs.

Each system is doing its job. The problem is that no single one holds the full operational context – the thread connecting a machine fault to the relevant maintenance history, the correct procedure and the quality record.

Manufacturers rarely have too little data. They have enormous volumes of it, trapped in systems that were never designed to work together in this way.

AI doesn’t remove that disconnection. It runs straight into it.

Now take FinTech.

Here, connecting AI to the data is only one part of the problem. Customer, transaction, risk, compliance and policy information can each carry different access requirements.

The goal was never to make everything searchable by everyone.

It’s to retrieve the right information for the right user under the right controls without crossing the boundaries around sensitive data.

In a regulated environment, a data foundation that ignores permissions isn’t a foundation at all.

The same principle sits underneath health, education, professional services and any knowledge-heavy operation: the organisation usually holds the answer somewhere.

The work is making it findable, trustworthy and appropriately governed.

A better model doesn’t fix a bad foundation

The common reaction to a disappointing result is:

“Maybe we need a more powerful model.”

Model capability does matter but there’s a hard limit to what it solves.

If the source information is wrong, a smarter model gives you a more articulate wrong answer.

If the right information is inaccessible, more capability doesn’t make it reachable.

If two sources conflict, a bigger model doesn’t automatically know which one the business considers authoritative.

If permissions are incorrect, a better model can simply enforce the mistake more efficiently.

If the AI is asking the wrong data for the right answer, a smarter model doesn’t solve the problem.

This is why, at GoRebel, we treat the model as one component of a larger system and why we built RebelCore around governed access to your data rather than around the model alone.

The intelligence is only as trustworthy as the foundation it sits on.

Five things to do before connecting AI to your business data

None of this requires a year-long data programme before you see any value.

It requires a sensible order of operations.

1. Identify the business question

What do you actually want AI to help people do?

2. Identify the information required

Which systems, documents and sources contain the information needed to answer that question?

3. Establish the authoritative source

When two sources disagree, which one should the system trust?

4. Define who can access what

Which users should be allowed to retrieve which information, and under what conditions?

5. Test against real data

Not a curated demo environment.

Use the messy, contradictory and imperfect information the production system will genuinely encounter.

That last step is one of the most important. It helps separate something that looks impressive in a meeting from something the business can actually rely on.

You don’t need to fix all your data first

Let me head off the obvious misreading, because it’s the one that makes businesses freeze.

“Data first” does not mean:

“Clean the entire organisation before you do any AI.”

That’s a recipe for never starting.

It means: pin down the specific business problem, work out which data it depends on, and get that data ready enough to support the solution.

If you’re building an internal HR knowledge assistant, you don’t need to repair the manufacturing database first.

You need the HR policies, employment documentation, procedures, permissions and clear ownership in good shape.

That’s it.

The scope of the data work is set by the scope of the problem which is exactly what keeps this practical rather than paralysing.

Before you change the AI, look underneath it

When an AI project disappoints, resist the reflex to immediately change platforms, models or vendors.

Look underneath the AI first.

Is the information available?

Can it be trusted?

Can the system access it?

And should the person asking the question be allowed to see the answer?

Get those things right and AI becomes dramatically more useful, often with the model you already have.

The smartest AI strategy usually doesn’t start with smarter AI. It starts with the data you already own.


Is your business data ready for AI?

Talk to GoRebel about how your existing systems, data and company knowledge can be prepared for secure, practical AI use.

Talk to GoRebel

Or start broader: Assess Your AI Readiness  · Related: “Your Business Doesn’t Need an AI Strategy

LinkedIn
Facebook
X
WhatsApp
Email

Related

AI Data

Your AI Problem Might Actually Be a Data Problem

Businesses often blame disappointing AI results on the model. But fragmented systems, unreliable information and unclear access controls are frequently the real constraint. Before investing in better AI, it’s worth looking closely at the data underneath it.