The AI Data Readiness Conversation That Needs to Happen Now

Summary

In a recent HBR Analytic Services survey, only seven percent of respondents called their data AI-ready. Before agents can act on their own, they need more than just clean data. They need a structure for unstructured data so they don’t keep reinterpreting relationships and meanings with every query. Sindhu Venkata explains why, and offers three questions to pose to your organization to test readiness.

[Estimated read time: 4 minutes]

That data’s not going to clean itself

Apparently, the fastest way to make enterprise data AI-ready is to buy an AI tool and hope the data gets intimidated into cleaning itself.

Unfortunately, the data does not appear to be cooperating.

A Harvard Business Review (HBR) Analytic Services survey sponsored by Cloudera put a number to a concern I’ve raised in client meetings for two years. Only seven percent of respondents call their data completely ready for AI. Not somewhat ready. Completely ready. I read that stat twice. Ninety-three percent of the organizations spending real budget on AI strategy right now are building on a foundation they already suspect is shaky.

I run into this constantly. Ask a client to show you the data behind their AI roadmap, and the room is noticeably silent. There’s plenty of data in most organizations. There’s rarely data a model can use, let alone an agent that’s supposed to act on its own.

Most leaders assume “ready” means clean and structured. That’s part of it. It’s not the part that worries me most. What worries me is that readiness for agentic AI is more than clean data. It’s whether the agent has a structure it can trust instead of text it has to continually reinterpret.

 

Why agents change the definition of ready

That distinction matters more once you’re talking about agents instead of chatbots. Generative AI is fundamentally probabilistic. It predicts a likely next answer; it doesn’t retrieve a guaranteed one. That’s tolerable when a person checks the output before anything happens. It’s much less tolerable once an agent has permission to actually do something like update a record, approve a request, or kick off a process on its own.

The structured-data problems are familiar by now: disconnected systems, no lineage, institutional knowledge locked in someone’s personal spreadsheet. Another HBR Analytic Services study sponsored by Hyland found something that gets missed constantly: 65 percent of leaders say their structured data is at least somewhat AI-ready, but only 39 percent say the same about unstructured data. That’s the contracts, emails, PDFs, and other working the documents that contain much of what an organization knows. If you’re only optimizing the tables, you’ve addressed maybe a third of the real problem.

 

Bring the right structure to unstructured data

And the problem with that unstructured knowledge isn’t only that it’s messy. It’s that it was written for people, not systems.

A wiki, a shared drive, or a stack of onboarding docs is written for a human to read and infer meaning from. An LLM working from that kind of source has to re-derive every relationship from scratch, on every query. Andrej Karpathy’s LLM Wiki framing describes this well. Without compiling the knowledge once and cross-referencing it, the model just keeps redoing the same interpretive work. And every time it reinterprets, it can interpret differently.

A knowledge graph does that compiling up front, defining meanings and relationships for the LLM. This policy applies to this entity. This field traces back to that source system. This exception overrides that rule.

Enterprise data architecture has its own name for the same idea. The ontology layer is the translation between raw rows in a database and the actual business objects a system needs to act on: this customer, this claim, this shipment, and how they connect. Skip it, and every team downstream ends up inventing its own private interpretation of what the data means.

So does every agent.

 

Moving toward completely ready

A recent Harvard Business Review article on agentic orchestration made a point that stuck with me. Most companies have deployed AI inside individual functions, but very few have built the cross-functional decision infrastructure that lets AI and people coordinate across procurement, finance, legal, and operations at the same time. That infrastructure only holds up if the data underneath has real, governed structure resulting in something deterministic the agent can rely on. Otherwise, the agent is referencing a paragraph it has to interpret without guardrails to ensure it did so correctly.

The organizations pulling ahead built real data platforms early and are investing in graph- or ontology-based structure instead of static documentation. The rest are retrofitting mid-project, at a much higher cost than if they’d done it up front.

A few questions worth running before your next AI investment. Can you trace a single data element back to its source system in under five minutes? Is your most important institutional knowledge sitting in a document, or in something a system can reliably query? If an agent had to act on this data tomorrow with no human checking its work, would it get it right?

If I’m honest about my own client base, I’d put most of them around a four or five out of ten on true readiness. That’s better than the survey average, worse than I’d like. I’m curious where you’d land.

About the author

Connect

Find out how our team can help you achieve great outcomes.

Insights delivered to your inbox