The Agentic AI Builder's Playbook

Your AI Strategy Has a Single Point of Failure: Your Data

By Jason Newell · ~5 min read · Strategy

Every organization investing heavily in AI right now is making the same structural mistake. They're pouring resources into models, pipelines, and infrastructure — and building all of it on a foundation that's already cracked.

The crack is your data.

The Stack That Most Teams Build

The typical enterprise AI stack looks like this, from top to bottom:

  • AI Models
  • Pipelines
  • Infrastructure
  • DATA (the foundation — and where it breaks)

The problem: teams spend months perfecting the layers above while ignoring the cracked foundation beneath. You can have the best model in the world sitting on top of duplicate, stale, siloed data — and the model will faithfully surface garbage, at scale, with confidence.

What Nobody's Fixing

A few uncomfortable data points:

  • 73% of enterprise data is never analyzed — it exists, but it's dark, untouched, and therefore useless to AI
  • Duplicate, stale, siloed data is the norm, not the exception — most organizations have the same data in multiple places, none of them authoritative
  • No data quality pipeline before AI pipeline — teams build the AI pipeline and assume the data going in is clean. It usually isn't.
  • Models trained on garbage produce garbage decisions — but with fluent, confident prose, which makes them harder to catch than obvious system failures
  • "We'll do metadata and schema governance later" — later never comes, and the technical debt compounds

AI Doesn't Have a Model Problem. It Has a Data Culture Problem.

This is the core insight. The failure mode isn't that GPT-4 vs Claude makes the wrong choice. The failure mode is that the best model in the world, when given incomplete, inconsistent, or incorrect data, produces wrong answers with high confidence.

The data maturity vs AI investment matrix:

  • High data maturity + High AI investment = Winners: AI works, produces genuine business value, compounds over time
  • Low data maturity + High AI investment = Expensive Demos: Impressive in a boardroom, broken in production, trust erodes fast

Most companies are in the bottom right quadrant. They're investing heavily in AI on top of data foundations that can't support it.

Fix the Foundation First

The prioritization should be:

  • Data quality pipeline before AI pipeline
  • Schema governance and metadata before model selection
  • Single source of truth before RAG implementation
  • Data access controls before agent deployment

This isn't a popular message when everyone wants to move fast and ship agents. But the teams that invest in data foundations early are the ones whose AI investments compound. The teams that skip it are the ones explaining to executives why the chatbot is hallucinating factual errors about their own products.

The Practical Starting Point

Three high-impact, low-cost actions:

  • Audit your dark data — identify the 73% that's never analyzed. Some of it is critical context your agents don't have.
  • Establish one authoritative source for the top 10 data entities your AI will reason about
  • Build a data quality gate before your AI ingestion pipeline — not after

Your AI strategy is only as good as your data strategy. Fix the foundation, and everything else gets easier. Skip it, and everything else gets harder — forever.

Jason Newell is an AI practitioner, builder, and writer covering agentic systems, developer tooling, and the future of AI engineering.

Agentic AISER-AGT-010

Related

Related field notes

Keep reading

More field notes

This piece is part of the MAX Research Collective library. Browse the rest, or connect on LinkedIn.