SeriesThe Data Mesh DiariesPart 3

Data Architecture Data Mesh

What Is a Data Product - and Why "Data as a Product" Is Not the Same Thing

Why did the industry need the idea of a 'data product' in the first place? Once you understand that, the distinction between a data product and data as a product becomes obvious.

The confusion

Two phrases circulate in every data programme today: a data product, and data as a product. They sound interchangeable. They are not, and most teams never define either one precisely enough to notice.

Data as a product is a philosophy - a commitment to treat data with the rigour a product team gives software. A data product is the artifact that philosophy produces when someone actually does the work. Confusing them is not a small vocabulary error. It is the reason organisations end up with a catalog full of things with a name and passing almost none of the characteristics of a product.

Why it matters

At some point, a CDO asks the team a direct question: "Do we have data products?"

The answer to this question is the basis for an important decision - investment, architecture and increasingly, what an AI agent is allowed to consumer unsupervised. Answer based on the wrong definition in your head and you have misguided your leadership.


67%

of data and analytics professionals don't fully trust their organisation's data for decision-making (Precisely & Drexel LeBow, 2025)

12%

say their data is of sufficient quality and accessibility for effective AI (Precisely & Drexel LeBow, 2025)

77%

rate their organisation's data quality as average or worse (Precisely & Drexel LeBow, 2025)


How data teams were organised before this

Before Dehghani made the term "data product" popular, data teams were organised around projects, not consumers. A request came in. A pipeline got built. Success meant the pipeline ran and the numbers landed in the report on time.

Nobody asked whether someone outside the team could pick up that output later and use it without help - because for a long time, almost nobody did. A pipeline had one consumer in mind: the report it fed. If a second consumer showed up, they asked the person who built it. That worked. It was inefficient, but it worked.

It was wrong for decades, then why did nobody fix it earlier?

Because it wasn't "wrong" yet. It was inefficient, and inefficiency at small scale doesn't create enough pain to demand a fix. When an organisation has twenty pipelines and every consumer can reach the two or three engineers who built them, a phone call is a perfectly functional interface. The central team absorbs the effort required and the effort is small enough to absorb. Nothing about that arrangement looks broken from the inside. It looks like normal operations - a bit slow, occasionally annoying, never urgent enough to redesign around.

The assumption that broke

The 'central' arrangement rested on one assumption: a central team can remain the living interpreter for everything it builds, indefinitely, for anyone who ever needs it.

That assumption is not linear. It is combinatorial. Ten pipelines and ten consumers is a hundred possible interactions the central team might field. A hundred pipelines and a hundred consumers is ten thousand. The team's headcount does not grow at the same rate as the number of people who need to ask it something. At some point, every senior engineer's calendar is entirely phone calls explaining fields in datasets they built two years ago, and no new pipeline gets shipped because the team is fully occupied answering questions about the old ones. Do you see the inefficiency in this arrangement? And why this very arrangement becomes a hurdle in data driven enterprise that your organisation aspires to be?

Software never carried this assumption in the first place. A software product's users are strangers by design, from day one - nobody expects the engineer who wrote the checkout flow to personally explain it to every customer. Data teams inherited no such discipline, because for most of their history, they didn't need it. The assumption was fragile the entire time. It simply hadn't been tested past its limit.

Why 2019, and not in 2010

Four things converged and tested it at once.

Cloud object storage made it cheap to hold far more data than before, so far more of it needed interpretation by people who hadn't built it.

Self-service analytics tools put dashboard-building directly in the hands of business users, multiplying the number of non-engineers trying to consume data independently.

Domain-oriented architecture - the same thinking already reshaping software into microservices - made "why isn't data organised this way too" an obvious question to ask.

And the earliest wave of enterprise AI started asking for data as a direct input, not a report someone had already sanity-checked.

All four pushed in the same direction: more consumers, arriving faster, while the central team's capacity to personally interpret for each one stayed exactly where it had always been. The combinatorial math from the previous section stopped being theoretical. Zhamak Dehghani named the failure in 2019 because 2019 was when enough organisations were living inside it to recognise the description.

The new philosophy: data as a product

"Data as a product" is the response - not a rebrand, a rejection of the old assumption. It is the commitment to build data the way a product team builds software: for a consumer who was never in the room, using rigour and ownership instead of a phone number.

The philosophy asks who the consumer is, what they need, what quality they can depend on, and who answers for it when it breaks. It is a mindset, not a deliverable. A team can hold it sincerely and have shipped nothing yet.

The artifact: a data product

A data product is what the philosophy produces when someone does the actual work: a self-contained unit built to serve analytical data to consumers, in a form they can use and trust, without depending on the team that built it.

Self-contained is the operative word. It packages everything a stranger needs to use the data independently:

None of this is exotic. That is the point. Dehghani imported the ordinary discipline of product management into a field that had never needed it before 2019.

Answering the CDO's question properly

The wrong way to answer "do we have data products?" is by inventory - counting catalog entries. That reports the failure pattern above as success. The right way is to score the assets carrying the name against the eight characteristics and report the count.

A useful answer sounds like: "We have 40 assets labelled data products. 3 of those 40 pass all eight characteristics of a data product. The most common failures are discoverability, self-describing documentation, and a committed quality SLO." That tells the CDO exactly where the organisation stands and what will it take to close that gap - ownership and contracts, not more platform.

The innocent answer - "yes, we have forty" - costs more than embarrassment. Investment gets allocated on it, and the gap surfaces later, at a worse time, in front of a consumer who depended on a "product" that was never one.

In my experience, most flagship assets fail at least five of the eight, and the failures cluster in the same three places: nobody can find the asset without asking, nobody can understand it without calling its builder, and no quality commitment exists. Those are the same three consumer-dependency problems this whole piece opened with. They just have new names on the catalog page.

AI exposes the problem. It does not create it.

For years, people absorbed the cost of this quietly. Analysts learned which datasets to trust, whom to call, which numbers to double-check before a board meeting. Human judgement was the workaround - every gap in the missing contract got filled by someone's memory and someone's phone call.

That workaround is what AI removes, not the underlying problem. When (not if) an agent consumes your data directly, there is no sceptical human left to catch the bad dataset. It does not know a table is "usually fine except at month-end." It reads what it is given and acts. A genuine data product - self-describing, quality-committed, addressable - is infrastructure an agent can safely build on. A relabelled pipeline was always a liability. It simply had a human standing in front of it, absorbing the risk. AI does not introduce a new weakness, it exposes what humans have been compensating for.

Only 12% of organisations report their data is ready for that. The eight characteristics were written years before agentic AI existed. They now read like prerequisites.

Conclusion

This is not a terminology problem. It is an operational one. If you cannot distinguish a data asset from a data product, you cannot accurately assess the maturity of your own data estate - and every decision built on that assessment, from budget to what you let an agent touch, inherits the error.

The confusion started with an assumption that held for decades and then didn't. The fix is not a better name. It is the contract, the owner, and the eight characteristics that make the name true. The next time someone asks whether you have data products, answer with a count.

Part 4 takes on the harder question this raises: if the test is this clear, why do so few organisations pass it? The answer is not technical. It is the eight failure modes - the specific, repeatable ways data mesh efforts die inside large enterprises.

Knowledge Check
0 / 4 ☆☆☆☆
Q1 Multiple Choice

Why, according to the article, did the interpretation-bottleneck problem go unfixed for decades before 'data product' became a term?

A
Data engineers lacked the necessary skills to build documentation
B
At small scale, a phone call to the original engineer was a functional interface - the problem was inefficient, not broken
C
Regulations prevented data teams from adopting product practices
D
The problem did not actually exist before 2019
💡 Insight

The article's core historical claim: the underlying assumption (a central team can stay the living interpreter for everything it builds) was fragile from the start but only becomes visibly broken once its combinatorial cost - consumers × pipelines - exceeds what a small team can absorb.

Q2 True / False

'Data as a product' is the artifact you build, and 'a data product' is the philosophy behind it.

True
False
💡 Insight

It is the reverse. 'Data as a product' is the philosophy - the commitment to product-grade rigour. 'A data product' is the concrete artifact the philosophy produces when applied seriously.

Q3 Multiple Choice

Which four forces does the article cite as converging around 2019 to break the old assumption?

A
Regulation, outsourcing, cost-cutting, and consolidation
B
Cloud object storage, self-service analytics, domain-oriented architecture, and early enterprise AI
C
Data warehousing, ETL tools, business intelligence, and dashboards
D
Data mesh, data fabric, data lakehouse, and data virtualization
💡 Insight

All four increased the number of people or systems trying to consume data directly, while the central team's capacity to personally interpret for each one stayed flat.

Q4 Multiple Choice

How does the article characterise AI's relationship to the data-product problem?

A
AI created an entirely new category of data quality problem
B
AI solved the problem by automating data cleaning
C
AI removed the human workaround (analyst judgement) that had been quietly absorbing the cost of poor data products, exposing a weakness that already existed
D
AI has no meaningful relationship to the data product problem
💡 Insight

The article's explicit framing: AI did not create the problem - it removed the sceptical human buffer that had been compensating for it.

Q5 Reflection

Your CDO asks tomorrow: 'Do we have data products?' Using the method in this article, draft your answer: pick the assets carrying the name, score them against the eight characteristics, and state the count, the most common failures, and what closing the gap requires.

Model Answer

A strong answer gives a specific count ('X labelled, Y pass'), names the failing characteristics precisely, and frames remediation as ownership and contracts rather than new platform purchases.

💡 Insight

This question tests whether the learner can distinguish between data products in name and data products in substance. The eight characteristics are the diagnostic instrument - labelling an asset a 'data product' without scoring it against discoverability, understandability, trustworthiness, natively accessible, interoperable, valuable on its own, secure, and addressable is exactly the superficiality the article warns against. The framing of remediation matters too - a learner who concludes 'we need a better catalog tool' has missed the point. The gaps almost always trace back to unclear ownership and absent contracts, not missing technology.

0
out of 4 correct

Question 1 of 5