The “Market for Lemons” Problem in Data Platforms
Organisations often measure the success of their data platforms by volume.
How many datasets have we onboarded? How many tables are available? How many data assets can users access?
The assumption is simple: more data means more value.
I think that assumption is often wrong. Past a certain point, adding technically accessible but poorly understood data does more than fail to create value. It makes every asset harder to assess, including the good ones. Curation, rather than volume, is what makes a data platform valuable.
Adding more data does not always lead to greater use. Users may still struggle to establish whether a table is current, what it was created for or who is responsible for it. Sometimes analysts create their own version because they cannot easily assess the one already available. The platform has more assets, but users are not necessarily getting more from it.
The lemons analogy, and where it does not quite fit
Economist George Akerlof explored a similar problem in his famous “market for lemons” theory. In a market where buyers cannot reliably distinguish a good used car from a bad one, known as a “lemon”, uncertainty affects the value of every car in the market.
It is worth being precise about the analogy. In Akerlof’s market, the information asymmetry is adversarial. Sellers know more about the quality of their cars than buyers and may have an incentive not to reveal what they know.
In data platforms, the asymmetry is rarely adversarial. More often, it is the result of organisational memory loss. Nobody is deliberately hiding the fact that a data asset is no longer maintained. The person who understood its history has moved on, the documentation explaining its limitations was never written, and ownership has changed hands more than once.
The circumstances are different, but users face the same difficulty: they cannot easily tell which assets they can trust or whether they are suitable for what they need to do. Making more data technically available does not resolve that uncertainty. In this case, the lemon problem arises not from conflicting incentives, but from unclear ownership, limited documentation and the gradual loss of organisational knowledge.
What “cost” actually means here
It is easy to refer to the “cost of assessment” without explaining what that cost includes. In practice, it tends to appear in several ways:
- Analyst time. Someone spends half a day confirming whether a table is still being updated rather than half an hour using it.
- Duplicated pipelines. Teams that do not trust an existing asset build their own version, increasing the maintenance burden rather than sharing it.
- Decisions based on stale data. Someone acts on an asset that appeared current but was not.
- Divergent answers. Two teams query what appears to be the same data and reach different conclusions because neither knows which version, definitions or caveats the other is using.
None of these costs appears in an “assets onboarded” metric. They surface elsewhere: in wasted time, duplicated effort, inconsistent decisions and, eventually, declining trust.
Where the problem becomes most visible
The problem is particularly visible during organisational handovers, when a team inherits a large portfolio of data assets without sufficient documentation, lineage or knowledge transfer.
The receiving team then faces a real choice, often under considerable pressure.
One option is to treat every asset as equally important and attempt to maintain, migrate or modernise all of them. This can feel like the safest choice because nothing is explicitly deprioritised.
Another approach is to start by asking which assets still provide meaningful value to users.
This is not only a technical question. Engineering teams need to know what should be maintained and what could be archived. Governance teams need to consider whether an asset still has regulatory or research importance. Users need clear information about what is reliable and available now.
Resolving these tensions requires agreed triage criteria before the handover. Otherwise, the default becomes “keep everything”, which is still a decision, just an unexamined one.
Maintaining data products has a real cost. Even an established asset can continue to require engineering effort, quality assurance, documentation, governance, domain expertise and service support.
There is an opportunity cost, too. Time spent maintaining every historical asset is time not spent on what users need now. A decision to maintain everything is therefore not neutral. It is also a decision about where the organisation will not invest.
Curation is a product decision.
This is why I increasingly see data curation as a core part of data product management.
A good data platform should not simply maximise the number of assets it exposes. It should help users find, understand and trust the right ones.
That requires deliberate decisions about which assets deserve continued investment. Not every criterion should carry the same weight. Some factors should determine whether an asset can remain an actively maintained product at all. Others should help prioritise among the assets that meet that minimum standard.
Gating factors
Before continuing to invest in a data product, an organisation should be able to answer three basic questions:
- Is the data good enough for its intended use?
- Is there a clear owner and someone who understands it?
- Can it be maintained and supported over time?
If the answer to any of these questions is no, the asset may still be retained, but it should not necessarily be treated as an actively maintained product.
Among the assets that pass those gates, investment can be prioritised according to:
- current or anticipated user and research value;
- strategic relevance;
- currency and expected future coverage;
- and uniqueness compared with existing assets.
Some assets will justify significant investment. Others may remain available as historical or legacy datasets, clearly labelled as such. Some may no longer justify continued investment at all.
That is not a failure of the data platform. It is portfolio management.
Treating every asset as equally valuable simply because it exists is also a product decision. It is just an implicit one.
The question is worth asking.
The aim is not necessarily to reduce the amount of data available. It is to clarify what is actively maintained, what is kept for historical purposes, and what may have limitations that users need to consider.
A smaller collection of well-understood, documented and supported data products may create more value than a vast catalogue whose quality users must investigate for themselves.
Perhaps the most important question for data organisations is not:
“How much data have we made available?”
but:
“How much of the data we make available can users confidently understand, trust and use?”
Because there comes a point when the question changes from “Which asset should I use?” to “Can I trust this platform at all?”