So if we stop trying to create a universal dictionary, what are we going to do instead?
A better approach would be one that takes advantage of the tactical point solutions that most enterprises seem to succeed with. The pejorative idea of a “data silo” should now form the basis of our solution.
We start by building systems and architectures based on the recognition that different contexts exist. I posit that such concepts as “schema on read” and the associated technologies invented to support the “Big Variety”* inside of our modern data platforms is a good place to start.
My proposal: create a “data thesaurus” capturing in meta data all aspects of the structure, meaning, flow, usage and context of data as it appears and moves about. This thesaurus would contain the obvious syntactic reference material of traditional data dictionaries, but then would include data describing the semantic relationships between different collections of data. By explicitly permitting multiple, varying DOMAINS to co-exist for different purposes and uses, the same datasets could be mapped to multiple uses without forcing a premature homogenization to a unified model.
Finally, it would explicitly include content indicating the pragmatics (usage patterns) by capturing in declarative form the transformations to apply as the data moves from one domain context to another. The scientific community has been working in this data management solution space for many years, although perhaps not recognizing its potential for the wider world.
Provenance systems point to a promising potential, given that they track lineage through and across transformation and manipulation. Every dataset, from the initial registration of data when first received, through every transformation and manipulation, must be tracked. While the system would not have to retain every variation, it should be able to reproduce any variation from any point in time.
To begin, build data dictionaries for individual systems or even small groups of files (as is often the full extent attempted and completed in most organizations). But instead of trying to extend these point solutions into a universal solution, links across contexts within the organization would be filled in only as practicality required, as the by-product of data integration projects or system consolidation efforts. In this way, the thesaurus will grow in content and utility, but not require unification in its entirety, at any stage.
* Big Variety is one of the three challenges leading to the adoption of Big Data, as described by M. Stonebraker. It describes the situation where the variety of data structure that must be integrated within an enterprise has ballooned to where the traditional practice of data management can no longer keep up.