catalog_evaluation.txt

    Choosing a Data Catalog for AI Readiness

    A platform evaluation scored on whether governance could be automated, not on feature count

    Discipline
    Platform evaluation
    Domains
    Healthcare · AI and Automation
    Timeframe
    2025 to 2026
    Status
    Approved as the evaluation approach. Scoring and the final recommendation moved to the governance program.
    what_was_this_project.txt

    What was this project?

    Choosing a Data Catalog for AI Readiness is an evaluation framework built to compare two enterprise data catalog platforms ahead of a governance program relaunch at a large health system.

    The organization was standing up a new version of its data governance program and wanted to know whether its current catalog was the right foundation for it. The two candidates were Alation, already in place, and data.world as the challenger.

    The timing mattered more than the shortlist. A catalog decision made before a governance program launches is a design decision. The same decision made after launch is a migration.

    what_problem_did_this_solve.txt

    What problem did this solve?

    Non technical staff could not reliably find trustworthy data, so they stopped trying, and the analytics and AI work downstream inherited both the delay and the doubt.

    Two failures compounded each other. Discovery was hard, so people asked a colleague instead of searching. Documentation was inconsistent, so when they did find a dataset they could not tell whether it was the governed one or somebody's copy.

    The consequence is trust, not time. A team that cannot verify a dataset's lineage will not put it under a model, and that is where a catalog problem quietly becomes an AI problem.

    what_i_observed.txt

    What did I observe?

    The interesting question was not which catalog had more features, because both had enough, but which one could take governance work off people and still leave an audit trail behind.

    Three things reframed the evaluation once I looked at the actual work.

    • The incumbent had recently shipped persona based homepages aimed at the exact usability complaint that had been raised before. Evaluating it on last year's weakness would have produced the wrong answer, so the framework had to test the current product rather than the remembered one.
    • Governance was going to fail on volume, not on policy. Approval routing, document triage, sensitive field identification, and term creation were all manual, and no amount of policy writing reduces that load. Automation of those four was the real requirement.
    • AI enablement meant two different things to the two vendors. One was automating stewardship. The other was semantic search and natural language query. Scoring them against a single AI category would have compared unlike things, so the framework had to separate the capability from the marketing word.

    Clinical system integration also turned out to be a discriminator rather than a checkbox, because a catalog that cannot see into the EHR data estate cannot govern the datasets that matter most in a hospital.

    my_role.txt

    What was my role?

    I designed the evaluation, which means I chose what would be measured, wrote the scoring method, built the comparison matrix, and specified the governance automation flow the winning platform would have to support.

    • Defined ten evaluation categories, including two the original request did not have: clinical system integration, and AI enablement as a capability rather than a label.
    • Wrote the scoring approach, rating each category on feature completeness, ease of use, and future readiness, so a platform could not win on features alone.
    • Built the side by side comparison matrix and wrote the key insights, including the risk considerations around migration and adoption.
    • Designed the five step governance automation flow the platform would need to support, with its two human decision points.
    • Planned the qualitative sessions, so the usability score came from watching people use the personas rather than from reading the release notes.

    I did not run the procurement, negotiate pricing, or make the final platform selection. The framework and the recommendation path were mine. The decision belonged to leadership.

    the_proposed_solution.txt

    What was the proposed solution?

    The proposed solution is a ten category scoring framework paired with a specified automation flow, so that the platform choice is made against the governance work the organization actually has to do rather than against a feature list.

    The framework covers data discovery and search, metadata management, collaboration and sharing, integrations, governance and compliance, user experience and navigation, scalability, AI enablement, clinical system integration, and total cost of ownership.

    Each is scored one to five on three axes: feature completeness, ease of use, and future readiness. Three axes rather than one is the whole point. A platform can be feature complete and unusable, or pleasant and a dead end, and a single score hides both.

    The second half of the deliverable is the automation flow. Five automatable behaviours were specified as requirements the platform has to meet:

    • Approval workflows for data access and stewardship requests.
    • Document triage for governance review.
    • Field identification for sensitive data classification.
    • Term creation when a new concept appears without a governed definition.
    • Stakeholder approval routing, with an audit trail attached to every decision.
    system_design.txt

    How was the system designed?

    The automation flow moves a new dataset from arrival to publication through two automated checks and one human gate, so that routine cases clear themselves and only the exceptions consume a steward's attention.

    The shape of the flow is the argument. Every automated step is a proposal, and the two points where an error would actually cost something, the compliance check and final publication, resolve to a named person rather than to a rule.

    The enrichment loop matters too. An incomplete asset is sent back to be enriched and rechecked rather than blocked or waved through, which is what keeps the catalog from filling with technically present but practically useless entries.

    Data catalog governance automation, proposed flowA decision flow. A new dataset arrives and the catalog checks whether a governed term already exists for it. If one does, the existing term is reused. If not, a term is generated through the catalog API. The flow then checks whether required metadata is complete; if it is not, an enrichment workflow is triggered and the check runs again. Once metadata is complete a data asset is created, and a compliance check runs. Passing assets publish to the catalog automatically. Failing assets route to a named human for approval, and only publish after that person approves. Every automated step is a proposal. The two gates that can cause harm, compliance and final publication, always end at a person.new datasetgoverned term exists?no, generate onegenerate termmetadata complete?no, enrich and recheckenrichcreate data assetcompliance checkpass or route to a personfailpasshuman approvalnamed stewardpublishwith audit trailAutomation proposes. A steward still approves.
    Choosing a Data Catalog for AI Readiness: system design. The full description is in the diagram's alternative text.
    why_this_design.txt

    Why was it designed this way?

    The framework is weighted toward future readiness and governance automation because the organization was buying a foundation for a program that did not exist yet, not a tool for the work it was doing that week.

    Design choices and the alternatives considered
    ChoiceAlternative consideredWhy the choice won
    Score on three axes per categoryA single weighted score per categoryA single number hides the difference between a platform that cannot do something and one that can do it badly. Those two failures have different remedies and different costs.
    Clinical system integration as its own categoryFolding it into general integrationsIn a hospital, the datasets that most need governing are the ones nearest the EHR. A platform that integrates broadly but not clinically fails at exactly the point where governance matters most.
    Split AI enablement from stewardship automationOne AI categoryThe two vendors meant different things by AI. Semantic search and natural language query help a person find data. Stewardship automation removes work. Scoring them together would have compared a discovery feature against an operational one.
    Human approval kept in the automated flowFully automated publication once checks passA compliance check that can be wrong needs somewhere for the wrongness to land. Routing failures to a named steward keeps accountability attached to a person and gives the audit trail something to record.
    Evaluate the current product, not the remembered oneScoring against the previously documented complaintsThe incumbent had shipped persona based homepages aimed squarely at the prior criticism. Judging it on the old version would have produced a recommendation that was already out of date.

    The risks I named were migration complexity, governance alignment, and stakeholder adoption. Adoption is the one that decides it. A catalog nobody opens is a subscription, not a system, which is why usability was scored from watching people rather than from a feature comparison.

    implementation_plan.txt

    How would it be implemented?

    The evaluation is sequenced so that the qualitative evidence is gathered before the scores are assigned, which prevents the framework from becoming a way to justify a decision already made.

    1. Run product experience sessions with real users on both platforms, including the persona based views, and capture qualitative usability feedback.
    2. Score both platforms against the ten categories on all three axes, with written justification per score rather than a bare number.
    3. Test the five automation behaviours against each platform's actual capability, not its roadmap.
    4. Present findings and a recommendation to leadership, including the migration cost and the adoption risk, not only the winner.
    5. Align the chosen platform's rollout to the governance program launch, so the catalog and the policy arrive together.
    expected_impact.txt

    What impact was expected?

    Every figure below is a projection from the design rather than a measured result, because this evaluation framework produced a recommendation and not a deployment I can report outcomes from.

    Projected effects and how each would be measured
    What would changeBaseline todayHow it would be measured
    Datasets get found and verified without asking a personDiscovery happens by asking a colleague. Trust is personal, not systemic.Time from search to a verified dataset, and the share of requests resolved without a human handoff.
    Governance stops scaling with headcountApproval routing, triage, classification, and term creation are all manual.Share of governance actions cleared automatically, and steward hours per hundred assets.
    Lineage becomes something a person can point atDocumentation is inconsistent, so lineage claims are unverifiable.Share of published assets with complete lineage and an owner of record.
    The platform decision survives the AI roadmapThe current catalog was chosen before AI enablement was a requirement.The future readiness score, revisited annually against what the AI programme actually needs.

    I did not project a cost saving. The honest position was that the two platforms had different pricing shapes, enterprise subscription against flexible tiers, and that comparing them required procurement numbers I did not have.

    what_happened.txt

    What happened?

    The evaluation framework and the automation flow were approved as the way to make the decision, and the scoring and final platform recommendation moved to the governance program.

    I built the instrument rather than the verdict. What I can say is that the framework changed the question from which catalog is better to which catalog can carry the governance load without adding headcount, and that reframing was the contribution.

    I do not know which platform was ultimately selected, and I am not going to guess in a portfolio entry.

    related.txt
    Working on something like this?

    Bring me the problem before the tool selection.