AnyLearn
All lessons
Businessintermediate

Who Owns the Data? Roles and Operating Models

Governance is accountability, and accountability needs names. Learn the core data governance roles, owner, steward, and custodian, and exactly who does what, then the operating models (centralized, decentralized, federated), why heavy committees fail, and how domain ownership and data-as-a-product push accountability to the teams closest to the data.

Updated · AI-authored, review-gated · how lessons are made

Not signed in: your progress and quiz score won't be saved.
Progress1 / 8

Accountability needs a name

The last lesson defined governance as decision rights and accountability. This lesson answers the question that makes it real: accountable to whom, exactly? Governance that belongs to "everyone" belongs to no one, so the heart of any governance program is assigning specific, named responsibility for specific data.

Think back to the opening scene of three departments reporting three customer counts. The deepest reason that happens is that no single person is accountable for the definition and quality of "customer" data. When something is everyone's shared concern and no one's explicit job, it drifts. Assigning clear ownership is the single most important structural move in governance.

The key concept is decision rights: for any given data, someone must have the authority to decide what it means, who can use it, and what quality it must meet, and someone must be answerable when those things go wrong. Without named decision rights, every disagreement about data becomes an unresolvable standoff, exactly the meeting that dissolved into argument.

So governance requires a structure of roles, a clear division of who is responsible for what. Over decades, the field has converged on a standard set of roles that separate business accountability from technical care-taking, and a set of operating models for how those roles are arranged across an organization. This lesson covers both: first the roles, then the models, then the modern evolution that is reshaping how ownership is assigned.

Full lesson text

All 8 steps on one page, for reading, reference, and search.

Show

1. Accountability needs a name

The last lesson defined governance as decision rights and accountability. This lesson answers the question that makes it real: accountable to whom, exactly? Governance that belongs to "everyone" belongs to no one, so the heart of any governance program is assigning specific, named responsibility for specific data.

Think back to the opening scene of three departments reporting three customer counts. The deepest reason that happens is that no single person is accountable for the definition and quality of "customer" data. When something is everyone's shared concern and no one's explicit job, it drifts. Assigning clear ownership is the single most important structural move in governance.

The key concept is decision rights: for any given data, someone must have the authority to decide what it means, who can use it, and what quality it must meet, and someone must be answerable when those things go wrong. Without named decision rights, every disagreement about data becomes an unresolvable standoff, exactly the meeting that dissolved into argument.

So governance requires a structure of roles, a clear division of who is responsible for what. Over decades, the field has converged on a standard set of roles that separate business accountability from technical care-taking, and a set of operating models for how those roles are arranged across an organization. This lesson covers both: first the roles, then the models, then the modern evolution that is reshaping how ownership is assigned.

2. The three core roles

Data governance distinguishes three roles that are constantly confused but do genuinely different jobs. Getting them straight is essential, because assigning them wrong is a classic cause of failure.

  • Data owner: a senior business person accountable for a data domain, for example the head of sales owning customer data. The owner is answerable for the data overall: they set the rules, approve access, decide definitions, and carry ultimate responsibility. Ownership is about authority and accountability, and it belongs to the business, not IT.
  • Data steward: the hands-on subject-matter expert who manages the data day to day. Stewards define and maintain the meaning of data elements, monitor quality, resolve issues, and act as the go-to person for questions about that data. If the owner sets direction, the steward does the ongoing work of keeping the data right. Stewardship is about day-to-day management and expertise.
  • Data custodian: the technical role, usually in IT, responsible for the safe transport, storage, and protection of the data in systems. Custodians implement the controls the owner requires, run the databases, manage backups and access mechanisms. Custody is about technical safekeeping, not deciding what the data means.

Above these sit governance bodies, a governance council or data governance office, that set enterprise-wide policy, arbitrate cross-domain disputes, and coordinate the owners and stewards.

The crucial split is between accountability (owner), expertise and daily work (steward), and technical safekeeping (custodian). Blur these and governance breaks: give definitions to the custodian and IT ends up guessing business meaning; leave stewardship unassigned and quality quietly rots with no one responsible for it.

3. Owner, steward, custodian: a worked example

Make the roles concrete with a single data element: a customer's email address, used across marketing, sales, and support. Watch how the three roles divide the work.

RoleWhoResponsibility for the email field
ownerVP of Salesaccountable overall; approves who may use it; sets the rule that it must be valid and current
stewardsales operations analystdefines "valid email"; monitors how many are missing or malformed; chases fixes; answers questions about it
custodianIT database teamstores it securely; enforces access permissions; encrypts it; runs backups

Now trace a problem. Support notices that many customer emails are bouncing. Who does what? The steward investigates, identifies that a web form was accepting malformed addresses, and quantifies the issue. The custodian may implement a technical validation check in the system. The owner decides the policy, that email must be validated at entry, and approves the resources to fix it. Each acts in their lane, and the problem gets resolved with clear accountability rather than finger-pointing.

Notice what would happen if the roles were missing. With no steward, no one would notice or diagnose the bouncing emails. With no owner, no one could authorize the fix or set the rule. With no custodian, the fix could not be implemented safely. The three roles are complementary, and a gap in any one leaves data problems to fester unowned.

This is why role assignment is the practical core of governance: it converts "someone should fix the data" into "this named person is responsible," which is the difference between governance that works and governance that is just documents.

4. Centralized, decentralized, federated

Once you have roles, you must decide how to arrange them across the organization. This is the operating model, and there are three basic shapes, each with a clear trade-off between consistency and agility.

  • Centralized: a single central team holds governance authority and makes the decisions for everyone. Strength: maximum consistency, one set of definitions and rules. Weakness: the central team becomes a bottleneck and often lacks the deep local knowledge of each business area, so it is slow and can feel disconnected from the front lines.
  • Decentralized: each business unit governs its own data independently. Strength: agility and local relevance, each area moves fast and knows its own data. Weakness: inconsistency, the same term ends up defined five different ways, recreating the original chaos across silos.
  • Federated: a hybrid, and the most common choice in practice. A central body sets enterprise-wide standards and policies (shared definitions, common rules, overall framework), while individual business domains own and manage their own data within those standards. Strength: it balances consistency with local ownership. Weakness: it requires ongoing coordination to keep the balance.

The federated model dominates because it resolves the core tension. Pure centralization is consistent but slow and out of touch; pure decentralization is fast but chaotic. Federation says: agree centrally on the rules of the road (what "customer" means, what quality is required, what policies apply), then let each domain drive within them.

Choosing a model is really choosing where on the spectrum between control and autonomy an organization wants to sit, and most land in the federated middle because both extremes fail in predictable ways.

5. The committee trap

A specific failure deserves its own warning because it is so common: governance that collapses into an endless series of committee meetings. This is the operating-model version of the bureaucracy failure from the last lesson.

The pattern: an organization stands up a data governance council, and governance becomes synonymous with attending meetings. A large group convenes regularly to discuss data, but real decisions are slow, diffuse, and often not made at all. Because a big committee owns everything jointly, no individual is truly accountable, and the group defaults to more discussion. Meanwhile the people doing actual work find the committee is a place requests go to be delayed, so they route around it. The council produces minutes, not outcomes.

The root problem is confusing coordination with decision-making. Councils are genuinely useful for coordination: setting shared standards, arbitrating cross-domain disputes, aligning direction. But they are terrible at the day-to-day decisions that governance actually requires, because those need a single accountable person who can decide quickly, not a group that must convene.

The fix is to push most decisions out of the committee and down to owners and stewards who have clear individual decision rights, reserving the council for the genuinely cross-cutting issues that need coordination. Governance should be measured by decisions made and problems solved, not meetings held.

The deeper principle: effective governance lives in operational, embedded decision-making by accountable individuals, not in a standing committee. The council coordinates; the owners and stewards decide and act. When an organization's governance is mostly meetings, it has confused the two, and that confusion is why so many programs feel busy while achieving little.

6. Domain ownership and data as a product

The most influential recent shift in governance thinking addresses a real weakness of traditional models: a central governance team, however federated, often sits far from the data and lacks the context to govern it well. The modern answer is domain ownership, most associated with the data mesh concept introduced by Zhamak Dehghani.

The core idea is to push data ownership to the domain teams closest to the data, the people who actually produce and understand it. The team that runs the payments system owns the payments data; the team that runs the product owns the usage data. They understand their data best, so they are best placed to govern its quality, meaning, and access, rather than a distant central team guessing at it.

Paired with this is a powerful reframing: data as a product. Instead of treating data as exhaust that flows out of systems for someone else to clean up, each domain treats the data it provides to others as a product it is responsible for, with the discipline that implies:

  • a clear owner (a data product owner),
  • defined quality and reliability the consumers can count on,
  • good documentation so others can find and understand it,
  • treating internal data consumers as customers to be served well.

This flips the incentive. Under the old model, the team creating data had little reason to care about its quality, that was someone else's downstream problem. Under data-as-a-product, the producing team is accountable for serving good data to its consumers, so quality is built in at the source rather than patched downstream, which connects directly to the 1-10-100 rule from the last lesson: fixing at the source is cheapest.

Domain ownership is not a rejection of governance but a redistribution of it: accountability moves to where the knowledge is, which the final step shows how to reconcile with the need for enterprise-wide consistency.

7. Reconciling autonomy and consistency

Domain ownership raises an obvious worry: if every domain governs its own data, do we not slide straight back into the decentralized chaos of five definitions of "customer"? The resolution is the key insight that makes modern governance work, and it echoes the federated model.

The answer is a division of labor sometimes called federated computational governance: domains own their data and agree to a set of global standards that make everything interoperate. Certain things are decided centrally and apply to everyone; everything else is left to the domains.

  • Global (agreed by all): shared definitions of cross-cutting concepts, common standards for quality and documentation, interoperability rules, and organization-wide policies like security and privacy classifications.
  • Local (owned by each domain): the specific data each domain produces, how they manage its quality day to day, and its internal details.

The "computational" part means these global standards are, wherever possible, built into the platform and automated, encoded as automatic checks and guardrails rather than enforced by committee review. Instead of a council manually approving each dataset, the platform automatically enforces that data meets the agreed standards, so consistency is achieved without a central bottleneck.

This is the synthesis the whole lesson builds to. The historic tension was control versus agility: centralize for consistency but be slow, or decentralize for speed but be chaotic. Domain ownership plus global standards plus automation dissolves the trade-off, local ownership and autonomy for agility, shared standards and automated enforcement for consistency.

The practical takeaway across all the models: governance is fundamentally about putting accountability where the knowledge and the work are, while keeping enough shared standard that the whole organization's data still fits together. Get that balance right, and you have the human structure that makes everything in the next lessons, quality, metadata, and control, actually happen.

8. Roles and the federated operating model

Business owners are accountable, stewards do the daily work, custodians safekeep the data technically; domains own their data locally while agreeing to automated global standards that keep everything consistent.

flowchart TD
  A["governance body: enterprise standards and coordination"] --> B["data owner: business accountability and decision rights"]
  B --> C["data steward: daily quality and definitions"]
  B --> D["data custodian: technical safekeeping in IT"]
  A --> E["global standards: shared definitions, quality, policy"]
  E --> F["automated enforcement in the platform"]
  B --> G["domain owns its data as a product"]
  F --> G

Check your understanding

The lesson ends with a 5-question quiz. Take it in the player above to see your score.

  1. What is the key difference between a data owner and a data steward?
    • They are the same role with different names
    • The owner is a senior business person accountable for the data and its rules; the steward is the hands-on expert who manages quality and definitions day to day
    • The owner is in IT; the steward is a database
    • The steward outranks the owner
  2. What does a data custodian do?
    • Decides what business terms mean
    • Sets policy and approves access
    • Handles the technical safekeeping, storing, protecting, backing up, and enforcing access in systems, usually in IT
    • Owns the data domain
  3. Why is the federated operating model the most common choice?
    • It requires no coordination
    • It removes the need for data owners
    • It is the cheapest to run
    • It balances the extremes: central standards give consistency while domains own their data for agility and local knowledge
  4. What is the 'committee trap' in data governance?
    • Governance collapses into endless meetings where no individual is accountable and real decisions are slow or never made
    • Having too few committees
    • Committees that make decisions too quickly
    • Letting IT run every meeting
  5. How does 'data as a product' with domain ownership improve data quality?
    • By moving all data to one central team
    • By eliminating data owners
    • The team that produces the data is accountable for serving good data to its consumers, so quality is built in at the source rather than patched downstream
    • By deleting low-quality data automatically

Related lessons

Business
intermediate

Why Principles Do Not Reach the Product

Almost every organisation has AI principles and almost none can point to a shipping decision they changed. This lesson covers why abstract commitments fail to bind, the specific gap between a value and a decision rule, ethics washing, and what a principle needs before it can affect anything.

8 steps·~12 min
Business
intermediate

What Actually Changes for a Manager

AI changes tasks rather than jobs, which means it redistributes work inside a role instead of removing the role. This lesson covers what that does to a manager: where review load lands, why self-reported productivity is unreliable, the skill-formation problem for junior staff, and which management assumptions stop holding.

8 steps·~12 min
Programming
intermediate

Standing Up the Function: First Ninety Days and Beyond

An AI QA function has to be built while features are already shipping. This lesson covers the order of work that produces something useful fast, where the function should sit, how to handle a team that resists a new gate, what to measure about the function itself, and the failure modes that quietly end it.

8 steps·~12 min
Business
intermediate

The AI Governance Function: What the Work Is and Who Does It

AI governance is a body of work before it is a job title, and most of it is done by people whose title says something else. This lesson sets out what the work consists of, how it splits across legal, risk, data protection and engineering, why a dedicated role appears at some scales and not others, and what the data protection officer precedent does and does not tell you.

9 steps·~14 min