Data Repository

A data repository is a named place — logical or physical — where an organization's information is stored, organized, and made available for use.

Definition

In business architecture, a data repository is the artifact that represents where information physically or logically resides in support of business capabilities and value streams. It sits at the intersection of business architecture and information/data architecture: business architects care about repositories not as technical storage engines, but as the containers that hold the information concepts (Customer, Product, Contract, Claim, and so on) that capabilities create, read, update, and consume. A repository can be as concrete as a customer master data platform or a claims database, or as logical as 'the system of record for policyholder data' before a specific technology has been chosen. The key boundary to hold onto is that a data repository is not the same as a database, a data warehouse, or an application. Those are technology implementations. The repository, in architecture terms, is the business-relevant concept of 'where this information lives and who owns it' — it can map to one database, span several, or be split across multiple systems that together represent a single logical repository. This distinction matters enormously in cross-mapping work: business architects map information concepts to repositories, and separately map repositories to the applications and technology platforms that implement them, giving leadership a clean line of sight from strategy to data to systems. Data repositories also differ from 'data domains' or 'information concepts' themselves. An information concept (e.g., Customer) is the business meaning of the data; the repository is where an instance of that concept is stored and governed. A single information concept frequently spans multiple repositories across an enterprise — which is precisely the redundancy and fragmentation problem that repository mapping is designed to expose.

Origin & Context

The concept draws from information architecture and data management practice long predating formal business architecture, but it was pulled into the BA discipline through the BIZBOK Guide's treatment of information mapping, where information concepts are cross-mapped to capabilities, value streams, and their underlying data stores. TOGAF's Data Architecture domain uses a closely related notion of data entities and data stores serving application and technology layers. Business architecture adopted 'data repository' as the bridge artifact that lets architects reason about information ownership and reuse without getting pulled into physical database design.

Why It Matters

CIOs and enterprise architects use repository mapping to find where the same information is duplicated across the estate — a direct driver of data quality issues, reconciliation costs, and failed regulatory reporting. Business architects use it to determine, during M&A integration or divestiture, which repositories must be consolidated, migrated, or kept separate, which materially shapes integration cost and risk. Data governance leaders rely on repository-to-capability mapping to assign clear data ownership and stewardship, closing the common gap where 'everyone touches the data, no one owns it.' Getting this right reduces redundant data investments and shortens the path from a strategic data initiative to an executable technology roadmap.

Common Misconceptions

Myth: A data repository is just another word for a database.
Reality: A database is a technology implementation; a data repository is the business-architecture representation of where information logically resides. One repository can be realized across several databases, and one database can house parts of several logical repositories — the mapping is deliberately technology-agnostic so the business view survives platform changes.
Myth: Mapping data repositories is a data architecture task, not something business architects need to touch.
Reality: Data architects own the physical and logical schema design, but business architects own the cross-mapping of information concepts and repositories to capabilities and value streams. Without that business-side mapping, data architecture decisions get made without visibility into which capabilities and stakeholders actually depend on the data — a frequent cause of governance and prioritization conflict.
Myth: Once you've documented your data repositories, the map is done.
Reality: Repository maps decay quickly as applications are retired, replaced, or consolidated. Treat the repository inventory as a living artifact tied to your capability model and application portfolio, reviewed whenever a major system change, acquisition, or data governance initiative occurs.

Practical Example

A regional insurer's business architecture team was asked to support a claims modernization initiative. The lead business architect first mapped the 'Claim' and 'Policyholder' information concepts against the capability map, then identified every repository touching those concepts — the legacy claims system, a regional policy administration platform, and two spreadsheet-based workarounds used by adjusters. The cross-mapping revealed that policyholder data existed in three separate repositories with no designated owner, causing adjusters to work from inconsistent contact records. Working with the data governance lead and the CTO's data architecture team, the business architect used this map to recommend consolidating policyholder data into a single system of record before building new claims capabilities on top of it — avoiding the common trap of automating a broken data foundation. The repository map became a standing reference for subsequent modernization phases.

Industry Applications

Financial Services
Mapping customer, account, and transaction data repositories to support Know Your Customer (KYC) obligations and to eliminate duplicate customer records across retail, wealth, and commercial lines of business.
Healthcare
Identifying which repositories hold patient, provider, and claims data to support interoperability initiatives and to clarify data ownership ahead of electronic health record consolidations.
Insurance
Cross-mapping policy, claims, and underwriting repositories during core system modernization to prevent new digital capabilities from being built on fragmented or duplicated data sources.
Manufacturing
Mapping product master and bill-of-materials repositories across ERP instances following an acquisition, to determine which system becomes the enterprise system of record.