DAMA Data Quality Specialist Questions and Answers
The best way to manage a data architecture roadmap is by using:
Options:
Integrated strategic reviews
An annual review
Senior management buy-in
Peer reviews
Results evaluation
Answer:
AExplanation:
Integrated strategic reviews provide the strongest mechanism for managing a Data Architecture roadmap because the roadmap must remain synchronized with enterprise strategy, business capability priorities, other architecture domains, projects, resources, and changing dependencies.
DAMA-DMBOK2 describes the Enterprise Data Architecture roadmap as the three-to-five-year path by which the target architecture becomes reality. Crucially, it states that this roadmap must be integrated into the overall Enterprise Architecture roadmap, including milestones, required resources, cost estimates, and business-capability workstreams. The roadmap should also reflect business requirements, current conditions, technical assessments, and organizational maturity.
This means roadmap management cannot sensibly be reduced to a once-a-year exercise. Strategic conditions, technology decisions, project sequencing, dependencies, and regulatory requirements may change throughout the roadmap period. Integrated reviews allow those changes to be evaluated in the context of the wider enterprise architecture rather than independently.
Senior-management support is necessary for authority and funding, while peer reviews and results evaluation are useful control activities. None, however, provides the same integrated strategic mechanism for keeping architectural direction aligned with enterprise priorities.
From a Data Quality perspective, such reviews also ensure that architecture changes preserve authoritative sources, lineage, integration controls, and enterprise quality requirements.
Reference Topics: DAMA-DMBOK2 Chapter 4 — Develop a Roadmap; Enterprise Architecture Integration; Data Dependencies; Lifecycle Reviews; Architecture Governance.
===============
Mapping requirements and rules for moving data from source to target enables:
Options:
Transformation
Backups
Load
Extract
Analysis
Answer:
AExplanation:
Source-to-target mapping enables Transformation. DAMA-DMBOK2 treats mapping as closely synonymous with transformation because a mapping defines how data in one source structure will be converted into the structure, format, representation, or value required by the target.
A mapping specification typically identifies the source attribute, target attribute, extraction conditions, target population rules, intermediate staging transformations, calculations, lookup requirements, and any changes required to make the source data conform to the target representation. DMBOK2 specifically explains that mapping sources to targets involves defining the rules for transforming information from one location and format into another.
Extraction simply retrieves data from the source. Loading places data into the target. Transformation is the activity that applies structural, syntactic, semantic, or value-level modifications between those stages.
The Data Quality connection is substantial. Mappings may standardize dates, convert units, harmonize codes, resolve reference values, remove duplicates, or enforce business rules. If mapping metadata is incomplete or incorrect, the transformation process can introduce rather than correct quality defects.
For this reason, source-to-target mapping should be governed, version-controlled, documented as metadata, traceable through lineage, and validated against agreed business definitions and Data Quality requirements.
Reference Topics: DAMA-DMBOK2 Chapter 8 — Map Data Sources to Targets; Transformation; ETL/ELT; Metadata Lineage; Chapter 13 — Data Cleansing and Standardization.
===============
A goal of reference and master data management is for data to ensure shared data is:
Options:
Secure, auditable, publicly available and free
Continuous, consistent, current and private
Complete, consistent, content and relevant
Secure, auditable, complete and relevant
Complete, consistent, current and authoritative
Answer:
EExplanation:
DAMA-DMBOK2 explicitly identifies a core goal of Reference and Master Data Management as ensuring that shared Master and Reference Data is complete, consistent, current, and authoritative across organizational processes.
Complete means required shared entities and attributes are sufficiently populated for their intended use. Consistent means equivalent data has compatible meaning and representation wherever it is consumed. Current means the information reflects an acceptably recent state. Authoritative means the organization recognizes a trusted source or governed process for determining the accepted value.
These characteristics are particularly important because Master and Reference Data is reused widely. An incorrect Product classification, Customer identifier, Country code, or Supplier status can therefore propagate defects across many systems and business processes.
DAMA also emphasizes that shared Master and Reference Data belongs to the organization rather than to a single application or department. This creates a strong requirement for enterprise stewardship and governance.
Reference and Master Data Management consequently interacts directly with Chapter 13. MDM can consolidate and distribute shared data, but it does not guarantee quality automatically. Matching, survivorship, validation, standardization, stewardship, and continuous monitoring are required to ensure the resulting records remain trustworthy.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Reference and Master Data Management Goals; Shared Data; Authoritative Sources; Stewardship; Chapter 13 — Completeness, Consistency and Currency.
===============
Data modelling tools and model repositories are necessary for:
Options:
Managing the enterprise data model in all levels
Designing and visualizing linkages between metadata, reference data and transactional data repositories
Visualizing and communicating database designs
Enabling data governance
Designing and visualizing organizational data artifacts
Answer:
AExplanation:
DAMA-DMBOK2 states that Data Modeling tools and model repositories are necessary for managing the Enterprise Data Model at all levels. This includes maintaining relationships among conceptual, logical, and physical representations and controlling how models created for different purposes fit into the wider enterprise architecture.
An enterprise modeling repository provides more than diagramming. It supports version control, definitions, lineage between model layers, comparison of changes, reuse of standard entities and attributes, enforcement of modelling conventions, and communication across projects. Many modelling tools also provide relationship and lineage capabilities that enable architects to trace how concepts represented at a high level are implemented in detailed models.
Option C describes one important capability—visualizing and communicating database designs—but it is narrower than the DMBOK2 purpose stated in the question. Similarly, modelling tools support governance and architectural visualization but are not solely established for either function.
For Data Quality, governed models provide the structural metadata from which rules can be derived. Keys support uniqueness; relationships support integrity; domains support validity; mandatory attributes support completeness; and standardized definitions support consistency.
The repository therefore becomes an important bridge between Data Architecture, Metadata Management, Data Governance, and Data Quality.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Data Modeling Tools and Repositories; Enterprise Data Model; Conceptual/Logical/Physical Models; Chapter 13 — Structural Quality Rules.
===============
A team uses a fishbone diagram during investigation of recurring missing customer email addresses. What is the primary purpose of the diagram?
Options:
Organize potential root causes into categories
Encrypt customer information
Calculate database storage
Create master records automatically
Answer:
AExplanation:
A fishbone diagram, also known as an Ishikawa or cause-and-effect diagram, is used to organize and explore potential root causes of a problem.
For recurring missing email addresses, categories might include People, Process, Technology, Data, Policy, Training, and External Sources. Potential causes could include optional form design, unclear business requirements, API mapping defects, missing supplier data, or inconsistent onboarding procedures.
The technique helps prevent teams from prematurely assuming that the most visible symptom is the actual cause. It encourages structured investigation across multiple dimensions of the process.
Once hypotheses are identified, evidence should be collected to confirm or reject them. The diagram itself does not prove causation.
In Data Quality programs, root-cause analysis is essential because repeated cleansing without process correction allows defects to recur. DAMA's Chapter 13 revision explicitly retains and strengthens the treatment of common causes and the improvement lifecycle.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Root-Cause Analysis; Cause-and-Effect Diagram; Issue Management; Continuous Improvement.
===============
According to the ISO/IEC 42010:2007 Software and Systems Engineering -Architecture Description, which of the following describes the definition of architecture:
Options:
The fundamental organisation of a system, and the principles governing its design and evolution
The fundamental responsibility for delivering the best systems at the lowest cost
The fundamental view of how the system should be built and how it will be maintained
The fundamental rules for ensuring the information captured in the architected solution is enforcing data quality and completeness
The fundamental collection of all artifacts that describes a system and how they work together
Answer:
AExplanation:
The correct definition is the fundamental organisation of a system, and the principles governing its design and evolution. DAMA-DMBOK2 incorporates the ISO/IEC 42010 architectural concept when establishing the theoretical basis for Data Architecture.
The fuller ISO formulation describes architecture in terms of the fundamental organization of a system embodied in its components, their relationships to one another and to the environment, together with the principles governing design and evolution. DAMA uses this concept to distinguish true architecture from a collection of diagrams or implementation specifications.
Option E is therefore incomplete: architectural artifacts document architecture, but the artifacts themselves are not the architecture. Option C focuses too narrowly on implementation and maintenance, while options B and D describe potential objectives or constraints rather than the definition.
Applied to enterprise data, this means architecture establishes the high-level organization of data assets, information flows, platforms, relationships, and guiding principles by which the data environment evolves.
Data Quality benefits because architectural principles establish where authoritative information originates, how it moves, how duplication is controlled, and where governance and quality controls must operate.
Reference Topics: DAMA-DMBOK2 Chapter 4 — Architecture Definition; ISO/IEC 42010; Enterprise Data Architecture; Architectural Principles; Chapter 13 — Quality by Design.
===============
Database monitoring tools measure key database metrics, such as:
Options:
Create, read, normalization, user access
Capacity, availability, backup instances, data quality
Create, read, update, delete
Capacity, design, normalization, user access
Capacity, availability, cache performance, user statistics
Answer:
EExplanation:
DAMA-DMBOK2 identifies capacity, availability, cache performance, and user statistics as representative metrics captured by database-monitoring tools. Database monitoring is an operational control mechanism used by Database Administrators and platform teams to understand whether database infrastructure is available, adequately sized, responsive, and being used as expected. DMBOK2 explicitly describes monitoring tools as automating the observation of metrics such as capacity, availability, cache performance, and user statistics.
Capacity metrics indicate resource consumption and future growth requirements. Availability measures whether data services remain accessible when required. Cache-performance measures help identify inefficient access patterns and bottlenecks, while user statistics provide information about workload and database consumption.
The other choices mix design concepts, CRUD operations, or Data Quality concepts with operational monitoring measures. For example, normalization is a modelling technique rather than a routine runtime performance metric. Similarly, “create, read, update, delete” describes basic data operations rather than monitoring indicators.
Although database performance and Data Quality are distinct disciplines, poor operational performance can affect Data Quality dimensions such as timeliness and availability. Monitoring therefore supports the technical environment in which governed, reliable information is delivered to users and applications.
Reference Topics: DAMA-DMBOK2 — Data Storage and Operations; Database Operations; Capacity and Availability Management; Monitoring; Chapter 13 — Timeliness and Operational Fitness for Purpose.
===============
A relationship that allows an address to be used by multiple people, and each person can have multiple addresses, can be resolved:
Options:
With an additional relationship describing the address usage
With a partnership entity called Person Address Usage and two, 'one to many' relationships
With an associative entity called Person Address Usage and two, 'one to many' relationships
By changing the primary keys on Person and Address to ensure referential integrity
By changing the role names of the foreign keys on Person and Address to ensure referential integrity
Answer:
CExplanation:
The scenario describes a many-to-many relationship: one Person may use several Addresses, while one Address may be associated with several Persons. In relational modelling, this should be resolved using an associative entity—here, Person Address Usage—between Person and Address. The original many-to-many relationship is thereby replaced with two one-to-many relationships. DAMA-oriented references identify this approach directly under the addition of associative entities.
The associative entity normally contains foreign keys referencing both parent entities and may also contain attributes describing the association itself, such as address type, usage purpose, effective date, end date, or primary-address indicator.
Simply modifying the primary keys of Person or Address does not resolve the semantic relationship. Changing foreign-key role names also changes terminology rather than cardinality. The term “partnership entity” is not the appropriate modelling construct.
This design has important Data Quality consequences. It permits explicit integrity constraints between the association and both parent entities, prevents ambiguous repeated columns, and allows the organization to validate rules such as valid effective periods and permitted address-use types.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Relationships; Cardinality; Associative Entities; Resolving Many-to-Many Relationships; Chapter 13 — Referential Integrity and Consistency.
===============
What is one of the Data Architecture artifacts which are usually captured within development projects, and then standardised and managed by data architects?
Options:
Business models
Enterprise data model
Data models
Company wide architectural blueprints
A Data Architecture roadmap
Answer:
CExplanation:
The correct answer is Data models. DAMA-DMBOK2 explicitly explains that data models and other Data Architecture artifacts are commonly created or captured within development projects and are subsequently standardized and managed by Data Architects.
Development projects routinely produce conceptual, logical, and physical representations of the data required by a solution. If each project develops those structures independently without architectural review, inconsistent terminology, duplicated entities, conflicting definitions, and incompatible relationship structures can accumulate across the enterprise. Data Architects therefore review project models, reconcile them with enterprise standards, and identify structures suitable for reuse.
An Enterprise Data Model is itself an important architectural artifact, but the wording of the question refers specifically to artifacts typically generated in individual development projects and later standardized. The DMBOK2 passage uses data models in precisely that context.
This process has a direct Data Quality benefit. Standardized models improve consistency of definitions, domains, keys, relationships, and constraints. They also provide metadata from which integrity, uniqueness, validity, and completeness rules can be derived.
Reference Topics: DAMA-DMBOK2 Chapter 4 — Data Architecture; Architectural Artifacts; Development Projects; Chapter 5 — Data Modeling and Design; Chapter 13 — Consistency and Integrity.
===============
A company has five systems that can update a customer's legal name, and no system is designated as authoritative. What is the greatest Data Quality risk?
Options:
Inconsistent values across systems
Excessive encryption
Insufficient disk compression
Slow printer performance
Answer:
AExplanation:
The principal risk is inconsistent customer-name values across systems. When multiple applications can independently update the same business fact without an authoritative source or coordinated Master Data process, conflicting representations are likely to emerge.
The problem is architectural and governance-related as much as it is a Data Quality issue. The organization must determine where the authoritative value is created, who is responsible for approving changes, how updates propagate, and how conflicts are resolved.
DAMA's public framework emphasizes a single trusted or consistent approach to shared core entities through Reference and Master Data Management.
The relevant Data Quality dimensions include Consistency, Accuracy, Currency, and potentially Uniqueness. An application may hold a technically valid legal name but still contain an outdated or conflicting value.
Metadata and lineage should document which systems create, master, distribute, or merely consume the attribute.
Reference Topics: DAMA-DMBOK2 — Authoritative Sources; Master Data Management; Data Architecture; Consistency; Lineage; Stewardship.
===============
An organization changes the definition of "Net Revenue". Before implementing the change, it wants to identify every report, transformation, and analytical model that depends on the existing definition. Which capability is most important?
Options:
Metadata lineage and impact analysis
Database compression
Password rotation
Physical server inventory
Answer:
AExplanation:
The required capability is Metadata lineage and impact analysis. Changing a business definition can affect far more than the glossary entry itself. The organization needs to know which physical fields, transformation rules, ETL jobs, calculations, reports, dashboards, analytical models, and downstream datasets depend on the current definition.
Lineage shows where data originates and how it flows and transforms. Impact analysis follows those relationships forward to determine what will be affected by a proposed change.
This prevents a common Data Quality failure in which a definition changes in one area while downstream systems continue applying the old interpretation. The result can be technically correct calculations that are semantically inconsistent.
A governed change should therefore update the business glossary, transformation metadata, Data Quality rules, reporting definitions, and affected documentation in a coordinated manner. Appropriate Data Stewards and Data Owners should approve the new meaning and effective date.
DAMA's framework describes Metadata Management as enabling understanding through definitions, lineage, and usage across systems, while the revised Chapter 13 explicitly strengthens the Data Quality–Metadata relationship.
Reference Topics: DAMA-DMBOK2 — Metadata Management; Data Lineage; Impact Analysis; Business Glossary; Data Quality Consistency.
===============
Business continuity is an aspect of Governance. What should a business continuity plan include?
Options:
Precedes business rules
Explains to external stakeholders why performance expectations are not being met
Outlines how a business will continue operating during an unplanned disruption in service
Provides explanation to customers during an unplanned disruption in service
Defines unplanned disruptions that may occur
Answer:
CExplanation:
A Business Continuity Plan is fundamentally concerned with maintaining essential business operations when normal services are disrupted. Therefore, the correct response is that it outlines how the business will continue operating during an unplanned disruption.
Within the DAMA-DMBOK2 framework, continuity concerns interact particularly with Data Governance, Data Storage and Operations, Data Security, and operational metadata. The DMBOK2 index explicitly identifies Business Continuity Plans and associates them with operational resilience and recovery concepts. A continuity plan must address how critical processes, information assets, personnel, technology, dependencies, and recovery arrangements will sustain or restore required business capability. The established definition of a BCP likewise centers on continued operation during unplanned service disruption.
Option E is incomplete: identifying potential disruptions belongs to risk assessment and continuity planning, but merely listing disruptions does not explain how operations will continue. Options B and D concern communications rather than continuity itself.
From a Data Quality perspective, continuity controls also protect availability, timeliness, integrity, and reliability. Recovery arrangements must ensure that restored data is complete, current enough for business use, and reconciled against authoritative sources after disruption.
Reference Topics: DAMA-DMBOK2 — Data Governance; Data Storage and Operations; Business Continuity; Disaster Recovery; Backup and Recovery; Chapter 13 — Timeliness, Integrity and Fitness for Purpose.
===============
Big data management requires:
Options:
More discipline than relational data management
Less discipline than relational data management
Big ideas with big budgets
No discipline at all
A certification in data science
Answer:
AExplanation:
DAMA-DMBOK2 establishes the guiding principle that Big Data management requires more discipline than relational data management. The reason is not simply data volume. Big Data environments combine large volumes with substantial variation in structure, source, velocity, semantics, and reliability. These characteristics increase the probability of uncontrolled duplication, inconsistent interpretation, poor provenance, inappropriate use, and undetected quality defects. DAMA's Big Data guidance specifically emphasizes disciplined management of metadata describing Big Data sources, their origins, and their value.
From a Data Quality perspective, scale does not make data fit for purpose. Large datasets still require defined quality expectations, profiling, monitoring, lineage, controlled transformations, ownership, and governance. Metadata becomes particularly important because analysts must understand where datasets originated, how they were transformed, their meaning, and whether they are appropriate for a particular analytical objective.
Data Governance must consequently establish accountability, acceptable-use policies, security requirements, quality expectations, and escalation mechanisms. Statistical and automated controls are also essential because manual inspection becomes impractical at Big Data scale.
Neither larger budgets nor data-science certification substitutes for sound data-management discipline. Likewise, reducing controls because data is large creates greater—not lower—operational and analytical risk.
Reference Topics: DAMA-DMBOK2 — Big Data and Data Science; Big Data Management Principles; Metadata Management; Data Governance; Chapter 13 — Fitness for Purpose and Data Quality Controls.
===============
Who should normally define the business meaning and acceptable quality requirements for a critical customer attribute?
Options:
A Business Data Steward working with relevant stakeholders
The network administrator alone
The database optimizer alone
The backup operator alone
Answer:
AExplanation:
A Business Data Steward, working with relevant business stakeholders and governance bodies, is the appropriate role to define or coordinate business meaning and quality expectations.
Quality requirements cannot be derived solely from database structures. The organization must understand what the attribute means, how it is used, what values are acceptable, which source is authoritative, and what consequences arise when it is wrong.
Data Stewards bridge business knowledge and formal Data Management controls. They commonly participate in maintaining glossary definitions, defining business rules, resolving issues, clarifying ownership, and establishing Data Quality expectations.
Technical specialists contribute implementation knowledge. A DBA may identify datatype constraints, while an integration specialist can implement transformations. Neither should independently determine business meaning.
DAMA's public framework describes Metadata Management as supporting definitions, lineage, and governance, while Reference and Master Data Management ensures consistency in shared core entities.
The strongest operating model therefore combines business accountability with technical execution.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Data Stewardship; Chapter 11 — Business Metadata; Chapter 13 — Data Quality Requirements; Critical Data Elements.
===============
A report displaying birth date contains possible, but incorrect values. What is a possible explanation?
Options:
Birth date is populated from two source systems, both of which record the birth date in the birth date field
Birth date is populated from a single source system, which does not contain birth date
Birth date is populated from a single source system, which contains missing values
Birth date is populated from a single source system, where the date field is an offset value of 1601
Birth date is populated from two source systems, one of which stores marriage date in the birth date field
Answer:
EExplanation:
The critical wording is “possible, but incorrect values.” This describes values that satisfy basic syntactic or domain validation—they look like legitimate dates—but do not accurately represent the real-world attribute defined by the field.
If two systems contribute data and one maps marriage date into the birth-date field, the resulting values can be perfectly valid calendar dates while being semantically incorrect as birth dates. This is principally an Accuracy defect, because DAMA defines accuracy in terms of how correctly data represents the real-world object or event it is intended to describe. It may also expose a consistency and integration-mapping problem between source systems. DAMA's quality framework distinguishes accuracy from completeness: data can be populated and formally valid while still being factually wrong.
Missing values would primarily produce a Completeness defect rather than populated-but-incorrect values. Two correctly mapped systems would not inherently explain the problem. An offset or technical date representation could create transformation problems, but the scenario most directly illustrates semantic mis-mapping between data elements.
The appropriate remediation is therefore not simple cleansing alone. Metadata mappings, source-to-target specifications, lineage, business definitions, and integration rules should be corrected at the root cause.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Accuracy, Completeness and Consistency; Root-Cause Remediation; Data Profiling; Metadata Management; Data Integration and Interoperability.
===============
An effective Data Governance communication program should include the following:
Options:
Regular newsletters
All answers
Events that encourage informal networking
A custom training program
A Data Governance Portal
Answer:
BExplanation:
An effective Data Governance communication program should employ multiple complementary communication mechanisms, making All answers correct. Governance changes how people define, create, use, approve, and resolve issues with data; consequently, sustained adoption requires more than publishing policies.
Regular newsletters keep stakeholders aware of progress, decisions, metrics, and upcoming activities. A Data Governance Portal provides a persistent location for policies, standards, stewardship information, glossaries, issue processes, and supporting materials. Custom training develops the capabilities required for individuals to understand their specific governance responsibilities. Informal networking events help establish relationships across business and technical groups, which is particularly important when resolving data ownership and definition conflicts.
DAMA-DMBOK2 treats communication and organizational change as core implementation considerations because governance depends on participation across functions rather than on a single technical team. Published CDMP material for this item identifies the combined response—newsletters, portal, training, and networking—as the intended answer.
For Data Quality, communication ensures that quality definitions, issue-management procedures, stewardship responsibilities, thresholds, and remediation decisions are understood and consistently applied.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Governance Communications; Organizational Change; Training; Data Governance Portal; Stewardship Engagement; Chapter 13 — Data Quality Culture.
===============
A customer's date of birth is stored as 12/04/1980. The value conforms to the database datatype and accepted date format, but the customer's verified date of birth is 21/04/1980. Which Data Quality dimension is primarily violated?
Options:
Completeness
Accuracy
Validity
Timeliness
Answer:
BExplanation:
The primary failure is Accuracy. Accuracy concerns whether a stored value correctly represents the real-world entity, event, or fact it is intended to describe. The recorded date is syntactically acceptable and conforms to the expected datatype and format, so it may satisfy Validity while still being factually wrong.
This distinction is fundamental in Data Quality assessment. A validity rule can determine whether a value is structurally permissible—for example, whether a date conforms to an approved pattern or whether a month lies between 1 and 12. It cannot by itself prove that the date belongs to the correct person. Government guidance based on DAMA dimensions makes the same distinction: a value can be valid while remaining inaccurate.
The strongest accuracy control would compare the value with an authoritative or independently verified source. Metadata should document the authoritative source and lineage, while governance should assign responsibility for resolving discrepancies.
The scenario therefore demonstrates why Data Quality programs must avoid assuming that format compliance equals correctness.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Accuracy; Validity; Authoritative Sources; Data Quality Rules; Metadata Lineage.
===============
The loading of country codes into a CRM is a classic:
Options:
Transaction data integration
Reference data integration
Analytics data integration
Fact data integration
Master data integration
Answer:
BExplanation:
Loading standardized country codes into a CRM is a classic example of Reference Data Integration. Country codes constitute controlled reference values used to classify or characterize other data. They differ from transactional information such as orders and payments and from master entities such as individual customers or products.
DAMA treats countries, currencies, geographic classifications, status codes, and similar controlled domains as typical Reference Data. The uploaded question therefore tests the distinction between integrating reusable code sets and integrating Master or Transaction Data. Published versions of the same DAMA item confirm Reference Data Integration as the intended answer.
In practice, an authoritative reference source may provide standardized country identifiers to CRM, ERP, MDM, analytical, and integration systems. Central management prevents different applications from independently inventing or changing equivalent codes.
This is also a Data Quality control. Reference integration improves Validity by restricting values to approved domains and Consistency by ensuring that different systems interpret equivalent countries identically. For example, an uncontrolled mixture of GB, UK, 826, and proprietary internal codes may create integration and reporting errors unless mappings are formally governed.
Metadata should document the authoritative source, code meaning, effective dates, mappings, ownership, and applicable standards.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Reference Data; Code Sets; Reference Data Integration; Authoritative Sources; Chapter 13 — Validity and Consistency.
===============
The Family Name of a Person is recorded in a system. The column name is Pname. Pname is an example of:
Options:
Data
Normalised data
Poor table design
Metadata
Megadata
Answer:
DExplanation:
Pname is Metadata because it is the name assigned to a database column and therefore describes the structure in which the actual Family Name values are stored. A value such as Smith or Patel would be data; Pname is structural information describing that data.
DAMA-DMBOK2 defines Metadata broadly as information about data and data-related processes. Technical metadata includes physical database characteristics such as table names, column names, datatypes, lengths, indexes, constraints, and related structural properties. The interpretation of Pname as metadata is also consistent with the published version of this exact certification question.
The fact that Pname may be an unclear abbreviation does not change its classification. It may represent a poor naming convention from a usability or governance perspective, but technically it remains metadata. A stronger physical name might be FamilyName, while the metadata repository could additionally map that field to the authoritative business glossary term.
This distinction matters to Data Quality because rules are frequently attached to metadata objects. For example, the Family Name attribute might have completeness, permitted-character, maximum-length, and standardization requirements.
Effective Metadata Management connects the technical column to its business meaning, system lineage, stewardship, and applicable Data Quality controls.
Reference Topics: DAMA-DMBOK2 Chapter 11 — Technical Metadata; Database Metadata; Business Glossary; Chapter 13 — Data Quality Rules and Metadata.
===============
A retail system accepts an order for 750,000 identical office chairs from an individual consumer. The value passes all datatype, domain, and mandatory-field checks. Which additional Data Quality dimension should detect the anomaly?
Options:
Completeness
Reasonableness
Uniqueness
Currency
Answer:
BExplanation:
The appropriate dimension is Reasonableness. A value may comply with technical and business-domain constraints while still being implausible in its operating context. An individual customer purchasing 750,000 office chairs is technically possible, but it is sufficiently unusual that it should trigger review.
Reasonableness controls evaluate whether values and combinations of values fall within credible expectations. These controls often depend on business context, historical behaviour, peer comparisons, statistical limits, or relationships between multiple fields.
The current DMBOK2 revision standardizes the term Reasonableness within its nine standard dimensions. The broader DMBOK2 dimension framework links reasonableness to whether data should be regarded as credible within the operational context.
Useful controls might compare order quantity against customer type, historical maximums, product availability, typical basket size, or statistical deviation from normal purchasing behaviour.
Reasonableness rules are especially valuable for identifying errors that conventional validation cannot detect. However, unusual data should not automatically be treated as incorrect; it should usually be flagged for investigation.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Reasonableness; Statistical Validation; Exception Detection; Business Rules.
===============
Emergency contact phone number would be found in which master data management program?
Options:
Product
Asset
Employee
Answer:
CExplanation:
An emergency contact phone number belongs to the Employee master-data domain. Employee Master Data describes relatively stable information about employees that needs to be consistently identified and reused across organizational processes and systems. Contact details associated with an employee—including emergency-contact information—therefore fall naturally within an Employee MDM program rather than Product or Asset MDM.
The broader DAMA Master Data framework identifies employees as one of the recurring enterprise entities that may require coordinated management alongside customers, products, suppliers, locations, and similar core subjects. Published versions of this specific DAMA certification item likewise identify Employee as the correct domain.
Managing this information centrally or through coordinated master-data processes helps prevent conflicting phone numbers, obsolete emergency contacts, duplicate employee identities, and inconsistent values across Human Resources and other authorized systems.
The data is also sensitive personal information. Governance should therefore define ownership, access rights, retention rules, and permitted usage. Quality controls should address completeness, validity, currency, and accuracy, because outdated emergency-contact information may fail precisely when it is needed.
Metadata Management should record definitions and classifications so that “emergency contact phone number” is consistently distinguished from an employee's own mobile or work telephone number.
Reference Topics: DAMA-DMBOK2 Chapter 10 — Master Data Domains; Employee Master Data; Data Governance; Chapter 13 — Accuracy, Completeness and Currency.
===============
A Data Quality incident affects several downstream systems. What should be determined before individual teams independently correct their copies?
Options:
The root cause and authoritative point of remediation
Which team has the largest budget
Which system contains the most records
Which copy is easiest to delete
Answer:
AExplanation:
The organization should determine the root cause and authoritative remediation point before multiple teams independently modify downstream copies.
If a defective value originates in a source system and is distributed to five downstream applications, correcting those five copies independently may provide temporary relief but leaves the upstream defect intact. The incorrect value may simply be redistributed again.
Lineage should be used to trace the defect through source systems, transformations, integration layers, Master Data hubs, warehouses, and reports. Once the origin is understood, governance can determine which system or process is authoritative and where correction should occur.
This approach minimizes inconsistent local fixes and reduces the risk that different teams apply incompatible interpretations of the same issue.
DAMA's Chapter 13 revision explicitly strengthens the relationship between Data Quality and Metadata Management, Reference/Master Data Management, Modeling, and Data Integration. Those connections are precisely what enable cross-system issue resolution.
Downstream correction may still be required after the authoritative source is fixed, but it should be coordinated.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Root-Cause Analysis; Issue Management; Metadata Lineage; Authoritative Sources; Data Integration.
===============
A customer table contains 50,000 records. Mandatory Tax Identification Number values are missing from 2,500 records. Which metric most directly measures the defect?
Options:
Consistency percentage
Completeness percentage
Uniqueness percentage
Reasonableness percentage
Answer:
BExplanation:
The defect is measured using Completeness. Completeness evaluates whether all required records or values that should be present are actually present. In this scenario, the mandatory Tax Identification Number is absent from 2,500 of 50,000 customer records.
A straightforward attribute-level completeness metric would be calculated as the number of populated required values divided by the number expected. Therefore, 47,500 of 50,000 records contain the required value, producing a completeness result of 95%.
Completeness does not establish that populated values are accurate. A record may contain a Tax Identification Number and therefore pass the completeness rule while containing the wrong number. This separation between completeness and accuracy is essential when designing Data Quality scorecards. DAMA-aligned guidance defines completeness as the presence of required values and explicitly distinguishes it from factual correctness.
Governance should also determine whether the attribute is genuinely mandatory for every customer type. Quality rules must reflect business applicability rather than blindly requiring every field for every record.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Completeness; Data Quality Metrics; Business Rules; Critical Data Elements; Scorecards.
===============
In data modelling practice, entities are linked by:
Options:
Indexes
Triggers
Cardinality
Relationships
Processes
Answer:
DExplanation:
In a data model, entities are linked by relationships. An entity represents a distinguishable business concept or thing about which the organization stores information, while a relationship expresses how one entity is associated with another.
For example, a Customer places an Order, an Order contains Order Lines, and a Product appears on an Order Line. Relationships therefore capture business semantics rather than merely physical implementation details. Cardinality is an important characteristic of a relationship because it specifies how many instances of one entity may or must be associated with instances of another; however, cardinality is not itself the general mechanism by which entities are linked. DAMA-oriented modelling material identifies relationships as the correct linkage concept.
Indexes and triggers are physical database mechanisms. Processes describe business activities and may interact with entities, but they do not define entity-to-entity structure within a data model.
The Data Quality implications are significant. Properly defined relationships establish expectations for referential integrity. For example, an Order referencing a nonexistent Customer indicates an integrity defect. Profiling can test orphaned foreign keys, invalid relationship cardinalities, and inconsistent associations.
Relationships and their cardinalities should therefore be documented as metadata and enforced through appropriate database, application, or quality controls.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Entities, Relationships and Cardinality; Logical Data Modeling; Chapter 13 — Integrity and Consistency; Metadata Management.
===============
Periodic archiving of transaction data from a production CRM system is critical for:
Options:
Training junior DBAs
Providing alternate sources for reporting systems
Managing deleted customer records
The maintenance of database performance
Enabling the distribution of transaction data across the enterprise
Answer:
DExplanation:
Periodic archiving is critical for maintaining database performance. As a production CRM accumulates historical transactions, active tables and indexes can become increasingly large. This increases storage consumption, backup duration, index-maintenance overhead, and the quantity of data that database engines must process during operational queries.
DAMA-DMBOK2 treats archiving as an important Data Storage and Operations activity. Historical information that remains subject to retention requirements but is no longer frequently needed for operational processing can be moved to suitable archival storage. DAMA-aligned guidance for this scenario specifically links periodic transaction archiving with maintaining production database performance.
Archiving is not the same as arbitrary deletion. Retention policies, legal obligations, recovery requirements, auditability, and business value determine how long data must remain accessible and where it should be stored. The archive must also be recoverable and appropriately secured.
Data Quality implications include maintaining integrity and traceability during migration to the archive. Records should remain complete, relationships should be preserved, and metadata should indicate retention status and archival location.
Providing reporting sources or managing deleted customers may be secondary considerations, but neither is the principal purpose described in the question.
Reference Topics: DAMA-DMBOK2 Chapter 6 — Data Storage and Operations; Archiving; Database Performance; Retention; Chapter 13 — Integrity and Historical Data.
===============
Profiling reveals that an Age field has a minimum value of -7 and a maximum value of 263. What should the Data Quality team do first?
Options:
Delete both records immediately
Treat the observations as potential exceptions and investigate the governing business rules
Increase the database datatype range
Replace both values with the dataset average
Answer:
BExplanation:
The values should first be treated as potential exceptions requiring investigation. Profiling identifies anomalies; it does not automatically prove that every unusual value is incorrect or specify the appropriate remediation.
Negative age and 263 years are highly improbable for a conventional person-age attribute, suggesting a validity or reasonableness defect. However, the Data Quality team should confirm the business definition, unit of measure, derivation logic, source mappings, special-value conventions, and intended population before changing the records.
Blindly replacing values with an average introduces fabricated data and destroys traceability. Deleting the records may also remove legitimate information elsewhere in the record. Increasing the datatype range addresses technical storage rather than semantic correctness.
A disciplined Data Quality process moves from profiling to rule validation, issue assessment, root-cause analysis, remediation, and monitoring. The business rule may ultimately specify an acceptable age range, but that rule should be governed and documented rather than inferred from individual anomalies.
Metadata is essential because it should describe whether Age is stored directly, calculated from Date of Birth, or encoded using another convention.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Profiling; Validity; Reasonableness; Root-Cause Analysis; Data Quality Rules; Metadata.
===============
A Data Quality team has identified 500 defects across several domains. Which factor should have the strongest influence on remediation priority?
Options:
Alphabetical order of the affected attributes
Business impact and risk
Which database contains the fewest records
Which issue was easiest to describe
Answer:
BExplanation:
Remediation should principally be prioritized according to business impact and risk. Not all defects have equivalent consequences, even when they occur at similar frequencies.
A defect affecting regulatory reporting, customer payments, safety-critical operations, executive reporting, or high-value master data may require immediate remediation. A larger number of defects affecting a low-impact optional field may legitimately receive lower priority.
A robust prioritization model may consider financial loss, regulatory exposure, operational disruption, customer impact, reputational damage, number of dependent systems, recurrence rate, remediation cost, and whether a Critical Data Element is involved.
This risk-based approach prevents Data Quality programs from becoming simple defect-count reduction exercises. The objective is not merely to maximize the number of corrected records but to improve fitness for purpose where poor data creates material consequences.
Governance should approve prioritization criteria and resolve conflicts where different business areas assign different importance to the same issue. Metadata and lineage provide evidence about downstream dependencies and affected processes.
DAMA's revision of Chapter 13 adds a clearer Critical Data Element concept and clarifies responsibility within the Data Quality Improvement Lifecycle, reinforcing risk-based prioritization.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Issue Prioritization; Business Impact; Critical Data Elements; Risk; Remediation.
===============
The primary reason for identifying Critical Data Elements is to:
Options:
Apply identical controls to every field in every system
Focus governance and quality resources on data with the greatest business impact
Eliminate the need for metadata
Replace all non-critical data
Answer:
BExplanation:
Critical Data Elements are identified so that organizations can concentrate governance and Data Quality effort where failure would create the greatest business impact.
Not every attribute deserves the same level of profiling, monitoring, stewardship, lineage documentation, control, and remediation. Applying identical controls to every field would usually be economically inefficient and may divert resources away from information that supports regulation, financial reporting, key operations, customer outcomes, or strategic decisions.
The DMBOK2 maintenance revision explicitly adds and clarifies the concept of a Critical Data Element within Chapter 13.
Once a CDE is identified, the organization can define applicable quality dimensions, measurable rules, thresholds, ownership, authoritative sources, lineage, issue escalation, and monitoring frequency.
Criticality is contextual. An attribute may be critical to one process and relatively insignificant to another. The decision should therefore reflect documented business use and risk rather than technical prominence.
Metadata remains essential because it links the CDE to definitions, systems, lineage, owners, and quality controls.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Critical Data Elements; Prioritization; Business Impact; Risk; Data Governance; Metadata.
===============
Who does the DMBoK consider to be generally responsible for developing business glossary content?
Options:
Business Data Stewards
Conceptual Data Modellers
Data Architects
Business Users
Coordinating Data Stewards
Answer:
AExplanation:
DAMA-DMBOK2 assigns primary responsibility for business glossary content to Business Data Stewards. Business Data Stewards are typically subject-matter experts who understand how data is defined, created, interpreted, and consumed within their business domain. DMBOK2 explicitly states that Data Stewards are generally responsible for business glossary content and identifies Business Data Stewards as professionals who work with stakeholders to define and control data.
A business glossary is not simply a technical dictionary. It establishes agreed business terminology, definitions, synonyms, business rules, responsible stewards, and relationships between business concepts. This makes stewardship involvement essential because definitions must reflect operational and business meaning rather than merely database structures.
Data Architects may contribute candidate definitions and structural context from subject-area and conceptual models, but they do not generally own the business meaning. Business users provide valuable input, while Coordinating Data Stewards help reconcile definitions across domains, yet the normal accountability remains with Business Data Stewards.
This relationship is particularly important for Data Quality. Quality rules depend on unambiguous definitions of data elements. If “Customer,” “Active Account,” or “Order Date” has inconsistent meanings, measurements of completeness, accuracy, and validity cannot be consistently interpreted.
Reference Topics: DAMA-DMBOK2 Chapter 3 — Develop a Business Glossary; Business Data Stewardship; Metadata Management; Chapter 13 — Data Quality Rules and Business Definitions.
===============
One way of defining ethics is:
Options:
Doing it right when no one is looking
Doing it wrong, and then expertly covering it up
Doing it wrong, and failing to covering it up
Doing it wrong, and then apologizing
Doing it right when someone is looking
Answer:
AExplanation:
The correct principle is “Doing it right when no one is looking.” The statement captures the essential distinction between ethical behaviour and mere compliance. Ethical conduct does not depend on observation, enforcement, or the probability of being caught; it requires responsible action because the action itself is appropriate. The answer wording is presented directly in the PDF on page 4.
DAMA-DMBOK2 treats data ethics as an essential component of professional Data Management because organizations routinely make choices about collecting, combining, analyzing, retaining, sharing, and monetizing data. A technically permissible activity may still create inappropriate effects on individuals or expose data to potential misuse.
Data ethics therefore extends beyond security and regulatory compliance. It requires consideration of the impact on people, potential for misuse, and economic value of data. Ethical handling also requires transparency concerning data provenance, intended use, quality limitations, and the consequences of decisions based on data.
Data Quality has an ethical dimension as well. Knowingly using inaccurate, incomplete, misleading, or poorly understood data in consequential decisions can harm customers and other stakeholders. Governance should therefore establish not only compliance controls but also principles for responsible use.
Reference Topics: DAMA-DMBOK2 Chapter 2 — Data Handling Ethics; Ethical Principles; Impact on People; Potential for Misuse; Chapter 13 — Reliable and Fit-for-Purpose Data.
===============
The seven Vs of Big Data are:
Options:
Volume, Velocity, Variety. Vocabulary, Viscosity. Volatility. Veracity
Volume, Vector. Variety. Vocabulary. Viscosity. Voiced, Veracity
Volume, Vector, Variety, Vocabulary, Viscosity, Volatility, Veracity
Volume, Vector, Variety, Vocabulary, Viscosity, Vexation. Veracity
Volume, Velocity, Variety. Variability, Viscosity. Volatility, Veracity
Answer:
EExplanation:
Within the DMBOK2 framing used by this certification material, the expanded characteristics are Volume, Velocity, Variety/Variability, Viscosity, Volatility, and Veracity. Option E is therefore the only choice that correctly contains the DAMA terms represented in the question. DMBOK-oriented study material explains that the original three Vs—Volume, Velocity, and Variety—were expanded to include Variability, Viscosity, Volatility, and Veracity in the broader characterization.
Volume concerns data quantity; Velocity concerns the rate of creation and processing; Variety/Variability concerns differing structures and changing representations; Viscosity describes difficulty in using or integrating the data; Volatility concerns how quickly usefulness or meaning changes; and Veracity addresses credibility and trustworthiness.
These characteristics have direct Data Quality implications. Increased variety creates semantic and structural consistency challenges. Velocity reduces the time available for traditional validation. Volatility affects currency and timeliness. Veracity is directly concerned with reliability.
DAMA's key point is that Big Data does not reduce the need for management discipline. Its scale and complexity increase the need for metadata, governance, automated profiling, lineage, and statistically driven quality controls.
Reference Topics: DAMA-DMBOK2 Big Data and Data Science — Big Data Characteristics; Volume; Velocity; Variety/Variability; Viscosity; Volatility; Veracity; Chapter 13 — Scalable Data Quality.
===============
A search engine database is populated by a web crawler or spider software that usually processes as:
Options:
Linearly independent, near exact (LINE)
Atomic, consistent, isolated, durable (ACID)
Continually repeating, always performing (CRAP)
Fundamentally available, consistently true (FACT)
Basically available, soft state, eventually consistent (BASE)
Answer:
EExplanation:
Web-scale search-engine repositories commonly operate according to BASE: Basically Available, Soft State, Eventually Consistent. BASE is associated with distributed and many NoSQL data architectures where availability, horizontal scalability, and tolerance of distributed processing are prioritized over the immediate transactional consistency expected from traditional ACID systems.
A web crawler collects and indexes content continuously across very large distributed environments. Different nodes or indexes may temporarily contain slightly different versions of information while updates propagate. The design is acceptable because the system is expected to converge toward consistency rather than guarantee that every replica reflects the identical state at every instant. DAMA-aligned material explicitly contrasts BASE with traditional ACID approaches for large-scale distributed data processing.
ACID—Atomicity, Consistency, Isolation, Durability—is generally associated with transaction-oriented relational systems where individual transactions require strict integrity guarantees. Search indexing has different workload characteristics and can tolerate temporary inconsistency.
The Data Quality implication is that “consistency” must be evaluated according to architectural context. Temporary divergence can be expected behavior in an eventually consistent platform rather than evidence of a defect, provided convergence occurs within acceptable business thresholds.
Reference Topics: DAMA-DMBOK2 — Data Storage and Operations; Distributed Databases; NoSQL; ACID versus BASE; Availability; Consistency.
===============
A minimal super key is:
Options:
Also known as a candidate key, it is a superkey without duplicated attributes.
Any set of attributes without duplicates that uniquely identifies an entity instance.
A synonym for a surrogate key.
A type of advanced index key structure, in the same family as Hash, Heap, B-Tree and Inverted.
Any set of attributes where each attribute that makes up the key is a foreign key in its own right.
An artificial key, made up of several meaningful components to help the reader understand the nature of the entity from the key alone.
Answer:
AExplanation:
A candidate key is a minimal super key. A super key is any set of attributes sufficient to uniquely identify an entity instance, but it may contain attributes that are unnecessary for uniqueness. A candidate key removes that redundancy: if any attribute is removed from the candidate key, the remaining attributes no longer uniquely identify the entity.
DAMA-DMBOK2 makes this distinction explicitly: a candidate key is a minimal set of one or more attributes identifying an entity instance, and “minimal” means that no subset of the candidate key can perform the same unique-identification function.
Option B is incomplete because a set can uniquely identify an entity while still containing unnecessary attributes; that would qualify as a super key but not necessarily a minimal super key. A surrogate key is different: it is an artificial identifier introduced primarily for technical identification. Foreign-key composition and physical index structures likewise do not define candidate-key minimality.
This concept has direct Data Quality implications. Correctly defined candidate and primary keys support uniqueness and integrity, prevent duplicate entity instances, enable reliable referential relationships, and improve entity matching within Master Data Management. Key definitions should also be captured as structural metadata so profiling and quality rules can consistently test duplicate and orphan conditions.
Reference Topics: DAMA-DMBOK2 Chapter 5 — Data Modeling and Design; Keys; Candidate Keys; Chapter 13 — Uniqueness and Integrity; Metadata Management; Master Data Management.
===============
A financial transaction is captured correctly at 9:00 AM but does not become available to the fraud-monitoring system until 6:00 PM, although the business requirement is availability within five minutes. Which Data Quality dimension is primarily violated?
Options:
Accuracy
Timeliness
Uniqueness
Completeness
Answer:
BExplanation:
The primary failure is Timeliness. The transaction may be completely accurate and complete, but it is not available within the period required by the consuming business process.
Timeliness evaluates whether data is available when needed for its intended use. The relevant threshold must therefore come from the business requirement rather than from an arbitrary technical target. In this scenario, the fraud-monitoring process requires the transaction within five minutes, while delivery occurs approximately nine hours later.
The root cause could exist in extraction frequency, integration queues, batch processing, network delays, source-system availability, or downstream ingestion. Lineage and operational metadata should be used to identify where the latency occurs.
Timeliness must also be distinguished from Currency. Currency asks whether information reflects a sufficiently recent real-world state; Timeliness asks whether data is delivered or available within the required period. A current transaction that arrives too late can therefore fail Timeliness even though the underlying value accurately represented reality when captured.
DAMA's revised DMBOK2 Chapter 13 recognizes both Timeliness and Currency as separate standard dimensions.
Reference Topics: DAMA-DMBOK2 Chapter 13 — Timeliness; Currency; Data Quality Requirements; Data Integration; Operational Monitoring.
===============
The implementation of a 'Master Data Repository’, which is integrated across the enterprise, is an example of which integration approach?
Options:
Hub and Spoke
Replication
Change Data Capture
Publish and Subscribe
Point to Point
Answer:
AExplanation:
An enterprise Master Data Repository serving multiple applications is a classic Hub-and-Spoke integration pattern. In this architecture, shared information is consolidated physically or virtually in a central hub, while participating applications interact with that hub rather than building a separate direct interface to every other application.
DAMA-DMBOK2 explicitly identifies Master Data Management hubs, Data Warehouses, Data Marts, and Operational Data Stores as familiar examples of data hubs. The hub-and-spoke model reduces the proliferation of point-to-point interfaces and can provide a consistent enterprise view of shared data.
This architecture is particularly appropriate for MDM because customer, product, supplier, location, or other master entities often need to be standardized, matched, governed, and redistributed across many systems. The central hub can apply survivorship rules, reference mappings, stewardship decisions, and quality controls before distributing trusted values.
Replication and Change Data Capture are techniques for moving or detecting changes in data, but neither describes the overall interaction topology. Publish-subscribe may operate alongside an MDM hub as a distribution mechanism, but the enterprise repository itself represents the hub within a hub-and-spoke architecture.
Reference Topics: DAMA-DMBOK2 Chapter 8 — Hub-and-Spoke; Integration Interaction Models; Chapter 10 — Master Data Management; Enterprise Master Data Hub; Data Quality Consistency.
===============
A data lake and a data warehouse are the same concepts in so far as:
Options:
They are not related in any way
They are both concerned with preparing data for reporting and analytics
They are both concerned with duplicating all the operational data
They are both concerned with producing star schemas and dimensional data models
They are both concerned with using different forms of data governance
Answer:
BExplanation:
Although Data Lakes and Data Warehouses have materially different architectures, both support the broader objective of making organizational data available for reporting, analysis, analytics, and decision support. Therefore, option B identifies their legitimate conceptual overlap.
A traditional Data Warehouse generally contains integrated, curated, structured data designed for repeatable Business Intelligence, reporting, and analytical workloads. A Data Lake typically retains larger volumes of raw or less-structured information and supports flexible processing, exploration, Data Science, and advanced analytics. Both can therefore participate in analytical data pipelines even though their approaches to schema, transformation, governance, and consumption differ.
They do not inherently duplicate every item of operational data. Nor must a Data Lake produce dimensional or star-schema models; that modeling approach is more closely associated with traditional warehouse implementations. Governance is required for both rather than being the defining difference between them.
From a Data Quality standpoint, warehouses typically apply substantial cleansing and conformity before consumption. Lakes may retain raw values and defer interpretation, which increases the importance of metadata, provenance, cataloging, profiling, and consumer awareness.
Reference Topics: DAMA-DMBOK2 — Data Warehousing and Business Intelligence; Big Data; Data Lakes; Analytical Data; Metadata; Chapter 13 — Data Preparation and Fitness for Purpose.
===============