VibeKoding / Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat / An Introduction to Data GovernanceAn Introduction to Data Governance
VK

An Introduction to Data GovernanceAn Introduction to Data Governance

๐Ÿ“š Ensiklopedia ยท Fondasi KuatEnsiklopedia ยท Fondasi Kuat ๐ŸŒ Dual Bahasa (ID / EN) โšก VibeKoding Native

Ensiklopedia VibeKoding: An Introduction to Data Governance.Ensiklopedia VibeKoding: An Introduction to Data Governance.

๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

Have you ever encountered this situation: the numbers on a report don't match the actual business, the same user's information is different across two systems, or analysis results are completely unreliable due to dirty data? Data governance is the systematic approach to solving these problems. In the era of "data-driven decision-making," data quality directly determines decision quality โ€” Garbage In, Garbage Out.Have you ever encountered this situation: the numbers on a report don't match the actual business, the same user's information is different across two systems, or analysis results are completely unreliable due to dirty data? Data governance is the systematic approach to solving these problems. In the era of "data-driven decision-making," data quality directly determines decision quality โ€” Garbage In, Garbage Out.

What will you learn in this article?What will you learn in this article?

After completing this chapter, you will gain:After completing this chapter, you will gain:

ChapterContentCore Concepts
Chapter 1Data Quality DimensionsCompleteness, accuracy, consistency, timeliness
Chapter 2Data Governance FrameworkOrganization, processes, technology, culture
Chapter 3Data Lineage TracingImpact analysis, root cause investigation, compliance auditing
Chapter 4Metadata ManagementTechnical metadata, business metadata, operational metadata
Chapter 5Data Layered ArchitectureODS, DWD, DWS, ADS
Chapter 6Governance Tools & PracticesGreat Expectations, dbt, DataHub

------

0. The Big Picture: Motivation for needing Data Governance0. The Big Picture: Motivation for needing Data Governance

Data governance is not a technical problem โ€” it's a management problem. It answers the core questions: Who is responsible for the data? What are the data standards? How do we ensure data remains trustworthy?Data governance is not a technical problem โ€” it's a management problem. It answers the core questions: Who is responsible for the data? What are the data standards? How do we ensure data remains trustworthy?

Imagine a company with 100 data tables, each maintained by different teams, with no unified naming conventions, no data dictionary, and no quality checks. The result: the same "monthly active users" metric yields 5 million from the marketing department and 3 million from the product department โ€” because the definitions are different.Imagine a company with 100 data tables, each maintained by different teams, with no unified naming conventions, no data dictionary, and no quality checks. The result: the same "monthly active users" metric yields 5 million from the marketing department and 3 million from the product department โ€” because the definitions are different.

๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

1. Organization: Define roles and responsibilities for data owners and data stewards 2. Processes: Establish standard procedures for data onboarding, changes, and decommissioning 3. Technology: Deploy data quality monitoring, metadata management, lineage tracing, and other tools 4. Culture: Make the entire company recognize that "data is an asset," not "data is a byproduct"1. Organization: Define roles and responsibilities for data owners and data stewards 2. Processes: Establish standard procedures for data onboarding, changes, and decommissioning 3. Technology: Deploy data quality monitoring, metadata management, lineage tracing, and other tools 4. Culture: Make the entire company recognize that "data is an asset," not "data is a byproduct"

------

1. The Six Dimensions of Data Quality1. The Six Dimensions of Data Quality

Data quality is not a vague concept โ€” it can be measured across six specific dimensions. Each dimension has clear definitions and detection methods.Data quality is not a vague concept โ€” it can be measured across six specific dimensions. Each dimension has clear definitions and detection methods.

DimensionDefinitionDetection MethodCommon Issues
CompletenessWhether data has missing valuesNull rate checkRequired fields are empty, related data missing
AccuracyWhether data is correctRule validation, sampling verificationNegative amounts, invalid dates
ConsistencyWhether multi-source data matchesCross-system comparisonCRM and order system have different usernames
TimelinessWhether data is updated promptlyUpdate time checkInventory data lagging, prices not synced
UniquenessWhether duplicate records existDeduplication checkSame user registered twice
ValidityWhether data conforms to format rulesRegex/range validationInvalid email format, negative age
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- $1: Validate data at the entry point to prevent dirty data from entering - $10: Clean existing dirty data in the data warehouse - $100: Loss from incorrect decisions caused by dirty data The earlier you detect and fix data quality issues, the lower the cost.- $1: Validate data at the entry point to prevent dirty data from entering - $10: Clean existing dirty data in the data warehouse - $100: Loss from incorrect decisions caused by dirty data The earlier you detect and fix data quality issues, the lower the cost.

------

2. Data Governance Framework: Full Lifecycle Management2. Data Governance Framework: Full Lifecycle Management

Data governance is not a one-time project but a continuous process that spans the entire data lifecycle. From data creation to destruction, each stage requires clear standards and responsible parties.Data governance is not a one-time project but a continuous process that spans the entire data lifecycle. From data creation to destruction, each stage requires clear standards and responsible parties.

StageCore OutputKey Roles
Define StandardsData dictionary, naming conventions, classification and grading standardsData Architect
Data IngestionOnboarding standards, validation rules, lineage recordsData Engineer
Storage ManagementLayered model, permission matrix, lifecycle policiesDBA / Platform Engineer
Data ConsumptionData catalog, masking rules, quality reportsData Analyst / Business Stakeholders
Archive & DestroyArchival policies, deletion records, audit logsSecurity & Compliance Team

2. Data Governance Framework2. Data Governance Framework

Data governance cannot be solved by simply buying a tool โ€” it requires a complete framework. The most commonly used reference framework in the industry is DAMA-DMBOK (Data Management Body of Knowledge).Data governance cannot be solved by simply buying a tool โ€” it requires a complete framework. The most commonly used reference framework in the industry is DAMA-DMBOK (Data Management Body of Knowledge).

Governance DomainCore ContentKey Outputs
Data ArchitectureDefine data models, data flows, storage strategiesData architecture diagrams, ER diagrams
Data StandardsUnified naming conventions, coding standards, metric definitionsData dictionary, metric library
Data QualityEstablish quality rules, monitoring alerts, remediation processesQuality reports, SLA dashboards
Data SecurityClassification and grading, access control, masking and encryptionSecurity policies, audit logs
Master Data ManagementUnified "golden records" for core entities like customers and productsMaster data hub
Data LifecycleManage the full process from data creation to archiving to destructionRetention policies, archival rules
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- Level 1 - Initial: No unified standards, each team operates independently - Level 2 - Repeatable: Basic documentation exists, but execution is inconsistent - Level 3 - Defined: Unified governance processes and tools exist, most teams comply - Level 4 - Managed: Quantified quality metrics and automated monitoring are in place - Level 5 - Optimizing: Continuous improvement, data governance integrated into daily development workflows- Level 1 - Initial: No unified standards, each team operates independently - Level 2 - Repeatable: Basic documentation exists, but execution is inconsistent - Level 3 - Defined: Unified governance processes and tools exist, most teams comply - Level 4 - Managed: Quantified quality metrics and automated monitoring are in place - Level 5 - Optimizing: Continuous improvement, data governance integrated into daily development workflows

------

3. Data Lineage: Where It Comes From, Where It Goes3. Data Lineage: Where It Comes From, Where It Goes

Data Lineage records the complete flow path of data from its source to final consumption. It's like data's "family tree," allowing you to trace the origins and destinations of any piece of data.Data Lineage records the complete flow path of data from its source to final consumption. It's like data's "family tree," allowing you to trace the origins and destinations of any piece of data.

Data lineage has three core application scenarios in practice:Data lineage has three core application scenarios in practice:

ScenarioQuestionHow Lineage Helps
Impact AnalysisIf I modify a field in the user table, which downstream reports will be affected?Trace all dependencies downstream along the lineage
Root Cause InvestigationToday's GMV report data is abnormal โ€” where did the problem occur?Trace upstream along the lineage at each step
Compliance AuditWhich systems has the user's phone number passed through? Is it masked everywhere?Track the full chain flow of sensitive fields
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- Active Collection: Parse SQL statements and ETL configurations to automatically extract table-level and field-level lineage relationships - Passive Collection: Intercept execution plans of query engines (such as Hive, Spark) through hooks to record lineage in real time Mainstream tools like Apache Atlas, DataHub, and OpenLineage all support automated lineage collection.- Active Collection: Parse SQL statements and ETL configurations to automatically extract table-level and field-level lineage relationships - Passive Collection: Intercept execution plans of query engines (such as Hive, Spark) through hooks to record lineage in real time Mainstream tools like Apache Atlas, DataHub, and OpenLineage all support automated lineage collection.

------

4. Metadata Management: "Data About Data"4. Metadata Management: "Data About Data"

Metadata is data about data. If data is the content of a book, metadata is the book's table of contents, author, publication date, and ISBN number. Without metadata, data is just a pile of incomprehensible numbers and strings.Metadata is data about data. If data is the content of a book, metadata is the book's table of contents, author, publication date, and ISBN number. Without metadata, data is just a pile of incomprehensible numbers and strings.

Metadata TypeDescriptionExamples
Technical MetadataPhysical storage information about dataTable name, field type, partitioning method, storage location
Business MetadataBusiness meaning of dataField display name, business definition, calculation methodology
Operational MetadataRuntime status of dataETL execution time, data volume, update frequency
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

A data dictionary is the most fundamental output of metadata management. A good data dictionary should include: - Field Name: English name and display name - Data Type: VARCHAR(50), INT, DATETIME, etc. - Business Definition: What does this field represent? How is it calculated? - Value Range: What are valid values? Are nulls allowed? - Owner: Who maintains this field? Who to contact for issues? Without a data dictionary, a new team member might need a week to understand a table's meaning; with one, 10 minutes is enough.A data dictionary is the most fundamental output of metadata management. A good data dictionary should include: - Field Name: English name and display name - Data Type: VARCHAR(50), INT, DATETIME, etc. - Business Definition: What does this field represent? How is it calculated? - Value Range: What are valid values? Are nulls allowed? - Owner: Who maintains this field? Who to contact for issues? Without a data dictionary, a new team member might need a week to understand a table's meaning; with one, 10 minutes is enough.

------

5. Data Layered Architecture: ODS โ†’ DWD โ†’ DWS โ†’ ADS5. Data Layered Architecture: ODS โ†’ DWD โ†’ DWS โ†’ ADS

A data warehouse doesn't just pile all data together โ€” it stores data in layers based on processing degree. Each layer has a clear responsibility, with upper layers depending on lower layers, gradually refining raw data into business-ready data.A data warehouse doesn't just pile all data together โ€” it stores data in layers based on processing degree. Each layer has a clear responsibility, with upper layers depending on lower layers, gradually refining raw data into business-ready data.

LayerFull NameResponsibilityData Characteristics
ODSOperational Data StoreSync business database as-isMost raw, unprocessed
DWDData Warehouse DetailClean, standardize, deduplicateClean detail records
DWSData Warehouse SummaryAggregate by subject (day/week/month)Pre-computed aggregate metrics
ADSApplication Data StoreOriented toward specific reports/APIsDirectly usable result data
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

- Reuse: DWD layer is cleaned once and shared by all upper layers, avoiding duplicate cleaning - Decoupling: Business database schema changes only affect the ODS layer, not reports - Performance: DWS layer pre-aggregates; reports read directly without real-time computation - Traceability: Each layer is preserved; issues can be investigated layer by layer- Reuse: DWD layer is cleaned once and shared by all upper layers, avoiding duplicate cleaning - Decoupling: Business database schema changes only affect the ODS layer, not reports - Performance: DWS layer pre-aggregates; reports read directly without real-time computation - Traceability: Each layer is preserved; issues can be investigated layer by layer

------

6. Governance Tools and Practices6. Governance Tools and Practices

ToolPositioningCore CapabilitiesUse Cases
Great ExpectationsData QualityDeclarative data validation rules, auto-generated quality reportsPython data pipelines
dbtData TransformationSQL-modeled development, built-in testing and documentation generationData warehouse modeling
DataHubMetadata ManagementData catalog, lineage tracing, data discoveryEnterprise data governance
Apache AtlasMetadata ManagementHadoop ecosystem lineage tracingBig data platforms
OpenMetadataMetadata ManagementOpen-source data catalog, supports multiple data sourcesSmall and medium teams
AmundsenData DiscoverySearch-based data discovery platformData democratization
๐Ÿ’ก Tips Praktis๐Ÿ’ก Pro Tip

If your team doesn't yet have data governance, we recommend following this sequence: 1. Build a data dictionary first: Document the meaning of existing tables and fields (even using Excel) 2. Add quality checks: Insert basic null and range validations in critical data pipelines 3. Unify metric definitions: Standardize the calculation methodology for core metrics like "DAU," "MAU," and "GMV" 4. Introduce tools: When manual management becomes too costly, adopt tools like DataHub or dbt 5. Establish processes: Data changes require review; quality issues have SLAs and alertsIf your team doesn't yet have data governance, we recommend following this sequence: 1. Build a data dictionary first: Document the meaning of existing tables and fields (even using Excel) 2. Add quality checks: Insert basic null and range validations in critical data pipelines 3. Unify metric definitions: Standardize the calculation methodology for core metrics like "DAU," "MAU," and "GMV" 4. Introduce tools: When manual management becomes too costly, adopt tools like DataHub or dbt 5. Establish processes: Data changes require review; quality issues have SLAs and alerts

------

SummarySummary

Data governance is the systematic engineering that transforms data from "usable" to "easy to use, trustworthy, and traceable." It is not a one-time project but an ongoing operational process.Data governance is the systematic engineering that transforms data from "usable" to "easy to use, trustworthy, and traceable." It is not a one-time project but an ongoing operational process.

Key takeaways from this chapter:Key takeaways from this chapter:

  1. Six Quality Dimensions: Completeness, accuracy, consistency, timeliness, uniqueness, validitySix Quality Dimensions: Completeness, accuracy, consistency, timeliness, uniqueness, validity
  2. Four Governance Pillars: Organization, processes, technology, culture โ€” all are essentialFour Governance Pillars: Organization, processes, technology, culture โ€” all are essential
  3. Data Lineage: Track data origins and destinations, supporting impact analysis and root cause investigationData Lineage: Track data origins and destinations, supporting impact analysis and root cause investigation
  4. Metadata Management: The data dictionary is the most fundamental and important governance outputMetadata Management: The data dictionary is the most fundamental and important governance output
  5. Layered Architecture: ODS โ†’ DWD โ†’ DWS โ†’ ADS, progressively refining data valueLayered Architecture: ODS โ†’ DWD โ†’ DWS โ†’ ADS, progressively refining data value
  6. Incremental Implementation: Start with a data dictionary and gradually introduce tools and processesIncremental Implementation: Start with a data dictionary and gradually introduce tools and processes
  7. Further ReadingFurther Reading

    • [DAMA-DMBOK](https://www.dama.org/cpages/body-of-knowledge) - Data Management Body of Knowledge, the "bible" of data governance[DAMA-DMBOK](https://www.dama.org/cpages/body-of-knowledge) - Data Management Body of Knowledge, the "bible" of data governance
    • [DataHub](https://datahubproject.io/) - LinkedIn's open-source metadata management platform[DataHub](https://datahubproject.io/) - LinkedIn's open-source metadata management platform
    • [Great Expectations](https://greatexpectations.io/) - Python data quality framework[Great Expectations](https://greatexpectations.io/) - Python data quality framework
    • [dbt](https://www.getdbt.com/) - Data transformation tool with built-in testing and documentation[dbt](https://www.getdbt.com/) - Data transformation tool with built-in testing and documentation
    • [Apache Atlas](https://atlas.apache.org/) - Metadata governance framework for the Hadoop ecosystem[Apache Atlas](https://atlas.apache.org/) - Metadata governance framework for the Hadoop ecosystem
    • [The Data Warehouse Toolkit](https://www.kimballgroup.com/data-warehouse-business-intelligence-resources/books/) - Kimball's classic on data warehouse modeling[The Data Warehouse Toolkit](https://www.kimballgroup.com/data-warehouse-business-intelligence-resources/books/) - Kimball's classic on data warehouse modeling