
Data Quality Checker
Generate dbt tests and Great Expectations suites to validate data pipelines
What You Can Do
You can analyze your data structure and content to identify quality risks, then generate production-ready dbt tests and Great Expectations validation suites. Claude creates schema conformance checks, uniqueness constraints, referential integrity validations, completeness gap detection, and data freshness monitoring—all in code-first formats (YAML and Python) that integrate seamlessly with modern data stacks.
Features
detects data type mismatches, unexpected columns, and structural anomalies
enforces foreign key relationships and cross-table dependencies
identifies missing values and incomplete records across datasets
flags stale data and latency issues in your pipelines
creates YAML-based tests for uniqueness, not-null, and custom SQL assertions
builds statistical and rule-based validation profiles in Python
enforces business logic constraints on numeric, categorical, and temporal data
sets up statistical thresholds for outliers and distribution changes
Example Output
Example 1: dbt Tests for Orders Table
version: 2
models:
- name: orders
columns:
- name: order_id
tests:
- unique
- not_null
- name: customer_id
tests:
- relationships:
to: ref('customers')
field: customer_id
- name: order_date
tests:
- dbt_utils.accepted_values:
values: ['2024-01-01', '2024-01-02']
Example 2: Great Expectations Checkpoint
validation_results = context.run_checkpoint(
checkpoint_name='orders_validation',
validations=[
{'expectation_suite_name': 'orders.full'},
{'batch_request': {'datasource_name': 'postgres', 'data_connector_name': 'default'}}
]
)
Example 3: Data Freshness Alert Detects orders table not updated in last 24 hours and flags upstream dependency failures in your pipeline orchestrator.
What's Included
- SKILL.md instruction file with comprehensive framework for data quality rule generation:
- dbt test templates for schema, uniqueness, referential integrity, and custom SQL validations:
- Great Expectations suite builder: Python configuration templates for statistical profiling
- Data freshness checklist: timestamp validation patterns and staleness detection queries
- Quality rules documentation: YAML and Python code-first reference formats for version control
Who It's For
- Data engineers building or maintaining dbt projects and analytics engineering workflows
- Analytics engineers establishing data governance and quality frameworks
- Data quality specialists implementing validation suites across enterprise pipelines
- Data platform teams inheriting legacy pipelines and establishing baseline quality controls
- DevOps and infrastructure engineers automating data validation in CI/CD workflows
Best For
- Creating dbt test suites for new fact and dimension tables in dimensional warehouses
- Generating Great Expectations validation profiles from existing data structures
- Implementing referential integrity and schema conformance checks in inherited pipelines
- Building data freshness monitors and staleness detection rules for real-time dashboards
- Establishing completeness and value range constraints for regulatory compliance workflows







