SkillsLib.ai

Data Quality Test Framework Builder

Generate Data Quality Tests & Anomaly Detection for dbt

0.0(0 reviews)
10+ downloads
Updated Sep 2026

What You Can Do

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

Features

dbt test suite generation

Automatically create YAML test configurations for freshness, completeness, uniqueness, and relationships

Anomaly detection protocols

Build statistical methods to flag outliers, value distributions, and unexpected data patterns

SQL validation patterns

Generate parameterized SQL queries that validate data ranges, constraints, and business logic

Data freshness checks

Create schedules and thresholds to monitor when tables were last updated and alert on staleness

Test documentation

Auto-generate test descriptions, ownership, SLAs, and remediation playbooks for your data quality standards

Coverage analysis

Identify which tables and columns lack tests and prioritize high-impact validation rules

Multi-warehouse support

Generate tests for dbt Cloud, Snowflake, BigQuery, PostgreSQL, and other SQL platforms

Test failure alerting

Build notification configurations that route data quality incidents to the right team members

Example Output

Example 1: dbt Test Suite Configuration

code
version: 2

models:
  - name: customers
    columns:
      - name: customer_id
        tests:
          - unique
          - not_null
      - name: email
        tests:
          - unique
          - matches_pattern:
              pattern: '^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}$'
    tests:
      - dbt_utils.recency:
          datepart: day
          interval: 1
          field_name: updated_at

Example 2: Anomaly Detection Query

code
WITH stats AS (
  SELECT
    AVG(order_value) as mean_value,
    STDDEV_POP(order_value) as std_dev
  FROM orders
  WHERE DATE(created_at) >= CURRENT_DATE - INTERVAL 30 DAY
)
SELECT
  order_id,
  order_value,
  ROUND((order_value - stats.mean_value) / stats.std_dev, 2) as z_score
FROM orders, stats
WHERE DATE(created_at) = CURRENT_DATE
  AND ABS((order_value - stats.mean_value) / stats.std_dev) > 3

Example 3: Test Report

code
✓ Completeness: 99.8% of required fields populated
✓ Freshness: Last update 2 hours ago (within 4-hour SLA)
✗ Anomalies: 12 outlier orders detected (values > 3σ)
⚠ Coverage: 18/22 columns tested (82%)

What's Included

  • SKILL.md: Full framework for building data quality tests
  • dbt test templates: Pre-built test patterns for common validation scenarios
  • SQL validation library: Reusable queries for constraint checks, distribution analysis, and freshness monitoring
  • Anomaly detection playbook: Statistical methods (z-score, IQR, seasonal decomposition) with implementation examples
  • Test documentation generator: Template for auto-generating test ownership, SLAs, and remediation runbooks
  • Multi-warehouse guide: Syntax and best practices for Snowflake, BigQuery, PostgreSQL, Redshift, and Databricks

Who It's For

  • Data engineers — Build robust data quality frameworks for ETL/ELT pipelines and dbt projects
  • Analytics engineers — Validate transformations and ensure data reliability for downstream analytics
  • QA engineers — Create automated data validation suites that test data pipelines like you'd test code
  • Data analysts — Document and enforce data quality standards to reduce data quality incidents
  • Analytics platform teams — Implement organization-wide data quality governance and monitoring

Best For

  • Setting up production dbt test suites from scratch
  • Building anomaly detection for high-volume transaction tables
  • Automating data validation across multiple warehouses
  • Creating SLA-driven data quality dashboards and alerts
  • Documenting and scaling data quality standards across teams

You might also like

Experiment Design & Statistical Analysis for Research Engineers
$45
Experiment Design & Statistical Analysis for Research Engineers

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

HEOR Evidence Synthesis & Dossier Builder
$35
HEOR3.6(5)
HEOR Evidence Synthesis & Dossier Builder

You can rapidly compile, organize, and format disparate health economic evidence—from clinical trials to cost-effectiveness analyses—into structured, regulatory-compliant dossiers. This skill maps your evidence to specific payer and HTA requirements, automatically generates evidence hierarchies, and produces submission-ready dossier outlines with formatting that meets regulatory standards for NICE, EUnetHTA, and other major bodies.

Payer Evidence Synthesis & HTA Builder
$30
Payer4.0(3)
Payer Evidence Synthesis & HTA Builder

You can structure comprehensive health technology assessments (HTAs) that organize clinical evidence, economic analyses, and regulatory considerations into evidence-based coverage recommendations. This skill helps you synthesize clinical trial data, health economic models, and real-world evidence into clear, defensible payer coverage determinations. You'll generate professional HTA reports that align with major frameworks like ICER, CADTH, and NICE standards.

Setup Agent Tail
$45
Monitoring4.4(48)
Setup Agent Tail

This skill detects your project framework (Vite, Next.js, plain Node, or monorepo) and automatically configures agent-tail to pipe dev server and browser console logs into unified log files. You'll get a proposed configuration tailored to your setup, install agent-tail with the correct plugins, and have logs immediately available for AI agents to consume and analyze.

Fine-Tuning Dataset Preparation & Validation for Claude
$45
Fine-Tuning Dataset Preparation & Validation for Claude

This skill walks you through the complete dataset preparation workflow for fine-tuning Claude models. You'll validate training data quality, detect and fix formatting errors, identify data imbalances, and generate validation reports to ensure your fine-tuned model performs reliably in production.

Statistical Analysis Guide
$35
Statistical Analysis Guide

You can confidently navigate complex statistical analyses by working through structured workflows that match your research question and data type to the right test. The skill guides you through assumption validation, helps you interpret results accurately without common pitfalls, and walks you through regression diagnostics with full transparency about what your findings actually mean.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$35.00