SkillsLib.ai

Data Pipeline Architect

Design scalable ETL/ELT data pipelines with orchestration and quality checks

4.5(51 reviews)
500+ downloads
Updated Oct 2026
Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.

What You Can Do

You can design complete data pipeline architectures that reliably move data from source systems to analytical or operational destinations. This skill helps you choose between ETL vs. ELT patterns, select appropriate orchestration tools (Airflow, Prefect, dbt), map complex data flows, handle schema evolution, embed data quality checks, and create maintainable pipeline specifications that scale with your organization.

Features

ETL vs. ELT pattern selection

Choose the right architecture based on your source systems, transformation complexity, and destination requirements

Tool recommendations

Evaluate and select from Airflow, Prefect, dbt, or custom solutions based on team skill, infrastructure, and pipeline complexity

Data flow mapping

Document source systems, transformation stages, dependencies, and target schemas with clear lineage

Schema evolution strategy

Manage breaking changes, backward compatibility, and versioning for evolving data models

Data quality framework

Embed validation rules, monitoring, alerting, and anomaly detection throughout pipeline stages

Error handling & retry logic

Design resilient pipelines with proper handling of failures, timeouts, and idempotency

Documentation templates

Create auditable, maintainable pipeline specifications for handoff and team collaboration

Example Output

Pipeline Architecture Example

For a multi-source analytics pipeline:

  • Architecture Decision: ELT pattern chosen (transform in warehouse) due to large data volume and existing Snowflake infrastructure
  • Tool Stack: Prefect for orchestration (flexible task dependencies), dbt for transformations, Snowflake for compute
  • Data Flow: Raw data lands in staging → dbt models apply business logic → marts feed BI tools
  • Quality Checks: Row count validation at each stage, null checks on key columns, referential integrity tests
  • Monitoring: Failed task alerts to Slack, daily lineage reports, 4-hour SLA enforcement

Schema Evolution Plan

  • Versioning Strategy: New columns added as v2 suffixes during transition, deprecated columns tracked in metadata
  • Backward Compatibility: Maintain mapping layer for 2 releases before removing old schemas
  • Breaking Change Protocol: 30-day notice to downstream teams, automated migration scripts provided

What's Included

  • SKILL.md: Complete data pipeline architecture framework with decision trees
  • Pipeline Design Template: ETL/ELT architecture canvas with source-transform-destination mapping
  • Tool Selection Checklist: Evaluation matrix for Airflow, Prefect, dbt, and custom solutions
  • Data Quality Specification: Validation rule templates and anomaly detection patterns
  • Schema Evolution Playbook: Breaking change management and versioning strategies
  • Architecture Documentation Template: Standardized specification for pipeline handoff and governance

Who It's For

  • Data Engineers — Design and implement production-grade pipelines with proper orchestration and monitoring
  • Data Architects — Plan enterprise-scale data infrastructure and integration strategies
  • Analytics Engineers — Architect ELT workflows that combine data loading with transformation in the warehouse
  • Data Leaders/Managers — Evaluate tool options and design team processes around data pipeline development
  • Platform Engineers — Build internal data infrastructure and self-service pipeline platforms

Best For

  • Multi-source data integration projects combining APIs, databases, and file systems
  • ETL/ELT pipeline design and architecture decisions
  • Tool selection and evaluation for data orchestration
  • Schema evolution and breaking change management strategies
  • Data quality framework and validation rule design
  • Pipeline documentation and governance specification

You might also like

Claude Fine-Tuning Optimization
$30
Claude Fine-Tuning Optimization

This skill provides a systematic framework for preparing training data, benchmarking model performance, identifying deployment risks, and calculating true ROI before committing fine-tuning investments. You'll validate datasets, run structured evaluations against baseline Claude models, test edge cases, and iterate toward production-ready fine-tuned versions. The framework ensures your fine-tuned models deliver meaningful improvements while maintaining safety and cost-effectiveness.

Experiment Design & Statistical Analysis for Research Engineers
$45
Experiment Design & Statistical Analysis for Research Engineers

You can design statistically valid experiments with proper power analysis, choose the right statistical tests for your data type, analyze results while controlling for multiple comparisons, and generate publication-ready reports with accurate interpretation of findings. Claude helps you avoid common statistical pitfalls and ensures your experimental claims are well-supported by evidence.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

Analytics Documentation Generator
$25
Analytics Documentation Generator

You can create comprehensive documentation for your data platforms, metrics, and analytics infrastructure that stakeholders actually understand. The skill generates clear data dictionaries, metric definitions, pipeline diagrams, and runbooks that bridge the gap between technical teams and business users. Your documentation stays consistent with your actual infrastructure while being instantly accessible to analysts, managers, and engineers.

HEOR Evidence Synthesis & Dossier Builder
$35
HEOR3.6(5)
HEOR Evidence Synthesis & Dossier Builder

You can rapidly compile, organize, and format disparate health economic evidence—from clinical trials to cost-effectiveness analyses—into structured, regulatory-compliant dossiers. This skill maps your evidence to specific payer and HTA requirements, automatically generates evidence hierarchies, and produces submission-ready dossier outlines with formatting that meets regulatory standards for NICE, EUnetHTA, and other major bodies.

Payer Evidence Synthesis & HTA Builder
$30
Payer4.0(3)
Payer Evidence Synthesis & HTA Builder

You can structure comprehensive health technology assessments (HTAs) that organize clinical evidence, economic analyses, and regulatory considerations into evidence-based coverage recommendations. This skill helps you synthesize clinical trial data, health economic models, and real-world evidence into clear, defensible payer coverage determinations. You'll generate professional HTA reports that align with major frameworks like ICER, CADTH, and NICE standards.

ML Infrastructure Failure Analysis & Optimization
$25
ML Infrastructure Failure Analysis & Optimization

This skill helps you systematically diagnose failures in distributed ML training and serving infrastructure. You provide system logs, metrics, and error traces, and Claude performs structured root-cause analysis to identify the underlying issue—whether it's resource exhaustion, distributed system deadlock, data pipeline corruption, or model serving misconfiguration. You get a detailed diagnosis with remediation steps ranked by likelihood and implementation effort.

$40.00