
Data Pipeline Architect
Design scalable ETL/ELT data pipelines with orchestration and quality checks
What You Can Do
You can design complete data pipeline architectures that reliably move data from source systems to analytical or operational destinations. This skill helps you choose between ETL vs. ELT patterns, select appropriate orchestration tools (Airflow, Prefect, dbt), map complex data flows, handle schema evolution, embed data quality checks, and create maintainable pipeline specifications that scale with your organization.
Features
Choose the right architecture based on your source systems, transformation complexity, and destination requirements
Evaluate and select from Airflow, Prefect, dbt, or custom solutions based on team skill, infrastructure, and pipeline complexity
Document source systems, transformation stages, dependencies, and target schemas with clear lineage
Manage breaking changes, backward compatibility, and versioning for evolving data models
Embed validation rules, monitoring, alerting, and anomaly detection throughout pipeline stages
Design resilient pipelines with proper handling of failures, timeouts, and idempotency
Create auditable, maintainable pipeline specifications for handoff and team collaboration
Example Output
Pipeline Architecture Example
For a multi-source analytics pipeline:
- Architecture Decision: ELT pattern chosen (transform in warehouse) due to large data volume and existing Snowflake infrastructure
- Tool Stack: Prefect for orchestration (flexible task dependencies), dbt for transformations, Snowflake for compute
- Data Flow: Raw data lands in staging → dbt models apply business logic → marts feed BI tools
- Quality Checks: Row count validation at each stage, null checks on key columns, referential integrity tests
- Monitoring: Failed task alerts to Slack, daily lineage reports, 4-hour SLA enforcement
Schema Evolution Plan
- Versioning Strategy: New columns added as
v2suffixes during transition, deprecated columns tracked in metadata - Backward Compatibility: Maintain mapping layer for 2 releases before removing old schemas
- Breaking Change Protocol: 30-day notice to downstream teams, automated migration scripts provided
What's Included
- SKILL.md: Complete data pipeline architecture framework with decision trees
- Pipeline Design Template: ETL/ELT architecture canvas with source-transform-destination mapping
- Tool Selection Checklist: Evaluation matrix for Airflow, Prefect, dbt, and custom solutions
- Data Quality Specification: Validation rule templates and anomaly detection patterns
- Schema Evolution Playbook: Breaking change management and versioning strategies
- Architecture Documentation Template: Standardized specification for pipeline handoff and governance
Who It's For
- Data Engineers — Design and implement production-grade pipelines with proper orchestration and monitoring
- Data Architects — Plan enterprise-scale data infrastructure and integration strategies
- Analytics Engineers — Architect ELT workflows that combine data loading with transformation in the warehouse
- Data Leaders/Managers — Evaluate tool options and design team processes around data pipeline development
- Platform Engineers — Build internal data infrastructure and self-service pipeline platforms
Best For
- Multi-source data integration projects combining APIs, databases, and file systems
- ETL/ELT pipeline design and architecture decisions
- Tool selection and evaluation for data orchestration
- Schema evolution and breaking change management strategies
- Data quality framework and validation rule design
- Pipeline documentation and governance specification







