
Csv & Spreadsheet Analyzer
Analyze CSV, Excel & JSON data: statistics, quality issues & Python code
What You Can Do
Systematically analyze any CSV, Excel, or JSON dataset to understand its structure, quality, and patterns. You'll receive detailed summary statistics (mean, median, standard deviation, quartiles), identify missing values, duplicates, and type mismatches, detect statistical outliers and anomalies, and get working Python/pandas code you can run or integrate directly into your data pipelines.
Features
Calculate mean, median, standard deviation, quartiles, and value counts for all numeric and categorical columns
Automatically detect missing values, duplicates, type inconsistencies, and cardinality issues across your dataset
Identify statistical anomalies using IQR, Z-score, and distribution analysis methods
Generate ready-to-run pandas code that replicates the analysis for integration into your workflows
Map null patterns and provide insights into data completeness by column and row
Identify relationships between numeric variables to spot patterns and dependencies
Analyze value distributions, skewness, and kurtosis to understand data shape and spread
Example Output
Input: CSV with 1,000 sales records
Output:
- Dataset has 1,000 rows × 8 columns; 2.3% missing values in 'region'
- Top 3 anomalies: 3 orders with negative quantities, 15 prices >3 std devs above mean
- Correlation matrix shows strong relationship (0.87) between order_value and quantity
- Python code block with pandas script to reproduce all findings
Input: Excel spreadsheet with employee data
Output:
- 145 employees; 12 duplicates by email found; salary range $35k–$250k (outlier: one record at $1.2M)
- Missing department values in 8 records; job_title has 42 unique values
- Recommended data cleaning steps with code snippet to handle duplicates and type conversions
What's Included
- SKILL.md: Complete system prompt with analysis methods and output structure
- Data Quality Checklist: Step-by-step checklist for validating datasets before analysis
- Python Code Templates: Reusable pandas snippets for statistics, outlier detection, and profiling
- Analysis Report Framework: Structured markdown template for documenting findings and recommendations
- Sample Datasets: Example CSV files to test the skill and practice exploratory data analysis
Who It's For
- Data Analysts — Quickly onboard and validate new datasets before modeling
- Data Engineers — Profile raw data and identify quality issues before pipeline processing
- Business Analysts — Understand data structure and anomalies for reporting and insights
- Data Scientists — Perform exploratory data analysis and detect outliers before feature engineering
- Database Administrators — Validate data integrity and completeness across source systems
Best For
- Exploratory Data Analysis (EDA) — Get a complete statistical overview of unfamiliar datasets
- Data Quality Assessment — Detect missing values, duplicates, type mismatches, and anomalies
- Outlier Detection — Identify statistical anomalies and unusual records using multiple methods
- Data Validation Workflows — Validate datasets before downstream processing or analysis
- Python Code Generation — Receive reproducible pandas code ready to integrate into pipelines







