
Test Flakiness
Detect flaky tests by analyzing CI logs and pass rate history
4.5(48 reviews)500+ downloadsUpdated Oct 2026Verified SafeSecurity VerifiedThis skill was analyzed by our AI security scanner for harmful content including data exfiltration, system manipulation, credential theft, and prompt injection. No threats were detected.
What You Can Do
Analyze CI run logs and test result history to detect tests that fail non-deterministically, calculate per-test pass rates across multiple runs, and distinguish between genuinely broken tests and intermittent failures. You get actionable recommendations on whether to quarantine or fix each flaky test, plus an updated registry to prevent masking real failures.
Features
Parse CI logs from GitHub Actions, GitLab CI, or standard test result formats (JUnit XML, NUnit)
automatically detects your CI platform
Aggregate pass rates across multiple CI runs to identify intermittent failures with statistical confidence
Distinguish flaky tests from consistently failing tests using pass-rate thresholds and failure pattern analysis
Generate quarantine recommendations with severity levels and likely root causes (timing issues, environment dependencies, race conditions)
Maintain a flaky test registry in regression-suite.md with history, status, and remediation guidance for each flagged test
Scan all available CI logs in standard directories (.github/, test-results/) or analyze a specific log file on demand
Provide remediation pathways: isolate vs. fix recommendations based on test complexity and failure frequency
Example Output
Input: CI logs from 15 recent runs
Output:
code
## Flaky Tests Detected (Pass Rate < 95%)
### High Priority (40-80% pass rate)
- `test_player_spawn_collision` — 65% pass rate — Likely: race condition in physics engine initialization
- **Recommendation:** Quarantine immediately, add synchronization to test setup
- **Severity:** Critical (masks real failures)
### Medium Priority (80-95% pass rate)
- `test_ui_animation_complete` — 92% pass rate — Likely: timing-dependent, platform-specific
- **Recommendation:** Increase wait timeout or add retry logic
- **Severity:** Moderate
What's Included
- `test-flakiness` SKILL.md instruction file with multi-platform CI log parsing:
- Flakiness detection algorithm with statistical thresholds:
- Quarantine registry template for regression-suite.md:
- Pass-rate calculation framework and failure pattern analyzer:
- Remediation recommendation matrix (quarantine vs. fix guidance):
Who It's For
- QA/Test Engineers — diagnose why CI is unreliable and decide quarantine vs. fix strategies
- Release Engineers — reduce false positives during release validation runs
- Software Engineers — identify which flaky tests to prioritize fixing during Polish phase
- Platform Teams — maintain test reliability and prevent test rot across CI/CD pipelines
- DevOps Engineers — monitor test suite health and flag systemic flakiness trends
Best For
- Detecting intermittent test failures across multiple CI runs without manual log review
- Quarantining known flaky tests to prevent masking real failures in CI
- Diagnosing root causes of test non-determinism (timing, race conditions, environment sensitivity)
- Building a historical registry of flaky tests and remediation status
- Polish phase regression testing when sufficient CI data has accumulated
You might also like
$40.00







