
RAG Pipeline Diagnostics & Optimization
Optimize RAG pipelines by diagnosing retrieval quality and context-prompt alignment
What You Can Do
You can systematically diagnose why your RAG system isn't returning relevant results and identify specific misalignments between queries and context. This skill walks you through analyzing retrieval quality, scoring context relevance, and optimizing prompt-retrieval alignment to improve your RAG accuracy. You'll get concrete, actionable recommendations for fixing bottlenecks in your pipeline.
Features
Score relevance of documents returned by your vector store and identify low-confidence results
Evaluate whether retrieved context actually answers the user's query using multi-dimensional scoring
Detect semantic gaps between how queries are phrased and how documents are indexed
Pinpoint whether failures come from indexing, query encoding, similarity thresholds, or context window limits
Analyze whether your embedding model captures semantic relationships relevant to your use case
Test different similarity score cutoffs to find optimal precision-recall trade-offs
Calculate ideal chunk sizes and overlap to maximize relevant context while minimizing noise
A/B test different retrieval strategies (BM25 vs dense, reranking, query expansion) side-by-side
Example Output
RAG Diagnostic Report
System: Customer support knowledge base RAG
Query: "How do I reset my password?"
Findings:
- ✅ Retrieval quality: 7/10 (4 of 5 results relevant)
- ⚠️ Context alignment: 5/10 (results don't match exact phrasing)
- ❌ Coverage: Missing articles on password reset via SSO
Bottleneck: Query embeds "reset" but documents use "change" or "update". Synonym expansion needed.
Recommendations:
- Add query expansion: ["reset", "change", "update", "recover"] → re-index
- Test similarity threshold 0.75 (current: 0.80) — gains recall without hurting precision
- Chunk password docs with consistent terminology (audit 12 docs)
Expected improvement: 9/10 relevance within 2 weeks
What's Included
- SKILL.md: Complete RAG diagnostic workflow with decision trees and checklists
- Retrieval quality template: Markdown rubric for scoring document relevance
- Alignment assessment checklist: Step-by-step checklist for prompt-retrieval analysis
- Performance metrics dashboard: CSV template for tracking retrieval latency and recall
- Embedding diagnostics worksheet: Guide to analyze embedding quality with sample queries
- Optimization playbook: Concrete fixes for 8 common RAG failure modes (missing context, low precision, slow queries, etc.)
Who It's For
- ML/AI engineers — Build and debug RAG systems with diagnostic insights into retrieval bottlenecks
- Data scientists — Optimize embedding models and retrieval strategies to improve RAG accuracy metrics
- Search/ranking engineers — Tune similarity thresholds, chunking strategies, and reranking pipelines
- LLM product managers — Understand why RAG quality varies by use case and prioritize fixes
- AI startup founders — Audit RAG performance before scaling to customers
Best For
- Debugging low-quality retrieval results in production RAG systems
- Optimizing vector similarity thresholds and context relevance scores
- Analyzing embedding model performance for your specific domain
- A/B testing different retrieval strategies and reranking approaches
- Scaling RAG systems by identifying and fixing bottlenecks before users see errors







