SkillsLib.ai

RAG Pipeline Optimization with Claude

Design and optimize RAG systems for production-grade accuracy and performance

0.0(0 reviews)
100+ downloads
Updated Sep 2026

What You Can Do

You can systematically evaluate and improve your RAG pipelines using Claude as a design partner. You'll analyze retrieval quality, identify bottlenecks in your embedding and chunking strategies, and receive actionable recommendations to reduce hallucinations and improve context relevance. By the end, you'll have a data-driven optimization plan tailored to your specific use case and performance metrics.

Features

Retrieval quality evaluation

Analyze precision, recall, and relevance scores across your document chunks

Query optimization strategies

Rewrite and decompose queries for better vector search results

Chunk size and strategy analysis

Find the optimal granularity for your domain and retrieval patterns

Embedding model recommendations

Compare embedding strategies and select models for your performance needs

Hallucination mitigation

Identify gaps between retrieved context and model outputs, with concrete fixes

Context window optimization

Design effective prompts that balance context length with accuracy

Vector similarity metrics analysis

Evaluate semantic distance calculations and similarity thresholds

Prompt engineering for RAG

Craft system prompts and instructions that minimize out-of-context outputs

Example Output

RAG Evaluation Report

  • ✅ Retrieval Quality Score: 78/100

  • Precision (correct docs returned): 82%

  • Recall (relevant docs captured): 74%

  • Mean reciprocal rank: 0.71

  • ❌ Key Issues Identified

  • Chunk size (512 tokens) too small for multi-step reasoning queries

  • Embedding model (text-embedding-3-small) loses semantic nuance for domain terminology

  • Query rewriting not active; queries like "What does the warranty cover?" map to irrelevant sections

  • 💡 Optimization Plan

  1. Increase chunk size to 1024 tokens with 200-token overlap
  2. Migrate to text-embedding-3-large with domain-specific fine-tuning
  3. Add query decomposition: break multi-part questions into sub-queries
  4. Set similarity threshold to 0.75 (currently 0.5) to reduce noise

Expected Impact: Precision +12%, Recall +8%, hallucination rate -35%

What's Included

  • SKILL.md: Full RAG optimization workflow with decision trees and evaluation frameworks
  • RAG Quality Checklist: Evaluation criteria for retrieval, ranking, and context relevance
  • Query Optimization Template: Strategies for decomposition, rewriting, and intent detection
  • Embedding Strategy Decision Tree: Select models, chunk sizes, and vector similarity thresholds
  • Performance Metrics Tracker: Template to log precision, recall, MRR, and latency across iterations
  • Hallucination Root-Cause Analyzer: Worksheet to identify when and why the model deviates from context

Who It's For

  • ML Engineers building and tuning RAG systems for production
  • Data Engineers designing document ingestion and vector storage pipelines
  • LLM Application Developers integrating retrieval into chat and search products
  • AI Product Managers optimizing RAG performance for user-facing features
  • Prompt Engineers evaluating context quality and reducing hallucinations

Best For

  • Evaluating retrieval quality and identifying bottlenecks
  • Optimizing chunk size, overlap, and embedding strategies
  • Designing multi-turn RAG workflows with query rewriting
  • Reducing hallucinations by improving context relevance
  • Benchmarking RAG performance metrics (precision, recall, MRR)

You might also like

Agentdb Vector Search
$20
RAG4.1(34)
Agentdb Vector Search

You can build production-grade vector search systems that retrieve semantically similar documents in sub-millisecond time using AgentDB's optimized HNSW indexing. The skill enables you to implement RAG pipelines, semantic search engines, and intelligent knowledge bases with configurable embedding dimensions, distance metrics (cosine, Euclidean, dot product), and similarity thresholds—all with built-in quantization and caching for massive performance gains.

BI Data Quality Investigator
$30
BI Data Quality Investigator

You'll systematically diagnose data quality problems by developing structured root cause analysis frameworks, calculating the true business impact, and creating reproducible validation tests. This skill walks you through hypothesis-driven investigation, data lineage analysis, and remediation planning — turning data issues into documented fixes and preventive measures.

Data Quality Test Framework Builder
$35
Data Quality Test Framework Builder

You'll build comprehensive data quality test suites that validate transformations, detect anomalies, and document standards across your dbt and SQL pipelines. This skill generates production-ready test configurations, anomaly detection protocols, and validation rules that catch data issues before they impact analytics.

Model Evaluation Suite
$30
Model Evaluation Suite

You can design multi-dimensional evaluation strategies tailored to your model's specific capabilities and use cases, create representative test sets that expose edge cases and failure modes, implement automated scoring mechanisms for reproducible results, and generate benchmark comparison reports that contextualize performance within industry standards. This skill transforms ad-hoc testing into systematic, evidence-based model assessment—essential for production deployment decisions and ongoing performance monitoring.

Analytics Report Builder: Executive-Ready Data Storytelling
$40
Analytics Report Builder: Executive-Ready Data Storytelling

You can convert complex datasets and business metrics into polished, executive-ready reports that stakeholders trust and act on. Claude generates data-driven narratives, executive summaries, actionable insights, and visualization recommendations tailored to your audience's priorities. Your reports will tell a cohesive story that connects metrics to business outcomes, eliminating confusion and accelerating decision-making.

Payer Evidence Synthesis & HTA Builder
$30
Payer4.0(3)
Payer Evidence Synthesis & HTA Builder

You can structure comprehensive health technology assessments (HTAs) that organize clinical evidence, economic analyses, and regulatory considerations into evidence-based coverage recommendations. This skill helps you synthesize clinical trial data, health economic models, and real-world evidence into clear, defensible payer coverage determinations. You'll generate professional HTA reports that align with major frameworks like ICER, CADTH, and NICE standards.

Setup Agent Tail
$45
Monitoring4.4(48)
Setup Agent Tail

This skill detects your project framework (Vite, Next.js, plain Node, or monorepo) and automatically configures agent-tail to pipe dev server and browser console logs into unified log files. You'll get a proposed configuration tailored to your setup, install agent-tail with the correct plugins, and have logs immediately available for AI agents to consume and analyze.

Analytics Documentation Generator
$25
Analytics Documentation Generator

This skill automatically documents your entire analytics infrastructure by analyzing data sources, transformations, and outputs. You'll generate production-ready data dictionaries with field definitions, lineage maps showing data flow across systems, and transformation documentation that explains logic and dependencies. Save weeks of manual documentation work while keeping your analytics stack discoverable as it evolves.

$40.00