RAG in Production: Embeddings That Don't Lie
RAG is the easiest architecture to demo and the hardest to keep honest. Chunking strategies, hybrid search (BM25 + embeddings), and an eval loop that measures groundedness before users ever do.
Chunking that respects structure
Naive fixed-size chunks break headings, tables and context. We chunk on semantic boundaries and keep a parent-child retriever so answers can cite whole sections.
Hybrid search wins
Pure vector search misses exact keywords. A reciprocal-rank fusion of BM25 and vector scores consistently beats either alone in our evals.
Evals in CI
Every prompt change runs a 300-case suite measuring answer relevance, groundedness and latency. If it regresses, the deploy fails.