Most engineering teams hit the same wall with retrieval-augmented generation. They follow the tutorial playbook: slice documents into fixed chunks of five hundred twelve or one thousand twenty-four tokens, push them through a single embedding model, and call a vector database with a simple top-k lookup. On a slide deck, this looks solid. In production, it falls apart.

Fixed chunks do not care about content. They will happily split a legal contract mid-sentence, leaving liability clauses dangling across two unrelated pieces of text. They will dump an entire API endpoint description into one bloated chunk so large that the specific parameter your user asked about drowns in noise. And when retrieval is slow, every millisecond of latency bleeds directly into the user experience. We learned this the hard way. Then we tore our retrieval layer down and rebuilt it. Our recall at ten jumped from seventy-eight percent to ninety-five percent. Latency did not increase. It collapsed.

The Problem with Copy-Paste RAG

The standard RAG stack has become a kind of default setting. Small chunks, one embedding model, vector search, done. That approach survives a demo because demos use clean questions and tidy documents. Production data is never tidy.

Legal documents have hierarchical structure. Sections contain subsections. Subsections contain clauses. Slice through them with a blunt token counter and you destroy the very relationships the model needs to reason about. API documentation has structure too, but it is different. A function signature, its parameters, its return value, and an example usage form a logical unit. Force that into a fixed token window and you either truncate the example or pad the chunk with unrelated functions. Support tickets are messy, conversational, and full of sudden topic shifts. Wikis are sprawling and cross-referenced. One chunking strategy cannot serve all of these, yet teams routinely deploy exactly that. We stopped pretending it could.

Strategic Chunking: Match the Method to the Material

We moved to content-aware chunking. For legal documents, we use recursive chunking that respects the document hierarchy. It keeps clauses intact and preserves the parent-child relationships between sections. For API documentation, we built function-aware chunking that treats each function or endpoint as a boundary. If a parameter description runs long, the chunk expands around that function, not around a token limit. For support tickets, we use semantic chunking that detects natural topic boundaries. When a customer suddenly switches from a billing complaint to a technical bug, the split happens at that turn. For wikis and unstructured knowledge bases, we use agentic chunking where a lightweight LLM evaluates the text and decides where a meaningful boundary should fall. This is slower to set up than a character split, but it is the difference between retrieval that works and retrieval that guesses.

Hybrid Retrieval: Why Vector Search Alone Is Not Enough

Vector search understands meaning, but it can miss exact matches. If a user pastes an error code like ERR_CONNECTION_RESET_0x5F3, semantic similarity might rank it below paragraphs that merely discuss network errors in general. BM25, on the other hand, finds exact strings but misses conceptual relatedness. You need both.

We run vector search and BM25 in parallel. Then we combine the results with Reciprocal Rank Fusion, or RRF, which normalizes the scores from the two different search spaces without forcing them into the same scale. After fusion, we send the top candidates through a cross-encoder reranker. This adds a small amount of latency, but the gain in precision is significant. The reranker reads the query and each candidate together and assigns a relevance score that is far more accurate than the cosine similarity of the initial embedding. In practice, this combination catches exact error codes that pure vector search misses, while still surfacing conceptually related troubleshooting steps that keyword search would ignore.

Query Expansion: Fixing User Input Before It Hits the Index

Users do not write perfect search queries. They ask multi-hop questions like why did my last deploy fail and how do I roll it back, which requires finding two separate bodies of knowledge and connecting them. Or they ask vague questions that map poorly to the index.

ਅਸੀਂ ਖੋਜ ਕਰਨ ਤੋਂ ਪਹਿਲਾਂ ਕੁਐਰੀਆਂ ਨੂੰ ਰੂਪਾਂਤਰਿਤ ਕਰਦੇ ਹਾਂ। ਇੱਕ ਮਲਟੀ-ਹੌਪ (multi-hop) ਸਵਾਲ ਨੂੰ ਉਪ-ਸਵਾਲਾਂ ਵਿੱਚ ਤੋੜ ਦਿੱਤਾ ਜਾਂਦਾ ਹੈ। ਇੱਕ ਅਸਪਸ਼ਟ ਇਰਾਦੇ ਨੂੰ ਕਈ ਖਾਸ ਸਰਚ ਕੁਐਰੀਆਂ ਵਿੱਚ ਵਧਾ ਦਿੱਤਾ ਜਾਂਦਾ ਹੈ। ਅਸੀਂ ਪਾਇਆ ਹੈ ਕਿ ਇੱਕ ਯੂਜ਼ਰ ਕੁਐਰੀ ਨੂੰ ਪੰਜ ਵੱਖ-ਵੱਖ ਸਰਚ ਕੁਐਰੀਆਂ ਵਿੱਚ ਵਧਾਉਣ ਨਾਲ ਰੀਕਾਲ (recall) ਸੱਤਰ-ਅੱਠ ਪ੍ਰਤੀਸ਼ਤ ਤੋਂ ਵਧ ਕੇ ਨੜਿਨਵੇਂ ਪ੍ਰਤੀਸ਼ਤ ਤੱਕ ਹੋ ਸਕਦਾ ਹੈ। ਇਹ LLM ਨੂੰ ਹੋਰ ਸਖ਼ਤ ਪ੍ਰੋਂਪਟ ਦੇਣ ਬਾਰੇ ਨਹੀਂ ਹੈ। ਇਹ ਰਿਟ੍ਰੀਵਲ ਸਿਸਟਮ ਨੂੰ ਸਹੀ ਸੰਦਰਭ ਲੱਭਣ ਲਈ ਵਧੇਰੇ ਮੌਕੇ ਦੇਣ ਬਾਰੇ ਹੈ। ਹਰੇਕ ਤਿਆਰ ਕੀਤੀ ਗਈ ਕੁਐਰੀ ਇੱਕ ਵੱਖਰੇ ਪਹਿਲੂ ਜਾਂ ਸ਼ਬਦਾਵਲੀ ਨੂੰ ਕੈਪਚਰ ਕਰਦੀ ਹੈ, ਅਤੇ ਮਿਲੇ ਹੋਏ ਨਤੀਜੇ ਇੱਕ ਪੂਰੀ ਤਸਵੀਰ ਪੇਸ਼ ਕਰਦੇ ਹਨ।

Bayesian Optimization: ਅੰਦਾਜ਼ੇ ਲਗਾਉਣਾ ਬੰਦ ਕਰੋ

ਇੱਕ ਵਾਰ ਜਦੋਂ ਤੁਹਾਡੇ ਕੋਲ ਕਈ ਚੰਕਿੰਗ ਰਣਨੀਤੀਆਂ (chunking strategies), ਹਾਈਬ੍ਰਿਡ ਰਿਟ੍ਰੀਵਲ, ਅਤੇ ਕੁਐਰੀ ਐਕਸਪੈਂਸ਼ਨ ਹੋ ਜਾਂਦੀਆਂ ਹਨ, ਤਾਂ ਤੁਹਾਡਾ ਸਾਹਮਣਾ ਇੱਕ ਨਵੀਂ ਸਮੱਸਿਆ ਨਾਲ ਹੁੰਦਾ ਹੈ। ਇੱਥੇ ਬਹੁਤ ਸਾਰੇ ਕੰਟਰੋਲ (knobs) ਹਨ। ਚੰਕ ਸਾਈਜ਼, ਓਵਰਲੈਪ ਪ੍ਰਤੀਸ਼ਤ, ਵੈਕਟਰ ਵੇਟ ਬਨਾਮ BM25 ਵੇਟ, ਰੀ-ਰੈਂਕਿੰਗ ਥ੍ਰੈਸ਼ਹੋਲਡ, ਅਤੇ top-k ਮੁੱਲ ਸਾਰੇ ਗੈਰ-ਰੇਖਿਕ (nonlinear) ਤਰੀਕਿਆਂ ਨਾਲ ਆਪਸ ਵਿੱਚ ਜੁੜੇ ਹੋਏ ਹਨ। ਮੈਨੂਅਲ ਟਿਊਨਿੰਗ ਇੱਕ ਅੰਦਾਜ਼ੇ ਦਾ ਖੇਡ ਬਣ ਜਾਂਦੀ ਹੈ।

ਅਸੀਂ ਅੰਦਾਜ਼ੇ ਲਗਾਉਣਾ ਬੰਦ ਕਰ ਦਿੱਤਾ। ਅਸੀਂ ... ਨੂੰ ...