મોટાભાગના RAG ડેમો લેપટોપ પર શાનદાર લાગે છે. એક સ્ક્રિપ્ટને વીસ પાનાની PDF આપો, એક પ્રશ્ન પૂછો, અને તે સાચા પેરાગ્રાફનો સંદર્ભ આપે છે તે જુઓ. પરંતુ તે જ પાઇપલાઇનને પ્રોડક્શનમાં લઈ જવું એ એવો તબક્કો છે જ્યાં વાસ્તવિકતા સામે આવે છે. કાયદાકીય દસ્તાવેજો વાક્ય સ્તરે અડધા કપાઈ જાય છે. જાડા API સંદર્ભો મહત્વના સિગ્નલને બોઇલરપ્લેટ નોઈઝમાં ડૂબાડી દે છે. લેટન્સી (વિલંબ) વધી જાય છે. વપરાશકર્તાઓ રાહ જુએ છે, અસ્વસ્થ થાય છે અને છોડીને જાય છે. અમે આ મુશ્કેલીનો સામનો કર્યો. તેથી અમે રિટ્રીવલ લેયરને સંપૂર્ણપણે તોડી નાખ્યું અને તેને એક માપન કરી શકાય તેવા, ટ્યુનેબલ સિસ્ટમ તરીકે ફરીથી બનાવ્યું. પરિણામ એવું પાઇપલાઇન આવ્યું જે યુઝર એક્સપિરિયન્સને સ્લાઇડશો બનાવ્યા વગર 95% રિકોલ (recall) પ્રાપ્ત કરે છે.
પ્રોડક્શનમાં RAG ડેમો કેમ નિષ્ફળ જાય છે
હોબી પ્રોજેક્ટ્સ અને પ્રારંભિક તબક્કાના ઉત્પાદનોમાં સ્ટાન્ડર્ડ સ્ટેક આશ્ચર્યજનક રીતે સમાન છે: ફિક્સ્ડ ટોકન ચંક્સ, ઓફ-ધ-શેલ્ફ એમ્બેડિંગ્સ અને એક સિંગલ વેક્ટર સર્ચ કોલ. આ સરળતા આકર્ષક છે, અને તે ત્યારે કામ કરે છે જ્યારે તમારો કોર્પસ (corpus) સ્વચ્છ, નાનો અને વ્યાકરણની દ્રષ્ટિએ અનુમાનિત હોય. પ્રોડક્શન ડેટા આમાંથી કંઈ જ નથી. 512 ટોકનનો એક ફિક્સ્ડ ચંક SaaS કોન્ટ્રાક્ટમાં વળતર આપવાના ક્લોઝ (indemnification clause) ની બરાબર વચ્ચેથી કપાઈ જશે. અચાનક તમારું રિટ્રીવલ લેયર લેંગ્વેજ મોડેલને કાયદાકીય જવાબદારીનો અડધો ભાગ આપશે અને તેને જવાબદારીના પ્રશ્નનો જવાબ આપવા કહેશે. મોડેલ હેલ્યુસિનેટ (hallucinate) કરે છે કારણ કે સંદર્ભ તૂટી ગયો છે.
મોટા ટેકનિકલ દસ્તાવેજો આ સમસ્યાને વધુ ગંભીર બનાવે છે. API ડોક્યુમેન્ટેશન ફંક્શન સિગ્નેચર્સ, ટેબલ્સ અને કોડ બ્લોક્સથી ભરેલું હોય છે. એક ફિક્સ્ડ વિન્ડો કદાચ TypeScript ઇન્ટરફેસની વચ્ચેનો ભાગ પકડી શકે છે પરંતુ તેની ઉપરના ફંક્શનનું નામ અને તેની નીચેના વપરાશના ઉદાહરણને ચૂકી શકે છે. એમ્બેડિંગ વેક્ટર વપરાશકર્તા જે ક્ષમતા વિશે પૂછી રહ્યા છે તેના બદલે સિન્ટેક્સ ફ્રેગમેન્ટ્સ અને ઇનલાઇન નોઈઝનું પ્રતિનિધિત્વ કરવા લાગે છે. ગાર્બેજ ઇન, હેલ્યુસિનેશન આઉટ (Garbage in, hallucination out).
ટોકન કાઉન્ટ દ્વારા નહીં, પણ સ્ટ્રક્ચર દ્વારા ચંકિંગ કરો
અમે કરેલો પહેલો ફેરફાર એ હતો કે ચંક્સને માત્ર ટોકન્સના થેલા તરીકે જોવાનું બંધ કર્યું. ચંક્સ એ સેમેન્ટિક યુનિટ્સ (અર્થપૂર્ણ એકમો) છે. યોગ્ય વ્યૂહરચના સંપૂર્ણપણે તમે શું ઇન્ડેક્સ કરી રહ્યા છો તેના પર આધાર રાખે છે.
કાયદાકીય દસ્તાવેજો માટે, અમે રિકર્સિવ ચંકિંગ (recursive chunking) અપનાવ્યું જે દસ્તાવેજ પદાનુક્રમ (hierarchy) નું સન્માન કરે છે. તે સેક્શન, સબ-સેક્શન અને ક્લોઝને સીમાઓ તરીકે ગણે છે. એક ક્લોઝ અખંડ રહે છે કારણ કે ક્લોઝ એ અર્થનો એક એકમ છે. જો તમે તેને વચ્ચેથી કાપો છો, તો કાયદાકીય તર્ક ખોવાઈ જાય છે.
API ડોક્યુમેન્ટેશન માટે, સ્ટ્રક્ચર-અવેર ચંકિંગ ફંક્શન્સ, ક્લાસ અને એન્ડપોઇન્ટ્સને એટમિક (atomic) તરીકે ગણે છે. એક ચંકમાં ફંક્શન સિગ્નેચર, તેના આર્ગ્યુમેન્ટ્સ અને તેની ડોકસ્ટ્રિંગ હોઈ શકે છે. તે માત્ર ટોકન કાઉન્ટ વધી ગયો છે એટલે પછીના યુટિલિટી ફંક્શનમાં અંધાધૂંધ ભળી જતું નથી. આનાથી એમ્બેડિંગ ચોક્કસ ક્ષમતા પર કેન્દ્રિત રહે છે.
સપોર્ટ ટિકિટ્સ વધુ અસ્તવ્યસ્ત હોય છે. તેઓ વાતચીત સ્વરૂપના, થ્રેડેડ અને નોન-લીનિયર હોય છે. ફિક્સ્ડ ચંક્સ એન્જિનિયરના સ્ટેટસ અપડેટ અને તે જ થ્રેડમાંથી ગ્રાહકની ફરિયાદ બંનેને પકડી લેશે અને એવો દેખાવ કરશે કે તેઓ એક સુસંગત એકમ બનાવે છે. અમે સેમેન્ટિક ચંકિંગ (semantic chunking) પર સ્વિચ કર્યું, જ્યાં ટોકન બજેટ ખતમ થાય ત્યારે નહીં, પરંતુ જ્યારે વિષય અથવા વક્તા બદલાય ત્યારે વિભાજન થાય છે.
ઇન્ટરનલ વિકિ ઘણીવાર સંસ્થામાં સૌથી અસ્તવ્યસ્ત ડેટા હોય છે. ફોર્મેટિંગ અસંગત હોય છે, હેડર્સ ખૂટે છે અને સેક્શન એકબીજામાં ભળી જાય છે. આ માટે, અમે LLM-આધારિત ચંકિંગનો ઉપયોગ કરીએ છીએ. એક નાનું મોડેલ એમ્બેડિંગ જનરેટ કરતા પહેલા જ આગળ વાંચે છે અને તાર્કિક સીમાઓ ઓળખે છે. તે કેરેક્ટર સ્પ્લિટ કરતા વધુ ખર્ચાળ છે, પરંતુ રિટ્રીવલની ગુણવત્તા તરત જ તેની કિંમત વસૂલ કરી લે છે.
હાઇબ્રિડ રિટ્રીવલ: સિગ્નલ્સને સંયોજિત કરો
વેક્ટર સર્ચ શક્તિશાળી છે પરંતુ તેમાં બ્લાઇન્ડ સ્પોટ્સ (blind spots) છે. જો તમે ERR_CONNECTION_REFUSED_0x800 જેવો ચોક્કસ એરર કોડ પેસ્ટ કરો છો, તો સિમિલારિટી સર્ચ કોઈ અસંબંધિત મોડ્યુલ માટે ટ્રબલશૂટિંગ ગાઇડ આપી શકે છે કારણ કે એમ્બેડિંગ સ્પેસે તેમને નજીક ક્લસ્ટર કર્યા છે. એક્ઝેક્ટ મેચ મહત્વના છે, અને માત્ર વેક્ટર સર્ચ તેમને અવગણી શકે છે.
BM25 સાથે કીવર્ડ સર્ચ એક્ઝેક્ટ-મેચની સમસ્યાને સુંદર રીતે ઉકેલે છે. પરંતુ તે વૈચારિક અંતર (conceptual distance) માં અટકી જાય છે. જો વપરાશકર્તા "heavy load હેઠળ પરફોર્મન્સમાં ઘટાડો" વિશે પૂછે છે, તો BM25 એ નિદાન નોંધ (diagnostic note) ચૂકી જશે જે "ટ્રાફિક સ્પાઇક્સ દરમિયાન ધીમી થ્રુપુટ"નું વર્ણન કરે છે, કારણ કે ત્યાં પૂરતો કીવર્ડ ઓવરલેપ નથી.
અમે કોઈ એક પક્ષ પસંદ કરવાનું બંધ કર્યું અને બંનેને સમાંતર રીતે ચલાવવાનું શરૂ કર્યું. વેક્ટર અને કીવર્ડ સર્ચ દરેક તેમની પોતાની રેન્ક્ડ લિસ્ટ આપે છે. અમે તેને Reciprocal Rank Fusion (RRF) સાથે મર્જ કરીએ છીએ. RRF તેની અસરકારકતામાં સરળ અને સચોટ છે. તે દરેક દસ્તાવનને તે દરેક લિસ્ટમાં ક્યાં છે તેના આધારે સ્કોર આપે છે. જે દસ્તાવતો બંને સિસ્ટમમાં ટોચ પર હોય છે તેમને મોટો બૂસ્ટ મળે છે. જે દસ્તાવતોને માત્ર એક જ એન્જિન પસંદ કરે છે, તેઓ પણ અંતિમ ઉમેદવાર સેટમાં સ્થાન મેળવે છે.
After fusion, we run the top candidates through a cross-encoder reranker. This is not free. It adds roughly 50 milliseconds of compute. It also increases recall by 15%. The cross-encoder evaluates the full query and each candidate chunk together, producing a relevance score far more nuanced than a bi-encoder embedding ever could. That extra 50 milliseconds is a bargain. It prevents you from shipping a garbage context window to the LLM and spending two seconds waiting for a confused or hallucinated answer.
Fix the Query Before You Search
Users do not write queries like search engineers. They type "app broken." They paste cryptic log fragments. They ask vague, ambiguous questions. If you send those raw strings straight to the index, you get garbage back.
We transform every query before it touches the retrieval engine.
First, query expansion. The system generates multiple search terms from a single short question. A user asks, "How do I fix the timeout?" The engine expands that to cover connection timeouts, read timeouts, gateway timeouts, and retry logic. This approach alone moved our recall from 78% to 96%.
Second, query decomposition. Complex questions get broken into smaller sub-questions. A query like "What's the refund policy for enterprise customers past 90 days and how does it differ from monthly plans?" becomes two focused searches rather than one bloated embedding lookup. Each sub-question hits the index independently. The results are stitched back together downstream. This keeps retrieval narrow and precise, which stops the dilution that happens when a single embedding tries to match a dozen concepts at once.
Let Bayesian Search Tune Your Pipeline
If you are still hand-tuning chunk size, overlap ratios, and retrieval weights, you are leaving performance on the table. We stopped guessing.
We defined a search space where chunk size, overlap percentage, vector-versus-BM25 weights, and reranker thresholds are all variables. Then we applied Bayesian optimization. Instead of grid-searching through hundreds of random configurations, Bayesian search builds a probabilistic model of what works. It proposes a configuration, observes the recall and latency, updates its beliefs, and proposes the next one. Over time it converges on balances a human would never stumble into manually.
It found combinations we never would have tried. Smaller chunks with heavier overlap. A slightly lower weight on dense vector search paired with a more aggressive reranker threshold. These non-obvious tradeoffs gave us both higher recall and lower latency.
This is not a one-time setup task. We re-run hyperparameter optimization monthly. Your corpus drifts. User behavior shifts. Your pipeline should adapt instead of rusting in place.
The Payoff
The raw output of that rebuild is hard to argue with.
Recall at position ten went from 78% to 95%. When the correct answer lives in our knowledge base, we surface it nineteen times out of twenty. Latency at the 95th percentile fell from 850 milliseconds to 320 milliseconds. The chat feels instant instead of ponderous.
Better retrieval gave the language model better grounding. Hallucination rate dropped from 12% to 3%. When the model has the right context in front of it, it stops inventing facts. Cost per query fell by 38%. Faster, sharper retrieval means fewer tokens wasted on irrelevant context, retry loops, and verbose but useless prompts.
Build It Like Infrastructure
If you are moving from prototype to production, treat retrieval as infrastructure code rather than a configuration
