Most teams build their first retrieval system the same way: slice every document into fixed 512-token chunks, push them into a vector database, and hope the embedding model does the hard work. That hope gets you through a demo. It does not survive contact with real users.
In production, a legal contract falls apart when you sever a liability clause from its exceptions. API documentation turns useless when a code sample gets detached from its function signature. A customer support thread becomes noise when you rip a single complaint out of its conversational history. The problem is rarely the language model sitting at the end of the pipeline. The problem is what you feed it.
We learned this the hard way. Our initial retrieval layer looked standard but behaved inconsistently. So we rebuilt it around a simple idea: treat retrieval as measured infrastructure, not magic. Here is exactly what changed, and how we pushed recall to 95 percent while cutting the 95th-percentile latency from 850 ms to 320 ms.
The Fixed-Chunk Trap
Uniform token counts are easy to code and easy to explain. That convenience masks a basic truth: documents have structure. When you ignore that structure, you destroy signal.
Consider a ten-page master service agreement. A fixed 512-token slice will land mid-obligation, splitting a clause from the very cap table that limits it. The retrieval step then returns half a thought. The generator hallucinates the rest. In API documentation, a chunk that is too large dilutes the embedding with boilerplate headers, burying the specific method a developer needs. In support tickets, a fixed window treats a conversation as a bag of sentences, stripping away the back-and-forth that reveals what actually failed.
We stopped treating chunk size as a hyperparameter we guessed. We started treating it as a mapping exercise between the document type and the information architecture inside it.
Match Your Chunking to the Data
The fix is not one perfect chunk size. The fix is three distinct strategies tuned to three distinct data shapes.
Legal documents now go through recursive splitting. The algorithm first looks for the largest natural boundaries—sections, then subsections, then numbered clauses—and only falls back to smaller splits when necessary. This keeps a termination clause attached to its survival conditions. The retrieval step sees complete logical units, which sharply reduces the model’s temptation to invent missing exceptions.
API and code documentation get structure-aware chunking. Markdown headers, code fences, and parameter tables are parsed as atomic units. We do not split inside a code block. We keep docstrings adjacent to their signatures. The result is that a query for a specific class method retrieves the full context a developer needs: the description, the typed parameters, and the working example.
Support and conversational data use semantic chunking. Instead of counting tokens, we look for shifts in topic or intent. If a customer describes a bug in message three and pastes a stack trace in message seven, we chunk by meaning, not by message index. The retrieval layer then returns the full arc of the problem rather than a orphaned sentence.
Why Vector Search Alone Fails
Even perfect chunks die in a pure vector search. Dense embeddings excel at capturing meaning and synonymy, but they are notoriously fuzzy on exact strings. If an engineer searches for the precise error code ERR_CONNECTION_REFUSED, vector similarity might return a dozen conceptual neighbors and miss the exact match buried at rank fourteen.
Keyword search with BM25 has the opposite problem. It finds exact tokens but misses semantic intent. A user asking “why is my database down” will never match a document that says “troubleshooting connection timeouts.”
We now run both. Vector and keyword results are fed into Reciprocal Rank Fusion, which blends the two ranked lists without requiring calibrated scores. The fused list is then passed through a cross-encoder reranker. The reranker is slower than the initial retrieval, but it is far more precise because it judges query-document relevance directly rather than through compressed embeddings. That hybrid pipeline alone lifted our recall by 15 percent.
Fixing Bad Queries Before They Hit the Index
Pengguna tidak menulis kueri pencarian yang ideal. Mereka menempelkan baris log yang terpotong. Mereka mengetik “ini rusak.” Mereka menggunakan jargon yang tidak pernah diadopsi oleh dokumentasi Anda. Jika Anda memercayai kueri mentah, Anda memercayai kebisingan (noise).
Kami sekarang memperluas setiap kueri yang masuk menjadi tiga hingga lima variasi sebelum mengirimkannya ke lapisan retrieval. Satu variasi mungkin berupa parafrasa langsung. Variasi lainnya mungkin berupa judul dokumen ideal yang hipotetis. Variasi ketiga menghilangkan kata pengisi percakapan dan mengisolasi kata kunci teknis. Setiap varian di-embed dan dicari. Kami kemudian melakukan deduplikasi dan menggabungkan kumpulan kandidat tersebut.
Ini tidak gratis. Panggilan embedding tambahan tersebut memakan biaya dan menambah beberapa milidetik. Namun efeknya terhadap recall sangat dramatis: kami bergerak dari 78 persen ke 96 persen dengan memperluas kueri sebelum retrieval. Karena retrieval yang lebih baik memperkecil jendela generasi dan mendasarkan model pada konteks yang benar, kami akhirnya menghemat biaya di hilir (downstream). Langkah retrieval yang sedikit lebih mahal lebih murah daripada langkah generasi yang panjang dan berhalusinasi.
Berhenti Menebak. Mulai Mencari.
Setelah kami menerapkan chunking, hybrid retrieval, dan query expansion yang tepat, kami masih menghadapi kekacauan kombinatorial. Ukuran chunk, overlap chunk, kedalaman retrieval top-k, cutoff reranker, dan bobot fusi semuanya saling berinteraksi. Pencarian grid manual akan memakan waktu berminggu-minggu dan tetap akan menyisakan kita pada nilai maksimum lokal (local maximum).
Kami beralih ke optimasi Bayesian untuk menjelajahi ruang tersebut. Alih-alih menguji setiap kombinasi secara menyeluruh, algoritma pencarian mempertahankan keyakinan tentang konfigurasi mana yang kemungkinan besar akan berkinerja baik dan secara progresif menyempit ke wilayah yang menjanjikan.
Hasilnya bukanlah satu pengaturan tunggal yang sempurna. Hasilnya adalah Pareto frontier dari pilihan-pilihan yang ada. Di satu sisi, kami memiliki konfigurasi ramping yang dioptimalkan untuk endpoint dukungan API high-throughput kami: inferensi cepat, recall moderat, dan latensi serendah mungkin. Di sisi lain, kami memiliki konfigurasi agresif untuk peninjauan hukum: retrieval yang lebih dalam, reranking yang lebih berat, dan overlap yang lebih ketat, menukar milidetik dengan ketelitian. Karena frontier tersebut eksplisit, kami dapat memilih titik yang tepat untuk produk tersebut alih-alih berpura-pura bahwa satu ukuran cocok untuk semua (one size fits all).
Seperti Apa Angka-angkanya Sebenarnya
Perubahan ini mengubah sistem dari prototipe yang rapuh menjadi pipeline produksi yang terukur.
Recall pada sepuluh meningkat dari 78 persen ke 95 persen. Itu berarti ketika jawaban yang benar ada dalam korpus kami, kami menemukannya sembilan belas kali dari dua puluh percobaan.
Latensi pada persentil ke-95 turun dari 850 ms menjadi 320 ms. Stack hybrid terdengar lebih berat di atas kertas, tetapi pengindeksan yang lebih cerdas, reranker yang lebih kecil, dan kemampuan untuk menyajikan chunk agresif hanya saat dibutuhkan membuat seluruh sistem menjadi lebih cepat.
Tingkat halusinasi—yang dilacak oleh anotator manusia pada golden dataset yang disisihkan—turun dari 12 persen menjadi 3 persen. Ketika model menerima konteks yang lengkap dan relevan, ia berhenti mengarang fakta.
Biaya per kueri turun dari $0,008 menjadi $0,005. Retrieval yang lebih baik berarti prompt LLM yang lebih pendek dan lebih terfokus serta lebih sedikit upaya pemulihan. Pengeluaran embedding tambahan untuk query expansion jauh lebih kecil dibandingkan penghematan pada tahap generasi.
Bangun Golden Dataset dan Perlakukan Retrieval Seperti Kode
Jika Anda mengambil satu hal dari sini, itu adalah disiplin pengukuran. Kami membangun golden dataset kecil berisi pertanyaan nyata dan lokasi jawaban yang terverifikasi. Sebelum perubahan apa pun masuk ke produksi, ia dijalankan terhadap dataset tersebut. Recall dan latensi dipantau secara real-time, bukan sekadar dilihat sekilas di notebook.
Retrieval bukanlah demo riset. Ia adalah infrastruktur. Ia layak mendapatkan unit test, benchmark regresi, dan optimasi otomatis sama seperti bagian lain dari stack Anda. Lakukan chunking berdasarkan struktur dokumen, bukan berdasarkan takhayul jumlah token. Gabungkan pencarian vektor dan kata kunci dengan reranker. Perluas kueri yang sebenarnya ditulis oleh pengguna Anda. Kemudian biarkan algoritma pencarian yang menyetel parameternya, bukan intuisi Anda.
Pipeline yang kami jelaskan bukanlah teori. Anda dapat membaca tulisan aslinya di sini, dan jika Anda ingin mendiskusikan rekayasa retrieval dengan komunitas yang peduli dengan hal ini, grup GyaanSetu AI terbuka untuk Anda.
