People keep writing obituaries for RAG. You have probably seen the headlines by now. Long context windows killed it. Agents replaced it. The whole pattern is obsolete. The truth is narrower and far more useful. RAG did not die. What actually collapsed was the comfortable illusion that you could split a pile of documents into chunks, feed them into a vector database, and suddenly own a reliable, truthful AI.

A few years ago, the pitch was seductive in its simplicity. Embed your knowledge base. Connect it to an LLM. Ask a question, and watch the model answer using only your retrieved data. For controlled demos and small FAQ bots, this honestly worked. A twenty-page help desk manual. A tidy internal wiki. The bot would more or less cite the right paragraph, and leadership would sign off on the pilot. But pilots are not production. Prototypes do not contain the scars of real business operations.

Production data is messy. It contains the same troubleshooting note copied across a dozen files, each with slightly different timestamps and conflicting status labels. It holds complex tables that bleed across pages, rendering nonsense when a splitter cuts them down the middle. It preserves contradictions without apology. The 2023 policy manual says one thing. The March 2024 amendment says another. The old PDF was never archived. The naive pattern of chunk, store, and retrieve treats every paragraph as an isolated island. It has no feel for hierarchy, version history, or conflict resolution. The model hallucinates not because the LLM is broken, but because the context it received was fragmented, orphaned, or outright wrong.

Some observers claim that million-token context windows make retrieval irrelevant. Their argument is straightforward. Just dump the entire corpus into the prompt and let the model read it all. This sounds elegant. It is also dangerously optimistic. A model may technically be able to ingest a volume of text equivalent to a short novel, yet locating one specific clause in the middle of that expanse remains an entirely different capability. Needles stay lost in haystacks. Long context windows expand the available canvas, but they do not solve the hard work of deciding what deserves a spot on that canvas. The problem was never simply retrieval. It is, and always has been, context assembly.

From Naive Retrieval to Context Engineering

In 2026, the field is maturing. We are moving past the phase where RAG was treated as a single linear pipeline and toward an architecture that treats context as a deliberately engineered product.

Hybrid search over pure semantics. Semantic similarity is excellent for understanding intent, but it is sloppy with precise identifiers. If an engineer queries a specific error code like ERR_CONNECTION_REFUSED or a software version like v3.2.1, pure vector search can dilute the exact match in a sea of conceptually similar but practically irrelevant results. The evolution here is straightforward. Modern systems combine dense vector retrieval with keyword search, using methods like BM25 or inverted indexes alongside embeddings. Exact names, error codes, version strings, and product IDs get caught by the keyword layer while conceptual nuance is handled by the vector layer.

Reranking before generation. Retrieval is naturally biased toward recall. You pull in forty or fifty chunks because you are terrified of missing the one golden paragraph. But feeding all of that noise into a large model wastes tokens and buries signal. Reranking solves this with a second, usually smaller model that scores each candidate for relevance against the specific query. The top five passages advance. The rest are discarded. It acts as a precision filter between retrieval and generation, ensuring the expensive reasoning model only reads what actually matters.

Contextual retrieval that preserves meaning. Chunking is a violent act. A splitter can sever a paragraph from its section header, its table caption, its surrounding legal disclaimer, or the footnote that modifies its meaning. Contextual retrieval mitigates this by enriching fragments before they ever reach the model. You prepend metadata indicating provenance: This excerpt belongs to the Q3 2024 incident report, Database Outage section, Severity Critical. The model sees not just a floating sentence but a situated piece of information. The fragment regains its bearings.

Định tuyến mô-đun theo ý định. Không phải câu hỏi nào cũng thuộc về một kho lưu trữ vector chứa đầy tài liệu. Một người dùng hỏi cách đặt lại mật khẩu có lẽ cần một bài viết hướng dẫn. Một người dùng hỏi tại sao doanh thu ở vùng Đông Bắc giảm trong quý trước cần truy vấn SQL vào kho dữ liệu, chứ không phải một đoạn văn có ý nghĩa tương tự về chiến lược bán hàng khu vực. Các hệ thống trưởng thành hiện nay định tuyến các truy vấn theo ý định, lựa chọn công cụ phù hợp. Tài liệu cho các quy trình. Cơ sở dữ liệu quan hệ cho phân tích có cấu trúc. Trình tổng hợp nhật ký (log aggregators) để gỡ lỗi vết (trace debugging). API để lấy trạng thái trực tiếp. Lớp truy xuất trở thành một bộ điều phối, thay vì là một hệ thống đơn nhất.

Các vòng lặp suy luận mang tính tác nhân. Một số câu hỏi không thể được trả lời chỉ bằng một bước tìm kiếm duy nhất. Chúng đòi hỏi sự diễn đạt lại. Một truy vấn ban đầu mơ hồ sẽ được làm rõ. Các khẳng định được truy xuất sẽ được đối chiếu với nguồn thứ hai. Nếu tài liệu mâu thuẫn với đặc tả API, hệ thống sẽ đánh dấu xung đột thay vì tự tạo ra một giải pháp trung hòa giả tạo. Mô hình quyết định khi nào cần tìm kiếm lại, khi nào cần tinh chỉnh truy vấn và khi nào đã thu thập đủ bằng chứng để trả lời. Đây không phải là truy xuất một lần (one-shot retrieval). Đó là quá trình suy luận có cấu trúc sử dụng tìm kiếm như một chương trình con.

GraphRAG cho các câu hỏi mang tính quan hệ. Một số câu hỏi kinh doanh xoay quanh các mối liên kết, chứ không phải các câu văn. Sự cố của thành phần nào đã kích hoạt các cảnh báo hạ nguồn nào? Nhà cung cấp nào cung cấp cho nhà máy nào, và lộ trình thay thế là gì? Ai trong tổ chức có quyền quyết định đối với dòng ngân sách cụ thể này? Các phân đoạn văn bản phẳng làm mất đi các mối quan hệ này vì chúng chưa bao giờ được thiết kế để bảo toàn cấu trúc topo. Đồ thị tri thức (Knowledge graphs) thì có. Khi câu hỏi liên quan đến tầm ảnh hưởng, nguồn gốc, các mô hình hoặc cấu trúc mạng lưới, việc duyệt qua một đồ thị sẽ cung cấp ngữ cảnh mà không có lượng truy xuất đoạn văn nào có thể sao chép được.

Những câu hỏi thực sự quan trọng

Các cuộc thảo luận xoay quanh RAG cần phải thay đổi. Đừng hỏi cách xây dựng một đường ống (pipeline) RAG chung chung nữa. Hãy bắt đầu hỏi mô hình phải giải quyết tác vụ cụ thể nào, chính xác dữ liệu nào nó cần để đạt được độ chính xác, và làm thế nào để bạn xác minh rằng ngữ cảnh được tập hợp là đầy đủ. Những câu hỏi này buộc bạn phải đi sâu vào thượng nguồn: chất lượng dữ liệu, thiết kế lược đồ (schema), các vòng lặp xác minh và nguồn gốc dữ liệu. Chúng bộc lộ liệu cơ sở tri thức của bạn có thực sự phù hợp để tiêu thụ tự động hay không.

RAG không còn là một quy trình tuyến tính duy nhất mà bạn chỉ cần cài đặt một lần rồi quên đi. Đó là một kỷ luật về việc lắp ráp ngữ cảnh phù hợp để mô hình có thể suy luận hiệu quả. Điều đó có nghĩa là coi việc truy xuất như một vấn đề thiết kế hệ thống, chứ không phải là việc nhập một thư viện (library import).

Các công cụ đang trở nên sắc bén hơn. Tìm kiếm là hỗn hợp (hybrid). Định tuyến là thông minh. Truy xuất được xếp hạng, làm giàu và xác minh. Những ảo tưởng đơn giản của năm 2022 đã phải sụp đổ để nhường chỗ cho những thứ thực sự hữu ích. Công việc của bạn bây giờ không chỉ đơn thuần là truy xuất văn bản từ một cơ sở dữ liệu. Đó là xây dựng các hệ thống biết mô hình cần gì trước khi mô hình bắt đầu tư duy.

Nếu bạn đang xây dựng trong lĩnh vực này, cộng đồng học tập GyaanSetu là nơi để trao đổi các ghi chú thực tế với những người đang giải quyết cùng một vấn đề: https://t.me/GyaanSetuAi