RAG-க்கு மரண அறிவிப்புகளை மக்கள் தொடர்ந்து எழுதி வருகின்றனர். நீளமான context windows அதை அழித்துவிட்டன, Agents அதை மாற்றீடு செய்துவிட்டன போன்ற தலைப்புகளை நீங்கள் இந்நேரம் பார்த்திருக்கலாம். உண்மை என்னவென்றால், அது மிகவும் குறுகிய மற்றும் மிகவும் பயனுள்ளது. RAG இறந்துவிடவில்லை. ஆவணங்களின் குவியலைத் துண்டுகளாகப் பிரித்து, அவற்றை ஒரு vector database-ல் செலுத்தி, திடீரென்று ஒரு நம்பகமான, உண்மையான AI-யைப் பெற்றுவிடலாம் என்ற வசதியான மாயையே உண்மையில் சரிந்துவிட்டது.

சில ஆண்டுகளுக்கு முன்பு, அதன் எளிமை மிகவும் வசீகரிப்பதாக இருந்தது. உங்கள் அறிவுத் தளத்தை (knowledge base) Embed செய்யுங்கள். அதை ஒரு LLM உடன் இணைக்கவும். ஒரு கேள்வியைக் கேளுங்கள், உங்கள் தரவைப் பயன்படுத்தி மாடல் பதிலளிப்பதைப் பாருங்கள். கட்டுப்படுத்தப்பட்ட டெமோக்கள் மற்றும் சிறிய FAQ bots-களுக்கு இது உண்மையாகவே வேலை செய்தது. இருபது பக்கங்கள் கொண்ட ஒரு உதவி மைய கையேடு அல்லது ஒரு சிறிய உள் விக்கி (internal wiki) போன்றவற்றுக்கு இது சரியாக இருக்கும். அந்த பாட் பெரும்பாலும் சரியான பத்தியைக் குறிப்பிடும், மேலும் நிர்வாகமும் அந்தத் திட்டத்திற்கு ஒப்புதல் அளிக்கும். ஆனால் சோதனைத் திட்டங்கள் (pilots) என்பது உற்பத்தி நிலை (production) அல்ல. முன்மாதிரிகள் (Prototypes) உண்மையான வணிகச் செயல்பாடுகளின் சவால்களைக் கொண்டிருக்காது.

உற்பத்தித் தரவு (Production data) மிகவும் சிக்கலானது. ஒரே சிக்கலைத் தீர்க்கும் குறிப்பு பல கோப்புகளில் வெவ்வேறு நேர முத்திரைகளுடனும், முரண்பட்ட நிலைகளுடனும் (status labels) நகலெடுக்கப்பட்டிருக்கலாம். பக்கங்களுக்கு இடையே சிதறிக்கிடக்கும் சிக்கலான அட்டவணைகள், ஒரு splitter அவற்றை நடுவில் வெட்டும்போது அர்த்தமற்ற தகவல்களை உருவாக்கும். முரண்பாடுகளை அது அப்படியே வைத்திருக்கும். 2023 கொள்கை கையேடு ஒன்றைச் சொல்கிறது, மார்ச் 2024 திருத்தம் வேறொன்றைச் சொல்கிறது. பழைய PDF ஒருபோதும் ஆவணப்படுத்தப்படாமல் இருக்கலாம். துண்டாக்குதல், சேமித்தல் மற்றும் மீட்டெடுத்தல் (chunk, store, and retrieve) என்ற எளிய முறை ஒவ்வொரு பத்தியையும் ஒரு தனித்த தீவாகவே treats செய்கிறது. அதற்கு படிநிலை (hierarchy), பதிப்பு வரலாறு (version history) அல்லது முரண்பாடுகளைத் தீர்ப்பது (conflict resolution) பற்றிய புரிதல் இல்லை. மாடல் தவறான தகவல்களை (hallucinate) உருவாக்குவது LLM பழுத்ததனால் அல்ல, மாறாக அதற்கு வழங்கப்பட்ட சூழல் (context) சிதறிய அல்லது தவறானதாக இருந்ததாலேயே.

மில்லியன் கணக்கான டோக்கன்களைக் கொண்ட context windows மீட்டெடுப்பிற்குத் தேவையில்லை என்று சிலர் கூறுகின்றனர். அவர்களின் வாதம் எளிமையானது: முழுத் தரவையும் (corpus) ப்ராம்ப்ட்டிற்குள் (prompt) தள்ளிவிடுங்கள், மாடல் அனைத்தையும் படிக்கட்டும். இது கேட்பதற்கு நேர்த்தியாகத் தோன்றலாம், ஆனால் இது ஆபத்தான ஒரு நம்பிக்கையாகும். ஒரு மாடலால் ஒரு சிறிய நாவலுக்கு இணையான உரையைத் தொழில்நுட்ப ரீதியாக உள்வாங்க முடியலாம், ஆனால் அந்தப் பரப்பிற்கு நடுவில் ஒரு குறிப்பிட்ட வாசகத்தைக் கண்டறிவது முற்றிலும் வேறுபட்ட திறன். வைக்கோல் குவியலில் ஊசியைத் தேடுவது போன்றது இது. நீளமான context windows கிடைக்கும் பரப்பளவை விரிவுபடுத்துகின்றனவே தவிர, அந்தப் பரப்பில் எதற்கு இடம் கொடுக்க வேண்டும் என்ற கடினமான வேலையைத் தீர்க்கவில்லை. பிரச்சனை மீட்டெடுப்பு மட்டுமல்ல; அது எப்போதும் சூழலை ஒருங்கிணைப்பதே (context assembly) ஆகும்.

எளிய மீட்டெடுப்பிலிருந்து சூழல் பொறியியல் வரை (From Naive Retrieval to Context Engineering)

2026-ல், இந்தத் துறை முதிர்ச்சியடைந்து வருகிறது. RAG என்பது ஒரு நேரியல் குழாய்முறை (linear pipeline) என்பதிலிருந்து மாறி, சூழலை ஒரு பொறியியல் ரீதியாக வடிவமைக்கப்பட்ட தயாரிப்பாகக் கருதும் கட்டமைப்பை நோக்கி நாம் நகர்கிறோம்.

தூய semantics-ஐ விட கலப்புத் தேடல் (Hybrid search over pure semantics). நோக்கத்தைப் புரிந்துகொள்ள semantic similarity சிறந்தது, ஆனால் துல்லியமான அடையாளக் குறிகளில் (identifiers) அது தடுமாறலாம். ஒரு பொறியாளர் ERR_CONNECTION_REFUSED போன்ற ஒரு குறிப்பிட்ட பிழை குறியீட்டை அல்லது v3.2.1 போன்ற மென்பொருள் பதிப்பைக் கேட்டால், தூய vector search, கருத்தியல் ரீதியாக ஒத்த ஆனால் நடைமுறையில் தேவையற்ற முடிவுகளின் கடலில் துல்லியமான பொருத்தத்தை மழுங்கச் செய்துவிடும். நவீன அமைப்புகள் dense vector retrieval உடன் BM25 அல்லது inverted indexes போன்ற முறைகளைப் பயன்படுத்தி keyword search-ஐ இணைக்கின்றன. துல்லியமான பெயர்கள், பிழை குறியீடுகள், பதிப்பு சரங்கள் (version strings) மற்றும் தயாரிப்பு ஐடிகள் ஆகியவை keyword அடுக்கு மூலம் கண்டறியப்படுகின்றன, அதே சமயம் கருத்தியல் நுணுக்கங்கள் vector அடுக்கு மூலம் கையாளப்படுகின்றன.

உருவாக்கத்திற்கு முன் மறுவரிசைப்படுத்துதல் (Reranking before generation). மீட்டெடுப்பு என்பது இயல்பாகவே அதிக தகவல்களைத் திரட்டும் (recall) நோக்கில் இருக்கும். ஒரு முக்கியமான பத்தியைத் தவறவிட்டுவிடுவோமோ என்ற பயத்தில் நாற்பது அல்லது ஐம்பது துண்டுகளை நீங்கள் எடுக்கக்கூடும். ஆனால் அந்த இரைச்சலை (noise) ஒரு பெரிய மாடலுக்கு வழங்குவது டோக்கன்களை வீணடிக்கும் மற்றும் முக்கியத் தகவலை மறைத்துவிடும். Reranking என்பது ஒரு இரண்டாவது, பொதுவாகச் சிறிய மாடலைப் பயன்படுத்தி ஒவ்வொரு முடிவையும் குறிப்பிட்ட வினாவிற்கு ஏற்ப அதன் பொருத்தத்திற்கு மதிப்பீடு செய்வதன் மூலம் இதைத் தீர்க்கிறது. முதல் ஐந்து பத்திகள் மட்டும் அடுத்த கட்டத்திற்குச் செல்லும், மற்றவை நீக்கப்படும். இது மீட்டெடுப்பிற்கும் உருவாக்கத்திற்கும் இடையில் ஒரு துல்லியமான வடிகட்டியாகச் செயல்பட்டு, விலையுயர்ந்த reasoning மாடல் உண்மையில் முக்கியமானவற்றை மட்டுமே படிப்பதை உறுதி செய்கிறது.

பொருளைப் பாதுகாக்கும் சூழல் சார்ந்த மீட்டெடுப்பு (Contextual retrieval that preserves meaning). துண்டாக்குதல் (Chunking) என்பது ஒரு வன்முறைச் செயல் போன்றது. ஒரு splitter ஒரு பத்தியைத் அதன் தலைப்பிலிருந்தோ, அட்டவணை விளக்கத்திலிருந்தோ, அதைச் சுற்றியுள்ள சட்டத் தகவல்களிலிருந்தோ அல்லது அதன் பொருளை மாற்றும் அடிக்குறிப்பிலிருந்தோ துண்டித்துவிடலாம். Contextual retrieval என்பது துண்டுகளை மாடலுக்கு அனுப்பும் முன்பே மெட்டாடேட்டாவை (metadata) இணைப்பதன் மூலம் இதைச் சரிசெய்கிறது. உதாரணமாக, அதன் மூலத் தன்மையைக் குறிக்கும் தகவல்களைச் சேர்க்கலாம்: இந்தக் குறிப்பு Q3 2024 சம்பவ அறிக்கை, Database Outage பகுதி, Severity Critical என்பதன் অন্তর্গত. இப்போது மாடல் ஒரு தனித்த வாக்கியத்தை மட்டும் பார்க்காமல், அது எந்தச் சூழலில் உள்ளது என்பதையும் பார்க்கிறது. அந்தத் துண்டுத் தகவல் மீண்டும் அதன் முழுப் பொருளைப் பெறுகிறது.

Modular routing by intent. Not every question belongs in a vector store full of documentation. A user asking how to reset a password probably needs a help article. A user asking why revenue dropped in the Northeast last quarter needs SQL against a data warehouse, not a semantically similar paragraph about regional sales strategy. Mature systems now route queries by intent, selecting the appropriate tool. Documentation for procedure. Relational databases for structured analytics. Log aggregators for trace debugging. APIs for live status. The retrieval layer becomes a dispatcher, not a monoculture.

Agentic reasoning loops. Some questions cannot be answered by a single search step. They require reformulation. A vague initial query gets clarified. Retrieved claims get cross-checked against a second source. If the documentation contradicts the API specification, the system flags the conflict instead of fabricating a middle ground. The model decides when to search again, when to refine its query, and when it has collected enough evidence to answer. This is not one-shot retrieval. It is structured reasoning that uses search as a subroutine.

GraphRAG for relational questions. Certain business questions are about connections, not sentences. Which component failure triggered which downstream alerts? Which supplier feeds which factory, and what is the alternate route? Who in the organization has decision rights over this specific budget line? Flat text chunks flatten these relationships because they were never designed to preserve topology. Knowledge graphs do. When the question is about influence, lineage, patterns, or network structure, traversing a graph provides context that no amount of paragraph retrieval can replicate.

The Questions That Actually Matter

The conversation around RAG needs to change. Stop asking how to build a generic RAG pipeline. Start asking what specific task the model must solve, what exact data it needs to be accurate, and how you verify that the assembled context is sufficient. These questions force you upstream into data quality, schema design, verification loops, and source provenance. They expose whether your knowledge base is even fit for automated consumption.

RAG is no longer a single linear process you install once and forget. It is a discipline of assembling the right context so a model can reason effectively. That means treating retrieval as a system design problem, not a library import.

The tools are getting sharper. Search is hybrid. Routing is intelligent. Retrieval is ranked, enriched, and verified. The simple illusions of 2022 had to collapse so that something genuinely useful could take their place. Your job now is not merely to retrieve text from a database. It is to build systems that know what the model needs before the model begins to think.

If you are building in this space, the GyaanSetu learning community is a place to trade practical notes with people solving the same problems: https://t.me/GyaanSetuAi