บทเรียน RAG ส่วนใหญ่มักจบลงแค่ใน notebook พวกเขาโหลดไฟล์ PDF ที่จัดระเบียบมาอย่างดีไม่กี่ไฟล์ แบ่งข้อความเป็นทุกๆ หนึ่งพันตัวอักษร ยัดชิ้นส่วนเหล่านั้นลงใน vector database แล้วเรียกมันว่าสถาปัตยกรรม ในบ่ายวันศุกร์ เดโมเหล่านั้นอาจทำงานได้อย่างสมบูรณ์แบบ แต่เมื่อใช้งานจริง (production) pipeline เดิมนั้นกลับกลายเป็นภาระอย่างเงียบเชียบ
คอขวดที่แท้จริงในระบบการสืบค้น (retrieval system) มักไม่ใช่ตัวโมเดลหรือ prompt แต่คือการนำเข้าข้อมูล (ingestion) pipeline ของ RAG จะสามารถดึงข้อมูลมาได้เฉพาะสิ่งที่มันถูกป้อนเข้าไปเท่านั้น และหากข้อมูลที่ป้อนเข้าไปนั้นมีสัญญาณรบกวน (noisy), ล้าสมัย (stale), หรือไม่สมบูรณ์ โมเดลก็จะให้คำตอบที่ผิดพลาดอย่างมั่นใจ เมื่อผู้ใช้บ่นว่าบอทเกิดอาการหลอน (hallucinated) บ่อยครั้งความผิดพลาดนั้นไม่ได้อยู่ที่ตัวโมเดล แต่อยู่ที่ต้นทางใน data pipeline ที่ไม่มีใครเฝ้าติดตามอย่างใกล้ชิด
กับดักบนไวท์บอร์ด
แผนผังโครงสร้าง (Architecture diagrams) มักทำให้การนำเข้าข้อมูลดูเหมือนลูกศรเพียงเส้นเดียวที่เขียนว่า “Documents → Vector DB” แต่ในความเป็นจริงนั้นยุ่งเหยิงกว่ามาก ระบบต้นทางมีการเปลี่ยนแปลงโดยไม่มีการแจ้งล่วงหน้า เลย์เอาต์ HTML ถูกออกแบบใหม่ URL ถูกเปลี่ยนเส้นทางไปยังหน้า landing page ทั่วไป หรือแม้แต่ JavaScript frameworks ที่เปลี่ยนเนื้อหาหลังจากตอบสนอง HTTP ครั้งแรก การมองว่าการนำเข้าข้อมูลเป็นงานที่ทำครั้งเดียวจบคือความผิดพลาดแรก แต่มันคือปัญหาด้านวิศวกรรมข้อมูล (data engineering) ที่ดำเนินไปอย่างต่อเนื่อง และสมควรได้รับการดูแลอย่างเข้มงวดพอๆ กับ pipeline ของ ETL ใดๆ
ทำไมความล้มเหลวของ RAG มักเกิดจากความล้มเหลวของการป้อนข้อมูล
ลองนึกภาพตามนี้: ผู้ใช้ถามผู้ช่วยภายในองค์กรเกี่ยวกับนโยบายการคืนเงินปัจจุบัน โมเดลไปดึงข้อมูลชิ้นส่วน (chunk) แรกสุดจาก vector store และระบุว่ามีระยะเวลา 30 วัน แต่จริงๆ แล้วนโยบายเพิ่งเปลี่ยนเป็น 60 วันเมื่อไตรมาสที่แล้ว LLM ไม่ได้แต่งคำตอบผิดขึ้นมาเอง แต่มันเชื่อข้อมูลที่ผิด ชั้นการสืบค้น (retrieval layer) ส่งหน้าเว็บเก่ามาให้ และเนื่องจาก embedding ดูมีความหมายใกล้เคียงกัน (semantically close) มากพอ โมเดลจึงถือว่าข้อมูลนั้นเป็นความจริงอ้างอิง (ground truth)
รูปแบบนี้เกิดขึ้นซ้ำแล้วซ้ำเล่า ทีมงานเสียเวลาหลายชั่วโมงไปกับการปรับค่า temperature และ top-k ทั้งที่คลังข้อมูล (corpus) ของพวกเขานั้นเต็มไปด้วยส่วนท้ายของหน้าเว็บ (navigation footers), ข่าวประชาสัมพันธ์ที่ซ้ำซ้อน, และชิ้นส่วนข้อมูลที่ตัดตารางออกเป็นสองส่วน ก่อนที่คุณจะเริ่มปรับแต่งการสร้างคำตอบ (generation) ให้ตรวจสอบก่อนว่าระบบของคุณได้รับอนุญาตให้รับรู้ข้อมูลอะไรบ้าง
7 กับดักที่ทำลายการนำเข้าข้อมูล
1. การรันครั้งแรกคือเรื่องโกหก
เครื่องหมายถูกสีเขียวจากการ crawl ครั้งแรกแทบไม่มีความหมาย ข้อมูลใน production นั้นมีชีวิต หน้าเอกสารมีการปรับปรุงโครงสร้าง (refactored), ลิงก์ถาวรของบล็อกเสีย, และ sitemaps อาจจะตัดบางส่วนออกไปอย่างเงียบๆ หากคุณตรวจสอบเพียงแค่ว่า pipeline ทำงานเสร็จสิ้นโดยไม่เกิด error คุณก็เหมือนกำลังบินแบบหลับตา คุณจำเป็นต้องตรวจสอบผลลัพธ์ (output) ตรวจสอบว่าเอกสารที่คาดหวังนั้นอยู่ครบถ้วนหรือไม่ โครงสร้างของมันยังสามารถ parse ได้เหมือนเดิมไหม และปริมาณข้อความทั้งหมดไม่ได้ลดฮวบลงเพียงเพราะแหล่งข้อมูลตัดสินใจเปลี่ยนวิธีการแบ่งหน้า (pagination) ผลลัพธ์
2. การ Crawl ไม่ใช่การนำเข้าข้อมูล
การดึง HTML เป็นส่วนที่ง่ายที่สุด การ crawl แบบดิบๆ (raw crawl) จะเก็บทุกอย่างมาหมด: ทั้งแบนเนอร์คุกกี้, แถบด้านข้าง "บทความที่เกี่ยวข้อง", บล็อกโฆษณา และประกาศลิขสิทธิ์ที่ส่วนท้ายหน้า หากคุณแบ่ง chunk จาก HTML ดิบๆ อย่างไม่ระวัง ทุกๆ ชิ้นส่วนของข้อความจะมีเศษเสี้ยวของเมนูนำทางติดมาด้วย เมื่อผู้ใช้ถามเกี่ยวกับ API rate limits ตัว retriever อาจจะดึงชิ้นส่วนที่มีลิงก์แถบด้านข้างอยู่ถึง 40 เปอร์เซ็นต์ขึ้นมา การสกัดข้อมูลที่สะอาด (clean extraction) จึงเป็นเรื่องสำคัญ คุณต้องระบุพื้นที่เนื้อหาหลัก, ตัดส่วนที่เป็น boilerplate ออก และลบองค์ประกอบที่ซ้ำกันในทุกหน้าออก มิฉะนั้นคุณไม่ได้กำลังสร้างฐานความรู้ (knowledge base) แต่คุณกำลังสร้าง search engine สำหรับส่วนประกอบตกแต่งเว็บไซต์ (website chrome)
3. การแบ่ง Chunk ทำลายความหมาย
การแบ่ง chunk แบบขนาดคงที่ (Fixed-size chunking) เป็นค่าเริ่มต้นในคู่มือ quickstart เกือบทุกฉบับ และมันอันตรายมาก หากคุณแบ่งเอกสารโดยใช้เพียงจำนวนตัวอักษร คุณจะตัดตารางขาดครึ่ง, แยกขั้นตอนที่ 4 และ 5 ในกระบวนการที่มีลำดับหมายเลขออกจากกัน, และทำให้ bullet points หลุดออกจากหัวข้อของมัน ชิ้นส่วนข้อมูลที่มีเพียงครึ่งหลังของตารางราคาจะไม่มีประโยชน์ในเชิงความหมายเลย การแบ่ง chunk แบบตระหนักถึงโครงสร้าง (structure-aware chunking) จะเคารพรูปแบบต้นฉบับ เช่น การ parse ลำดับชั้นของหัวข้อ, พยายามรักษาตารางไว้ให้ครบถ้วน, แบ่งที่ขอบเขตของย่อหน้าภายใต้ H2 หรือ H3 เดียวกัน และรักษา list ให้อยู่ใน chunk เดียวกันหากมันสั้นพอ เป้าหมายไม่ใช่การสร้างบล็อกที่มีขนาดเท่ากัน แต่คือการสร้างหน่วยความหมายที่สอดคล้องกัน (coherent units of meaning)
4. ปัญหาเรื่องความสดใหม่ของข้อมูล
A static snapshot of an internal wiki is simple mode. Continuously ingesting from the live web is hard. You need to know when a page was last gathered, whether it has changed since then, and how long the information remains valid. Stale data does not always mean a visibly old date. Sometimes a page updates its text but keeps the same URL, so your system never notices without content hashing. Build clear refresh rules based on source volatility. A financial data feed might need hourly checks. A company about page might need quarterly checks. Record timestamps and set time-to-live boundaries, especially if your domain involves regulated or safety-critical guidance where old facts can cause real harm.
5. Duplicate Pollution
Websites are full of repetition. The same product description appears on the category page, the product page, and a promotional landing page. The same press release lives under /news/, /press/, and /blog/. Vector search does not deduplicate automatically. If ten nearly identical chunks sit in your database, they can crowd out diverse, relevant results in your top-k retrieval. You need canonical tracking or content deduplication before embedding. If two chunks say the same thing, keep the authoritative source and drop the copies. Your retriever has limited slots. Do not let them go to waste.
6. Missing Metadata
A vector database without metadata is just a dense text search engine with no memory of context. Smart retrieval depends on filtering and ranking signals that raw embeddings cannot provide. Store the source URL, the capture date, the document category, and the version number. If you ingest API documentation, versioning is essential. Without it, a query might blend v1 and v2 specs into the same answer. If you ingest HR policies, tagging by region or department lets you filter results before they ever reach the model. Metadata turns a text dump into a curated knowledge system.
7. JavaScript Gaps
Modern sites do not ship their content in the first HTML payload. They send a skeleton and hydrate it with JavaScript calls. A basic HTTP request might see nothing but a loading spinner and a layout shell. If your pipeline cannot execute JavaScript, you will ingest blank pages or partial fragments and never realize anything is wrong. Using a headless browser solves the rendering problem but introduces new ones: heavier memory use, slower throughput, and bot detection walls. Choose your trade-offs deliberately, but do not pretend a simple curl equivalent is enough for every source.
A Practical Ingestion Checklist
If you are building or reviewing a RAG feed, start here:
- Validate source coverage and pagination. A sitemap might list only the first ten articles in a category. Crawl deep and verify that paginated or dynamically loaded content is actually captured.
- Remove boilerplate before chunking. Strip navigation, ads, footers, and repeated legal disclaimers. If a phrase appears on every page, it is noise.
- Use structure-aware chunking. Respect headings, bullet lists, and tables. Split on semantic boundaries, not character counts.
- Attach rich metadata. Include URL, capture date, content category, and version. Make these fields filterable in your retrieval queries.
- Set refresh frequencies based on data volatility. High-change sources need frequent re-crawls. Static archives do not.
- Monitor the corpus, not just the job status. A pipeline can exit with code zero while producing garbage. Audit samples of stored chunks regularly for drift and quality.
- Define rules for versioning and deletions. When a source page is removed, delete its chunks. When it updates, overwrite or version them. Orphaned data is a silent killer.
The Hard Truth About Embeddings
No embedding model, no matter how advanced, can repair a missing document. It cannot guess that a page was updated last week if your feed still holds last year's copy. It cannot infer the context of a table row that got separated from its header by a bad chunk boundary. Embeddings compress meaning, but they do not create meaning where the ingestion layer failed to preserve it.
Retrieval quality starts at the ingestion layer. That layer decides whether your RAG system is a useful tool or just a confident liar with a vector database behind it.
The Real Takeaway
เลิกวัดความสมบูรณ์ของการนำเข้าข้อมูลด้วยแดชบอร์ดของไพป์ไลน์เพียงอย่างเดียว งานที่ขึ้นสถานะสีเขียวและล็อกที่สะอาดไม่ได้การันตีว่าคอร์ปัสของคุณจะสะอาด ลองเปิดฐานข้อมูลแล้วอ่านชิ้นส่วนข้อมูล (chunks) จริงๆ ที่ผู้ใช้จะดึงไปใช้ หากข้อความเต็มไปด้วยประกาศลิขสิทธิ์ ตารางที่ถูกตัดแบ่ง และหน้าประกาศนโยบายที่ล้าสมัย ปัญหาของคุณไม่ใช่ที่ LLM แก้ไขที่แหล่งข้อมูลก่อน เพราะนอกเหนือจากนั้นก็เป็นเพียงการปรับจูนบนกองขยะเท่านั้น
