โมเดลภาษาขนาดใหญ่ (LLM) ที่รันในเครื่องจะให้ความรู้สึกว่าเร็วปานสายฟ้าแลบในช่วงแรก คุณโหลดโมเดลขนาด 7B หรือ 13B พารามิเตอร์ขึ้นมา ส่ง prompt สั้นๆ เข้าไป แล้วโทเค็นก็จะไหลผ่านหน้าจอด้วยความเร็วที่น่าพอใจ แต่เมื่อคุณวางโค้ดชุดยาวๆ หรือประวัติการแชทเริ่มสะสมจนยาวเหยียด โมเดลก็จะเริ่มอืดลง ความล่าช้าที่เกิดขึ้นมักจะไม่ใช่การค่อยๆ ช้าลงอย่างนุ่มนวล แต่มันคือการตกหน้าผา ชั่วขณะหนึ่ง GPU ยังคงประมวลผลโทเค็นได้อย่างรวดเร็ว แต่ในวินาทีถัดมา ระบบ monitor ของคุณจะแสดงให้เห็นว่าความกดดันของหน่วยความจำ (memory pressure) กำลังเพิ่มสูงขึ้น และการสร้างข้อความก็เริ่มติดขัด คุณไม่สามารถทำนายได้แน่ชัดว่าจะเกิดสิ่งนี้ขึ้นเมื่อไหร่ด้วยสูตรคำนวณที่ตายตัว สิ่งเดียวที่เป็นเครื่องนำทางที่เชื่อถือได้คือตัวฮาร์ดแวร์เอง
ต้นทุนที่ซ่อนอยู่ของ Context
ทุกๆ โทเค็นที่คุณสร้างขึ้นจะเพิ่มสถานะ (state) เข้าไปใน KV cache ซึ่ง cache นี้จะเก็บค่า keys และ values ที่ถูกคำนวณระหว่างขั้นตอน prefill และ generation โดยมันจะอาศัยอยู่ในหน่วยความจำร่วมกับ model weights, attention buffers และ runtime overhead อื่นๆ ใน GPU สำหรับผู้บริโภคทั่วไปที่มี VRAM ขนาด 12 GB หรือ 16 GB ในที่สุด KV cache จะต้องแย่งพื้นที่กับส่วนประกอบอื่นๆ เมื่อหน่วยความจำวิดีโอ (dedicated video memory) เต็ม ระบบปฏิบัติการจะไม่แจ้งข้อผิดพลาดแล้วหยุดทำงาน แต่มันจะแอบส่งข้อมูลส่วนที่ล้นออกไปยัง shared memory แทน โดยการส่งข้อมูลสลับไปมาระหว่าง GPU และ system RAM ผ่าน PCIe bus ซึ่งบัสนี้อาจจะเร็วสำหรับการถ่ายโอนไฟล์ แต่เมื่อเทียบกับแบนด์วิดท์หน่วยความจำ (memory bandwidth) ภายในตัวการ์ดจอแล้ว มันช้าเหมือนเต่า ผลลัพธ์ที่ได้จึงไม่ใช่แค่ประสิทธิภาพที่ลดลงเล็กน้อย แต่มันคือการพังทลายของระบบ
3 สัญญาณเตือนว่าคุณกำลังตกหน้าผา
ให้สังเกตฮาร์ดแวร์ monitor ในขณะที่โมเดลกำลังทำงาน คุณจะเห็นสัญญาณที่ชัดเจน 3 อย่างเมื่อประสิทธิภาพเริ่มตกหน้าผา
- Shared VRAM เพิ่มสูงขึ้น: นี่คือหน่วยความจำที่ GPU driver ได้ผลักออกจาก dedicated video RAM ไปยัง pool ที่จัดการโดยระบบปฏิบัติการ ทันทีที่ตัวเลขนี้สูงกว่าศูนย์ แสดงว่าคุณได้ข้ามเส้นแบ่งนั้นมาแล้ว
- การใช้งาน System RAM พุ่งสูงขึ้น: ข้อมูลที่ล้นออกมาต้องมีที่ลง และปลายทางนั้นก็คือหน่วยความจำหลักของคุณ หากการใช้งาน RAM เพิ่มขึ้นในขณะที่โมเดลกำลังสร้างโทเค็น แสดงว่าข้อมูลกำลังถูกถ่ายโอนออกจาก GPU
- ความเร็ว Eval speed ลดลงครึ่งหนึ่งหรือมากกว่านั้น: หากความเร็วลดลง 10% อาจหมายถึงปัญหาความร้อน (thermal throttling) หรือมีโปรเซสอื่นทำงานอยู่เบื้องหลัง แต่ถ้าความเร็วลดลง 50% หรือมากกว่านั้น หมายความว่าคอขวดได้เปลี่ยนจาก tensor cores ไปอยู่ที่ memory bandwidth และ PCIe latency แล้ว เมื่อคุณเห็นความเร็วในการสร้างโทเค็นตกลงจากเลขสองหลักมาเหลือเลขหลักเดียว แสดงว่าคุณได้ตกหน้าผาเรียบร้อยแล้ว
ทำไมการทำ Benchmark แบบเร็วๆ ถึงอาจหลอกคุณได้
การทดสอบแบบเร็วๆ (smoke test) จะให้ความมั่นใจที่ผิดพลาดแก่คุณ หากคุณทำ benchmark โมเดลด้วย prompt เพียงร้อยโทเค็น เห็น throughput ที่ดี แล้วสรุปผลเลย นั่นหมายความว่าคุณเพิ่งวัดผลแค่ในช่วง "ฮันนีมูน" เท่านั้น เพราะในตอนนั้น KV cache ยังเกือบว่างเปล่า เลเยอร์ต่างๆ ยังไม่ถูกใช้งานหนักจากการ prefill ที่ยาวนาน รอยเท้า (footprint) ที่แท้จริงจะปรากฏออกมาก็ต่อเมื่อโมเดลได้ประมวลผล prompt ขนาดใหญ่และ cache ถูกเติมจนเต็มขนาดการทำงานจริงแล้ว คุณต้องทดสอบด้วยการทำ deep prefill และการรัน generation ที่ยาวนาน ปล่อยให้ context สะสมไปเรื่อยๆ เมื่อนั้นความกดดันของหน่วยความจำจึงจะคงที่และแสดงให้เห็นขีดจำกัดที่แท้จริงของคุณ
การหาขีดจำกัดของคุณด้วย llama.cpp
หากคุณรันโมเดลผ่าน llama.cpp คุณสามารถวัด "กำแพง" ของคุณได้ด้วยการคำนวณง่ายๆ และการทดสอบอย่างอดทน
1. วัดการใช้งาน shared memory
บันทึกค่า baseline ของ dedicated VRAM ด้วย prompt ที่สั้นที่สุด จากนั้นรันงานที่มี long-context และจดค่าสูงสุด (peak) ไว้ นำค่า peak ลบด้วย baseline ส่วนต่างที่ได้คือปริมาณข้อมูลที่ล้นออกจาก GPU ไปยัง shared system memory
2. คำนวณค่า RAM delta
ใช้วิธีการลบแบบเดียวกันกับ system RAM โดยนำค่า RAM baseline ลบด้วยค่า RAM สูงสุดระหว่างการรันงานยาวๆ ตัวเลขนี้จะบอกคุณอย่างแม่นยำว่ามีข้อมูลถูกผลักจากวิดีโอการ์ดไปยังหน่วยความจำหลักมากน้อยเพียงใด ซึ่งเป็นการวัดปริมาณการรั่วไหลผ่านบัส
3. จับเวลาความเร็ว eval speed ที่ตกลง
เปรียบเทียบอัตรา tokens-per-second ในช่วง baseline กับอัตราหลังจากที่โมเดลอ่านเอกสารยาวๆ จบ คุณอาจเห็นโมเดลวิ่งด้วยความเร็ว 17 tokens per second เมื่อ context ยังใหม่ แต่กลับเหลือเพียง 2 tokens per second เมื่อ cache บวมขึ้น การลดลงถึง 15 tokens นั้นคือสัญญาณเตือนภัยล่วงหน้า (canary in the coal mine) ของคุณ
การระบุจุดแตกหักอย่างแม่นยำ
To map the curve accurately, do not settle for one lonely data point. Run three distinct trials at 16,000 tokens, 32,000 tokens, and 65,000 tokens. Two points might suggest a line, but two dots are just a guess. The third point proves whether you are looking at measurement noise or a real memory wall. Subtract the results between runs to calculate how much extra memory each additional thousand tokens consumes on your specific combination of model, quantization layer, and GPU.
Once you have that slope, you can project forward. Take your per-token cost, multiply it by the target context length, divide by 1024 to move between units, and add the result to your base model VRAM load. The equation looks like this:
Model VRAM load + (tokens × memory per token ÷ 1024) = Theoretical VRAM usage
This projection is not prophecy. It is a guidepost derived from actual behavior. Use it to estimate your ceiling before you commit to a full production run.
Why Paper Formulas Fail, and What Quantization Can Fix
Textbook formulas ignore the messy reality of local inference. Different architectures allocate attention buffers differently. Your operating system reserves VRAM for the display driver, compositor, and CUDA context. Driver versions change how aggressively they use shared memory. A theoretical equation cannot know how much VRAM is actually free on your machine at 2:00 PM with a browser full of tabs open. You have to run the model on your specific hardware and watch the meters.
Quantization offers partial relief. Moving the KV cache from f16 to q8_0 halves its memory footprint while keeping precision high enough for nearly all practical tasks. That change buys you headroom. It does not grant immunity. The cache still grows linearly with every token you feed in. Eventually, even the reduced size overwhelms your available dedicated memory and the spillover to system RAM begins. The pressure only stops when the context window is capped or the data stops moving.
The Real Takeaway
Do not trust marketing slides, parameter counts, or back-of-the-envelope math. Load the model. Open your system monitor. Run a 65,000-token thread, watch the RAM climb, and count the tokens per second. The numbers that appear on your specific screen, on your specific GPU, are the only numbers that matter. Context always wins. Your job is to know exactly when it wins on your machine.
