ਸ਼ੁਰੂ ਵਿੱਚ ਲੋਕਲ ਲਾਰਜ ਲੈਂਗੂਏਜ ਮਾਡਲ ਬਹੁਤ ਤੇਜ਼ ਮਹਿਸੂਸ ਹੁੰਦੇ ਹਨ। ਤੁਸੀਂ ਇੱਕ 7B ਜਾਂ 13B ਪੈਰਾਮੀਟਰ ਵਾਲਾ ਮਾਡਲ ਲੋਡ ਕਰਦੇ ਹੋ, ਇੱਕ ਛੋਟਾ ਪ੍ਰੋਂਪਟ ਦਿੰਦੇ ਹੋ, ਅਤੇ ਟੋਕਨ ਸਕ੍ਰੀਨ 'ਤੇ ਆਰਾਮਦਾਇਕ ਰਫ਼ਤਾਰ ਨਾਲ ਚੱਲਦੇ ਹਨ। ਫਿਰ ਤੁਸੀਂ ਕੋਡ ਦਾ ਇੱਕ ਲੰਬਾ ਬਲਾਕ ਪੇਸਟ ਕਰਦੇ ਹੋ, ਜਾਂ ਤੁਹਾਡੀ ਚੈਟ ਹਿਸਟਰੀ ਕਈ ਗੇੜਾਂ (turns) ਤੱਕ ਵਧ ਜਾਂਦੀ ਹੈ, ਅਤੇ ਮਾਡਲ ਦੀ ਰਫ਼ਤਾਰ ਬਹੁਤ ਘੱਟ ਹੋਣ ਲੱਗਦੀ ਹੈ। ਇਹ ਸੁਸਤੀ ਕਦੇ ਵੀ ਹੌਲੀ-ਹੌਲੀ ਨਹੀਂ ਆਉਂਦੀ। ਇਹ ਇੱਕ ਅਚਾਨਕ ਖੱਡ ਵਾਂਗ ਹੁੰਦੀ ਹੈ। ਇੱਕ ਪਲ GPU ਟੋਕਨ ਬਣਾ ਰਿਹਾ ਹੁੰਦਾ ਹੈ; ਅਗਲੇ ਹੀ ਪਲ, ਤੁਹਾਡਾ ਸਿਸਟਮ ਮੋਨੀਟਰ ਮੈਮੋਰੀ ਦਾ ਦਬਾਅ ਵਧਦਾ ਦਿਖਾਉਂਦਾ ਹੈ ਅਤੇ ਜਨਰੇਸ਼ਨ (generation) ਅਟਕ-ਅਟਕ ਕੇ ਚੱਲਣ ਲੱਗਦੀ ਹੈ। ਤੁਸੀਂ ਕਿਸੇ ਸਹੀ ਫਾਰਮੂਲੇ ਨਾਲ ਇਹ ਪੂਰੀ ਤਰ੍ਹਾਂ ਨਹੀਂ ਦੱਸ ਸਕਦੇ ਕਿ ਇਹ ਕਦੋਂ ਹੋਵੇਗਾ। ਤੁਹਾਡਾ ਇੱਕੋ ਇੱਕ ਭਰੋਸੇਯੋਗ ਮਾਰਗਦਰਸ਼ਕ ਹਾਰਡਵੇਅਰ ਖੁਦ ਹੈ।

ਕੰਟੈਕਸਟ (Context) ਦੀ ਲੁਕੀ ਹੋਈ ਕੀਮਤ

ਤੁਹਾਡੇ ਦੁਆਰਾ ਜਨਰੇਟ ਕੀਤਾ ਗਿਆ ਹਰ ਟੋਕਨ KV cache ਵਿੱਚ ਸਟੇਟ (state) ਜੋੜਦਾ ਹੈ। ਇਹ ਕੈਸ਼ prefill ਅਤੇ generation ਪੜਾਵਾਂ ਦੌਰਾਨ ਗਣਨਾ ਕੀਤੇ ਗਏ keys ਅਤੇ values ਨੂੰ ਸਟੋਰ ਕਰਦਾ ਹੈ, ਅਤੇ ਇਹ ਤੁਹਾਡੇ ਮਾਡਲ ਵੇਟਸ (weights), attention buffers, ਅਤੇ runtime overhead ਦੇ ਨਾਲ ਮੈਮੋਰੀ ਵਿੱਚ ਰਹਿੰਦਾ ਹੈ। 12 GB ਜਾਂ 16 GB VRAM ਵਾਲੇ ਇੱਕ ਆਮ ਕੰਜ਼ਿਊਮਰ GPU 'ਤੇ, KV cache ਅੰਤ ਵਿੱਚ ਬਾਕੀ ਸਾਰੀਆਂ ਚੀਜ਼ਾਂ ਨਾਲ ਜਗ੍ਹਾ ਲਈ ਮੁਕਾਬਲਾ ਕਰਦਾ ਹੈ। ਜਦੋਂ ਡੈਡੀਕੇਟਡ ਵੀਡੀਓ ਮੈਮੋਰੀ ਭਰ ਜਾਂਦੀ ਹੈ, ਤਾਂ ਓਪਰੇਟਿੰਗ ਸਿਸਟਮ ਕੋਈ ਐਰਰ (error) ਨਹੀਂ ਦਿੰਦਾ ਅਤੇ ਰੁਕਦਾ ਨਹੀਂ ਹੈ। ਇਹ ਚੁੱਪਚਾਪ ਵਾਧੂ ਡੇਟਾ ਨੂੰ shared memory ਵਿੱਚ ਭੇਜ ਦਿੰਦਾ ਹੈ, ਅਤੇ PCIe bus ਰਾਹੀਂ GPU ਅਤੇ system RAM ਦੇ ਵਿਚਕਾਰ ਡੇਟਾ ਨੂੰ ਆਵਾਜਾਈ ਕਰਦਾ ਹੈ। ਉਹ ਬੱਸ ਫਾਈਲ ਟ੍ਰਾਂਸਫਰ ਲਈ ਤੇਜ਼ ਹੈ, ਪਰ ਗ੍ਰਾਫਿਕਸ ਕਾਰਡ ਦੇ ਅੰਦਰਲੀ ਮੈਮੋਰੀ ਬੈਂਡਵਿਡਥ (memory bandwidth) ਦੇ ਮੁਕਾਬਲੇ ਇਹ ਬਹੁਤ ਹੀ ਹੌਲੀ ਹੈ। ਨਤੀਜਾ ਪ੍ਰਦਰਸ਼ਨ (performance) ਵਿੱਚ ਮਾਮੂਲੀ ਗਿਰਾਵਟ ਨਹੀਂ ਹੈ। ਇਹ ਇੱਕ ਪੂਰੀ ਤਰ੍ਹਾਂ ਦੀ ਬੈਖ਼ੀ (collapse) ਹੈ।

ਤਿੰਨ ਸੰਕੇਤ ਜੋ ਦੱਸਦੇ ਹਨ ਕਿ ਖੱਡ ਆ ਗਈ ਹੈ

ਜਦੋਂ ਮਾਡਲ ਚੱਲ ਰਿਹਾ ਹੋਵੇ ਤਾਂ ਆਪਣੇ ਹਾਰਡਵੇਅਰ ਮੋਨੀਟਰਾਂ 'ਤੇ ਨਜ਼ਰ ਰੱਖੋ। ਜਦੋਂ ਪ੍ਰਦਰਸ਼ਨ ਵਿੱਚ ਅਚਾਨਕ ਗਿਰਾਵਟ ਆਉਂਦੀ ਹੈ, ਤਾਂ ਤੁਹਾਨੂੰ ਤਿੰਨ ਸਪਸ਼ਟ ਸੰਕੇਤ ਦਿਖਾਈ ਦੇਣਗੇ।

  • Shared VRAM ਵਧਦਾ ਹੈ। ਇਹ ਉਹ ਮੈਮੋਰੀ ਹੈ ਜਿਸ ਨੂੰ GPU driver ਨੇ dedicated video RAM ਤੋਂ ਕੱਢ ਕੇ host operating system ਦੁਆਰਾ ਪ੍ਰਬੰਧਿਤ ਪੂਲ (pool) ਵਿੱਚ ਪਾ ਦਿੱਤਾ ਹੈ। ਜਿਸ ਪਲ ਇਹ ਮੈਟ੍ਰਿਕ ਜ਼ੀਰੋ ਤੋਂ ਉੱਪਰ ਜਾਂਦਾ ਹੈ, ਸਮਝੋ ਤੁਸੀਂ ਲਕੀਰ ਪਾਰ ਕਰ ਲਈ ਹੈ।
  • System RAM ਦੀ ਵਰਤੋਂ ਵਧਦੀ ਹੈ। ਵਾਧੂ ਡੇਟਾ ਨੂੰ ਕਿਤੇ ਨਾ ਕਿਤੇ ਜਾਣਾ ਹੀ ਪੈਂਦਾ ਹੈ, ਅਤੇ ਉਹ ਜਗ੍ਹਾ ਤੁਹਾਡੀ ਮੁੱਖ ਮੈਮੋਰੀ ਹੈ। ਜੇਕਰ ਮਾਡਲ ਟੋਕਨ ਜਨਰੇਟ ਕਰਦੇ ਸਮੇਂ ਤੁਹਾਡੀ RAM ਦੀ ਵਰਤੋਂ ਵਧਦੀ ਹੈ, ਤਾਂ ਇਸਦਾ ਮਤਲਬ ਹੈ ਕਿ ਡੇਟਾ GPU ਤੋਂ ਹਟਾਇਆ ਜਾ ਰਿਹਾ ਹੈ।
  • Eval speed ਅੱਧੀ ਜਾਂ ਇਸ ਤੋਂ ਵੱਧ ਘਟ ਜਾਂਦੀ ਹੈ। 10% ਦੀ ਸੁਸਤੀ ਦਾ ਮਤਲਬ thermal throttling ਜਾਂ ਬੈਕਗ੍ਰਾਊਂਡ ਪ੍ਰੋਸੈਸ ਹੋ ਸਕਦਾ ਹੈ। 50% ਦੀ ਗਿਰਾਵਟ, ਜਾਂ ਇਸ ਤੋਂ ਵੀ ਮਾੜਾ, ਇਸਦਾ ਮਤਲਬ ਹੈ ਕਿ ਰੁਕਾਵਟ (bottleneck) tensor cores ਤੋਂ ਹਟ ਕੇ memory bandwidth ਅਤੇ PCIe latency ਵੱਲ ਚਲੀ ਗਈ ਹੈ। ਜਦੋਂ ਤੁਸੀਂ ਦੇਖਦੇ ਹੋ ਕਿ generation ਦੋ ਅੰਕਾਂ (double digits) ਤੋਂ ਇੱਕ ਅੰਕ (single digits) ਤੱਕ ਡਿੱਗ ਗਈ ਹੈ, ਤਾਂ ਤੁਸੀਂ ਪਹਿਲਾਂ ਹੀ ਖੱਡ ਵਿੱਚ ਡਿੱਗ ਚੁੱਕੇ ਹੋ।

ਤੁਹਾਡਾ ਤੇਜ਼ ਬੈਂਚਮਾਰਕ (Benchmark) ਸ਼ਾਇਦ ਝੂਠ ਬੋਲ ਰਿਹਾ ਹੈ

ਇੱਕ ਛੋਟਾ ਜਿਹਾ ਟੈਸਟ ਤੁਹਾਨੂੰ ਗਲਤ ਭਰੋਸਾ ਦੇ ਸਕਦਾ ਹੈ। ਜੇਕਰ ਤੁਸੀਂ ਸੌ-ਟੋਕਨ ਵਾਲੇ ਪ੍ਰੋਂਪਟ ਨਾਲ ਮਾਡਲ ਦਾ ਬੈਂਚਮਾਰਕ ਕਰਦੇ ਹੋ, ਚੰਗੀ ਰਫ਼ਤਾਰ ਦੇਖਦੇ ਹੋ, ਅਤੇ ਸਮਝ ਲੈਂਦੇ ਹੋ ਕਿ ਸਭ ਠੀਕ ਹੈ, ਤਾਂ ਤੁਸੀਂ ਸਿਰਫ਼ ਸ਼ੁਰੂਆਤੀ ਸੁਖਦ ਪੜਾਅ (honeymoon phase) ਨੂੰ ਮਾਪਿਆ ਹੈ। KV cache ਲਗਭਗ ਖਾਲੀ ਹੁੰਦਾ ਹੈ। ਲੰਬੇ prefill ਕਾਰਨ ਲੇਅਰਾਂ 'ਤੇ ਕੋਈ ਦਬਾਅ ਨਹੀਂ ਪੈਂਦਾ। ਅਸਲ ਪ੍ਰਭਾਵ ਉਦੋਂ ਹੀ ਸਾਹਮਣੇ ਆਉਂਦਾ ਹੈ ਜਦੋਂ ਮਾਡਲ ਇੱਕ ਵੱਡੇ ਪ੍ਰੋਂਪਟ ਨੂੰ ਪ੍ਰੋਸੈਸ ਕਰ ਲੈਂਦਾ ਹੈ ਅਤੇ ਕੈਸ਼ ਆਪਣੇ ਅਸਲ ਕੰਮ ਕਰਨ ਵਾਲੇ ਆਕਾਰ ਤੱਕ ਭਰ ਜਾਂਦਾ ਹੈ। ਤੁਹਾਨੂੰ ਡੂੰਘੇ prefill ਅਤੇ ਲੰਬੇ generation ਰਨ ਨਾਲ ਟੈਸਟ ਕਰਨਾ ਚਾਹੀਦਾ ਹੈ। ਕੰਟੈਕਸਟ ਨੂੰ ਅਸਲ ਵਿੱਚ ਇਕੱਠਾ ਹੋਣ ਦਿਓ। ਉਦੋਂ ਹੀ ਮੈਮੋਰੀ ਦਾ ਦਬਾਅ ਸਥਿਰ ਹੋਵੇਗਾ ਅਤੇ ਤੁਹਾਨੂੰ ਅਸਲ ਸੀਮਾ ਦਿਖਾਏਗਾ।

llama.cpp ਨਾਲ ਆਪਣੀ ਸੀਮਾ ਲੱਭਣਾ

ਜੇਕਰ ਤੁਸੀਂ llama.cpp ਰਾਹੀਂ ਮਾਡਲ ਚਲਾ ਰਹੇ ਹੋ, ਤਾਂ ਤੁਸੀਂ ਸਧਾਰਨ ਗਣਿਤ ਅਤੇ ਇੱਕ ਧੀਰਜ ਵਾਲੇ ਟੈਸਟ ਰਨ ਨਾਲ ਆਪਣੀ ਸੀਮਾ ਮਾਪ ਸਕਦੇ ਹੋ।

1. Shared memory ਦੀ ਵਰਤੋਂ ਮਾਪੋ।
ਇੱਕ ਨਿਗਰਨੀ ਪ੍ਰੋਂਪਟ ਨਾਲ ਆਪਣੀ ਬੇਸਲਾਈਨ (baseline) dedicated VRAM ਨੂੰ ਰਿਕਾਰਡ ਕਰੋ, ਫਿਰ ਇੱਕ ਲੰਬੇ ਕੰਟੈਕਸਟ ਵਾਲਾ ਕੰਮ ਚਲਾਓ ਅਤੇ ਸਭ ਤੋਂ ਵੱਧ (peak) ਮਾਪ ਨੂੰ ਨੋਟ ਕਰੋ। Peak ਵਿੱਚੋਂ baseline ਨੂੰ ਘਟਾਓ। ਇਹ ਅੰਤਰ ਉਹ ਹੈ ਜੋ ਤੁਹਾਡੇ GPU ਤੋਂ shared system memory ਵਿੱਚ ਜਾ ਗਿਆ ਹੈ।

2. ਆਪਣਾ RAM delta ਗਣਨਾ ਕਰੋ।
System RAM ਲਈ ਵੀ ਇਹੀ ਘਟਾਓ ਕਰੋ। ਲੰਬੇ ਰਨ ਦੌਰਾਨ peak RAM ਵਿੱਚੋਂ ਆਪਣੀ baseline RAM ਨੂੰ ਘਟਾਓ। ਇਹ ਸੰਖਿਆ ਤੁਹਾਨੂੰ ਬਿਲਕੁਲ ਦੱਸੇਗੀ ਕਿ ਕਿੰਨਾ ਡੇਟਾ ਵੀਡੀਓ ਕਾਰਡ ਤੋਂ ਤੁਹਾਡੀ ਮੁੱਖ ਮੈਮੋਰੀ ਵਿੱਚ ਭੇਜਿਆ ਗਿਆ ਹੈ। ਇਹ ਬੱਸ ਰਾਹੀਂ ਹੋ ਰਹੀ ਲੀਕੇਜ ਨੂੰ ਮਾਪਦਾ ਹੈ।

3. Eval speed ਦੀ ਗਿਰਾਵਟ ਦਾ ਸਮਾਂ ਦੇਖੋ।
ਆਪਣੇ ਬੇਸਲਾਈਨ tokens-per-second ਦੀ ਰਫ਼ਤਾਰ ਦੀ ਤੁਲਨਾ ਉਸ ਰਫ਼ਤਾਰ ਨਾਲ ਕਰੋ ਜੋ ਇੱਕ ਲੰਬੇ ਦਸਤਾਵੇਜ਼ ਨੂੰ ਪ੍ਰੋਸੈਸ ਕਰਨ ਤੋਂ ਬਾਅਦ ਮਿਲਦੀ ਹੈ। ਤੁਸੀਂ ਦੇਖ ਸਕਦੇ ਹੋ ਕਿ ਜਦੋਂ ਕੰਟੈਕਸਟ ਨਵਾਂ ਹੁੰਦਾ ਹੈ ਤਾਂ ਮਾਡਲ ਸਤਾਰਾਂ (17) ਟੋਕਨ ਪ੍ਰਤੀ ਸੈਕਿੰਡ ਦੀ ਰਫ਼ਤਾਰ ਨਾਲ ਚੱਲਦਾ ਹੈ, ਪਰ ਕੈਸ਼ ਭਰ ਜਾਣ ਤੋਂ ਬਾਅਦ ਸਿਰਫ਼ ਦੋ ਟੋਕਨ ਪ੍ਰਤੀ ਸੈਕਿੰਡ ਹੀ ਦਿੰਦਾ ਹੈ। ਉਹ ਪੰਦਰਾਂ (15) ਟੋਕਨਾਂ ਦੀ ਗਿਰਾਵਟ ਤੁਹਾਡੇ ਲਈ ਇੱਕ ਚੇਤਾਵਨੀ ਹੈ।

ਬ੍ਰੇਕਿੰਗ ਪੁਆਇੰਟ (Breaking Point) ਦਾ ਪਤਾ ਲਗਾਉਣਾ

To map the curve accurately, do not settle for one lonely data point. Run three distinct trials at 16,000 tokens, 32,000 tokens, and 65,000 tokens. Two points might suggest a line, but two dots are just a guess. The third point proves whether you are looking at measurement noise or a real memory wall. Subtract the results between runs to calculate how much extra memory each additional thousand tokens consumes on your specific combination of model, quantization layer, and GPU.

Once you have that slope, you can project forward. Take your per-token cost, multiply it by the target context length, divide by 1024 to move between units, and add the result to your base model VRAM load. The equation looks like this:

Model VRAM load + (tokens × memory per token ÷ 1024) = Theoretical VRAM usage

This projection is not prophecy. It is a guidepost derived from actual behavior. Use it to estimate your ceiling before you commit to a full production run.

Why Paper Formulas Fail, and What Quantization Can Fix

Textbook formulas ignore the messy reality of local inference. Different architectures allocate attention buffers differently. Your operating system reserves VRAM for the display driver, compositor, and CUDA context. Driver versions change how aggressively they use shared memory. A theoretical equation cannot know how much VRAM is actually free on your machine at 2:00 PM with a browser full of tabs open. You have to run the model on your specific hardware and watch the meters.

Quantization offers partial relief. Moving the KV cache from f16 to q8_0 halves its memory footprint while keeping precision high enough for nearly all practical tasks. That change buys you headroom. It does not grant immunity. The cache still grows linearly with every token you feed in. Eventually, even the reduced size overwhelms your available dedicated memory and the spillover to system RAM begins. The pressure only stops when the context window is capped or the data stops moving.

The Real Takeaway

Do not trust marketing slides, parameter counts, or back-of-the-envelope math. Load the model. Open your system monitor. Run a 65,000-token thread, watch the RAM climb, and count the tokens per second. The numbers that appear on your specific screen, on your specific GPU, are the only numbers that matter. Context always wins. Your job is to know exactly when it wins on your machine.