AI ਗੱਲਬਾਤ ਵਿੱਚ ਗੁੰਮ ਹੋਈ ਕੜੀ

ਹਰ ਕੋਈ AI agents ਬਾਰੇ ਗੱਲ ਕਰ ਰਿਹਾ ਹੈ। ਕਿਸੇ ਵੀ ਟੈਕ ਫੀਡ ਨੂੰ ਦੇਖੋ, ਤੁਹਾਨੂੰ ਅਜਿਹੇ ਦਰਜਨਾਂ ਡੈਮੋ ਮਿਲ ਜਾਣਗੇ ਜੋ ਇੱਕ ਵੱਡੇ ਲੈਂਗੂਏਜ ਮਾਡਲ (LLM) ਨੂੰ ਇੱਕ ਹੀ ਸ਼ਾਨਦਾਰ ਗੱਲਬਾਤ ਵਿੱਚ ਫਲਾਈਟ ਬੁੱਕ ਕਰਦੇ, ਕੋਡ ਲਿਖਦੇ, ਜਾਂ ਸਪੋਰਟ ਟਿਕਟਾਂ ਦਾ ਜਵਾਬ ਦਿੰਦੇ ਦਿਖਾਉਂਦੇ ਹਨ। ਇਸ ਦਾ ਅਸਲ ਸੰਦੇਸ਼ ਸਪੱਸ਼ਟ ਜਾਪਦਾ ਹੈ: ਜੇਕਰ ਤੁਸੀਂ ਇੱਕ ਯੂਜ਼ਰ ਨੂੰ LLM ਨਾਲ ਜੋੜ ਦਿੰਦੇ ਹੋ, ਤਾਂ ਜਾਦੂ ਹੋ ਜਾਂਦਾ ਹੈ।

ਉਹ ਭਰਮ ਪੰਜ ਮਿੰਟ ਦੇ ਡੈਮੋ ਲਈ ਬਹੁਤ ਵਧੀਆ ਕੰਮ ਕਰਦਾ ਹੈ। ਪਰ ਜਿਵੇਂ ਹੀ ਅਸਲੀ ਯੂਜ਼ਰ, ਅਸਲੀ ਡੇਟਾ ਅਤੇ ਅਸਲੀ ਪੈਸਾ ਸਾਹਮਣੇ ਆਉਂਦਾ ਹੈ, ਇਹ ਭਰਮ ਟੁੱਟ ਜਾਂਦਾ ਹੈ। ਪ੍ਰੋਡਕਸ਼ਨ (production) ਵਿੱਚ, ਰਿਸ਼ਤਾ ਕਦੇ ਵੀ ਸਿਰਫ਼ User ↔ LLM ਨਹੀਂ ਹੁੰਦਾ। ਇਹ User ↔ ਇੱਕ ਗੁੰਝਲਦਾਰ ਸਿਸਟਮ ਹੁੰਦਾ ਹੈ ਜਿਸ ਵਿੱਚ ਇੱਕ LLM ਸ਼ਾਮਲ ਹੁੰਦਾ ਹੈ। ਉਸ ਸਿਸਟਮ ਦਾ ਉਹ ਹਿੱਸਾ ਜਿਸ ਬਾਰੇ ਕੋਈ ਗੱਲ ਨਹੀਂ ਕਰਦਾ, ਉਹ ਹੈ harness—ਉਹ ਢਾਂਚਾ (scaffolding) ਜੋ ਮਾਡਲ ਦੇ ਆਲੇ-ਦੁਆਲੇ ਸਭ ਕੁਝ ਚੁਣਦਾ ਹੈ, ਰੂਟ ਕਰਦਾ ਹੈ, ਸੁਰੱਖਿਅਤ ਕਰਦਾ ਹੈ ਅਤੇ ਸੰਚਾਲਿਤ (orchestrate) ਕਰਦਾ ਹੈ। ਇਸ ਤੋਂ ਬਿਨਾਂ, ਤੁਹਾਡੇ ਕੋਲ ਕੋਈ ਪ੍ਰੋਡਕਟ ਨਹੀਂ ਹੈ। ਤੁਹਾਡੇ ਕੋਲ ਸਿਰਫ਼ ਇੱਕ ਪ੍ਰੋਟੋਟਾਈਪ (prototype) ਹੈ।

ਸਧਾਰਨ ਲੂਪ ਕਿਉਂ ਟੁੱਟ ਜਾਂਦਾ ਹੈ

ਇੱਕ ਡੈਮੋ ਇੱਕ ਨਿਯੰਤਰਿਤ ਮਾਤਾਵਲ (controlled environment) ਹੁੰਦਾ ਹੈ। ਸਵਾਲ ਛੋਟੇ ਹੁੰਦੇ ਹਨ, ਸੰਦਰਭ (context) ਸੀਮਤ ਹੁੰਦਾ ਹੈ, ਅਤੇ ਜੋਖਮ ਘੱਟ ਹੁੰਦੇ ਹਨ। ਡਿਵੈਲਪਰ ਇੱਕ ਸਿੰਗਲ API ਕਾਲ ਕਰਦਾ ਹੈ, ਇੱਕ ਸੁਚਾਰੂ ਜਵਾਬ ਪ੍ਰਾਪਤ ਕਰਦਾ ਹੈ, ਅਤੇ ਦਰਸ਼ਕ ਤਾੜੀਆਂ ਵਜਾਉਂਦੇ ਹਨ। ਪਰ ਪ੍ਰੋਡਕਸ਼ਨ ਬਹੁਤ ਉਲਝਣ ਭਰਿਆ ਹੁੰਦਾ ਹੈ। ਯੂਜ਼ਰ ਅਸਪਸ਼ਟ ਫਾਲੋ-ਅੱਪ ਸਵਾਲ ਪੁੱਛਦੇ ਹਨ। Third-party APIs ਟਾਈਮ-ਆਊਟ ਹੋ ਜਾਂਦੇ ਹਨ। ਇੱਕ ਮਾਡਲ ਜੋ ਕੱਲ੍ਹ ਤੱਕ ਬਿਲਕੁਲ ਸਹੀ JSON ਬਣਾ ਰਿਹਾ ਸੀ, ਅਚਾਨਕ ਮਾਰਕਡਾਊਨ (markdown) ਦੇਣ ਲੱਗ ਜਾਂਦਾ ਹੈ। Context windows ਭਰ ਜਾਂਦੀਆਂ ਹਨ। ਸਭ ਤੋਂ ਮਾੜੇ ਸਮੇਂ 'ਤੇ rate limits ਲਾਗੂ ਹੋ ਜਾਂਦੀਆਂ ਹਨ।

ਇੱਕ ਸਧਾਰਨ prompt-response ਲੂਪ ਕੋਲ ਇਹਨਾਂ ਵਿੱਚੋਂ ਕਿਸੇ ਵੀ ਚੀਜ਼ ਦਾ ਕੋਈ ਜਵਾਬ ਨਹੀਂ ਹੁੰਦਾ। ਇਸ ਨੂੰ ਨਹੀਂ ਪਤਾ ਹੁੰਦਾ ਕਿ ਕਿਸੇ ਖਾਸ ਕੰਮ ਲਈ ਮਾਡਲ ਦਾ ਕਿਹੜਾ ਵੇਰੀਐਂਟ (variant) ਵਰਤਣਾ ਚਾਹੀਦਾ ਹੈ। ਇਸ ਨੂੰ ਯਾਦ ਨਹੀਂ ਰਹਿੰਦਾ ਕਿ ਤਿੰਨ ਪਲ ਪਹਿਲਾਂ ਕੀ ਹੋਇਆ ਸੀ। ਇਹ ਅਸਫਲ ਕਾਲ ਨੂੰ ਦੁਬਾਰਾ ਕੋਸ਼ਿਸ਼ (retry) ਨਹੀਂ ਕਰ ਸਕਦਾ, ਖਰਚੇ ਵਧਣ 'ਤੇ ਰਿਕਵੈਸਟਾਂ ਨੂੰ ਘਟਾ (throttle) ਨਹੀਂ ਸਕਦਾ, ਜਾਂ ਤੁਹਾਡੇ ਡੇਟਾਬੇਸ ਵਿੱਚ ਜਾਣ ਤੋਂ ਪਹਿਲਾਂ ਆਉਟਪੁੱਟ ਨੂੰ ਸਾਫ਼ (sanitize) ਨਹੀਂ ਕਰ ਸਕਦਾ। ਇਹ ਕੋਈ ਅਸਧਾਰਨ ਮਾਮਲੇ (edge cases) ਨਹੀਂ ਹਨ। ਇਹ ਅਸਲ ਦੁਨੀਆ ਦੇ ਸਾਫਟਵੇਅਰ ਦੀਆਂ ਮੁੱਖ ਵਿਸ਼ੇਸ਼ਤਾਵਾਂ ਹਨ। ਇਹਨਾਂ ਨੂੰ ਸੰਭਾਲਣਾ ਹੀ harness ਦਾ ਕੰਮ ਹੈ।

Harness ਅਸਲ ਵਿੱਚ ਕੀ ਕਰਦਾ ਹੈ

harness ਨੂੰ ਇੱਕ ਇੰਜੀਨੀਅਰਿੰਗ ਲੇਅਰ ਵਜੋਂ ਸਮਝੋ ਜੋ ਇੱਕ ਲੈਂਗੂਏਜ ਮਾਡਲ ਨੂੰ ਇੱਕ ਚਲਾਕ ਟੈਕਸਟ ਜਨਰੇਟਰ ਤੋਂ ਇੱਕ ਭਰੋਸੇਯੋਗ ਸਰਵਿਸ ਕੰਪੋਨੈਂਟ ਵਿੱਚ ਬਦਲ ਦਿੰਦੀ ਹੈ। ਇਸ ਦੀਆਂ ਜ਼ਿੰਮੇਵਾਰੀਆਂ ਸਪੱਸ਼ਟ ਅਤੇ ਆਮ ਹਨ, ਜਿਸ ਕਾਰਨ ਅਕਸਰ ਇਹਨਾਂ ਵੱਲ ਧਿਆਨ ਨਹੀਂ ਦਿੱਤਾ ਜਾਂਦਾ।

ਮੌਜੂਦਾ ਕੰਮ ਲਈ ਮਾਡਲ ਦੀ ਚੋਣ। ਹਰ ਇੰਟਰੈਕਸ਼ਨ ਲਈ ਉਪਲਬਧ ਸਭ ਤੋਂ ਸ਼ਕਤੀਸ਼ਾਲੀ ਫਾਊਂਡੇਸ਼ਨ ਮਾਡਲ ਦੀ ਲੋੜ ਨਹੀਂ ਹੁੰਦੀ। ਕੁਝ ਕੰਮਾਂ ਲਈ ਡੂੰਘੀ ਤਰਕ ਸ਼ਕਤੀ (reasoning power) ਦੀ ਲੋੜ ਹੁੰਦੀ ਹੈ; ਦੂਜਿਆਂ ਨੂੰ ਸਿਰਫ਼ ਤੇਜ਼ੀ ਅਤੇ ਘੱਟ ਲਾਗਤ ਦੀ ਲੋੜ ਹੁੰਦੀ ਹੈ। ਇੱਕ ਵਧੀਆ ਤਰੀਕੇ ਨਾਲ ਬਣਾਇਆ ਗਿਆ harness ਰਿਕਵੈਸਟਾਂ ਨੂੰ ਸਮਝਦਾਰੀ ਨਾਲ ਰੂਟ ਕਰਦਾ ਹੈ। ਉਦਾਹਰਨ ਲਈ, ਇੱਕ ਕਸਟਮਰ ਸਪੋਰਟ ਏਜੰਟ ਆਉਣ ਵਾਲੇ ਸੁਨੇਹੇ ਦੇ ਇਰਾਦੇ (intent) ਨੂੰ ਸ਼੍ਰੇਣੀਬੱਧ ਕਰਨ ਲਈ ਇੱਕ ਤੇਜ਼, ਸਸਤੇ ਮਾਡਲ ਦੀ ਵਰਤੋਂ ਕਰ ਸਕਦਾ ਹੈ—ਜਿਵੇਂ ਕਿ ਰਿਫੰਡ ਦੀ ਬੇਨਤੀ ਬਨਾਮ ਸ਼ਿਪਿੰਗ ਸਵਾਲ। ਜੇਕਰ ਇਰਾਦਾ ਕਿਸੇ ਗੁੰਝਲਦਾਰ ਨੀਤੀ ਦੇ ਵਿਵਾਦ ਦਾ ਸੰਕੇਤ ਦਿੰਦਾ ਹੈ, ਤਾਂ harness ਇਸ ਕੰਮ ਨੂੰ ਇੱਕ ਭਾਰੀ reasoning ਮਾਡਲ ਨੂੰ ਸੌਂਪ ਦਿੰਦਾ ਹੈ। ਜੇਕਰ ਯੂਜ਼ਰ ਨੂੰ ਸਿਰਫ਼ ਟ੍ਰੈਕਿੰਗ ਲਿੰਕ ਚਾਹੀਦਾ ਹੈ, ਤਾਂ ਹਲਕਾ ਮਾਡਲ ਤੁਰੰਤ ਜਵਾਬ ਦਿੰਦਾ ਹੈ ਅਤੇ ਤੁਹਾਡਾ ਖਰਚਾ (burn rate) ਕੰਟਰੋਲ ਵਿੱਚ ਰਹਿੰਦਾ ਹੈ।

ਡੇਟਾ ਫਲੋ (data flow) ਨੂੰ ਸੰਭਾਲਣਾ। ਅਸਲੀ ਐਪਲੀਕੇਸ਼ਨਾਂ ਖਾਲੀਪਣ ਵਿੱਚ ਨਹੀਂ ਰਹਿੰਦੀਆਂ। ਇੱਕ AI ਏਜੰਟ ਨੂੰ ਅਕਸਰ ਇੱਕ ਵੈਕਟਰ ਸਟੋਰ (vector store) ਤੋਂ ਦਸਤਾਵੇਜ਼ ਕੱਢਣ, CRM ਨੂੰ ਕੁਐਰੀ ਕਰਨ, ਹਾਲੀਆ ਯੂਜ਼ਰ ਗਤੀਵਿਧੀ ਪੜ੍ਹਨ, ਅਤੇ ਫਿਰ ਉਹਨਾਂ ਸਭ ਨੂੰ ਇੱਕ ਸੁਮੇਲ ਜਵਾਬ ਵਿੱਚ ਬਦਲਣ ਦੀ ਲੋੜ ਹੁੰਦੀ ਹੈ। Harness ਉਸ ਪ੍ਰਕਿਰਿਆ ਦਾ ਪ੍ਰਬੰਧਨ ਕਰਦਾ ਹੈ। ਇਹ ਸਹੀ context chunks ਲੈਂਦਾ ਹੈ, ਇਹ ਜਾਂਚ ਕਰਦਾ ਹੈ ਕਿ ਉਹ ਪ੍ਰਸੰਗਿਕਤਾ ਗੁਆਏ ਬਿਨਾਂ ਟੋਕਨ ਸੀਮਾਵਾਂ ਦੇ ਅੰਦਰ ਹਨ ਜਾਂ ਨਹੀਂ, ਉਹਨਾਂ ਨੂੰ ਮਾਡਲ ਲਈ ਤਿਆਰ ਕਰਦਾ ਹੈ, ਅਤੇ ਨਤੀਜੇ ਨੂੰ ਚੇਨ ਵਿੱਚ ਅਗਲੇ ਸਿਸਟਮ ਨੂੰ ਭੇਜ ਦਿੰਦਾ ਹੈ। ਇਸ ਸੰਚਾਲਨ (orchestration) ਤੋਂ ਬਿਨਾਂ, ਮਾਡਲ ਕੋਲ ਜਾਂ ਤਾਂ ਸੰਦਰਭ ਦੀ ਕਮੀ ਹੁੰਦੀ ਹੈ ਜਾਂ ਉਹ ਬਹੁਤ ਜ਼ਿਆਦਾ ਫਾਲਤੂ ਜਾਣਕਾਰੀ ਵਿੱਚ ਡੁੱਬ ਜਾਂਦਾ ਹੈ।

ਗਲਤੀਆਂ (errors) ਦਾ ਪ੍ਰਬੰਧਨ ਕਰਨਾ। LLMs ਉਹਨਾਂ ਤਰੀਕਿਆਂ ਨਾਲ ਅਸਫਲ ਹੁੰਦੇ ਹਨ ਜਿਵੇਂ ਰਵਾਇਤੀ ਸੇਵਾਵਾਂ ਨਹੀਂ ਹੁੰਦੀਆਂ। ਉਹ ਗਲਤ ਜਾਂ ਕਲਪਨਾਤਮਕ (hallucinate) ਸਟ੍ਰਕਚਰਡ ਆਉਟਪੁੱਟ ਦਿੰਦੇ ਹਨ। ਉਹ ਖਾਲੀ ਜਵਾਬ ਵਾਪਸ ਕਰਦੇ ਹਨ। ਜਿਵੇਂ ਹੀ ਅਸਲ ਮਾਡਲ ਵਰਜ਼ਨ ਵਿੱਚ ਥੋੜ੍ਹਾ ਜਿਹਾ ਬਦਲਾਅ ਆਉਂਦਾ ਹੈ, ਉਹ ਫਾਰਮੈਟਿੰਗ ਨਿਰਦੇਸ਼ਾਂ ਦੀ ਉਲੰਘਣਾ ਕਰਦੇ ਹਨ। Harness ਇਹਨਾਂ ਅਸਫਲਤਾਵਾਂ ਨੂੰ ਹੈਰਾਨੀ ਦੀ ਬਜਾਏ ਇੱਕ ਉਮੀਦ ਕੀਤੀ ਗਈ ਪ੍ਰਕਿਰਿਆ ਵਜੋਂ ਲੈਂਦਾ ਹੈ। ਇਹ schemas ਦੀ ਜਾਂਚ ਕਰਦਾ ਹੈ, ਗਲਤ ਆਉਟਪੁੱਟ ਨੂੰ ਫੜਦਾ ਹੈ, exponential backoff ਦੇ ਨਾਲ retry logic ਲਾਗੂ ਕਰਦਾ ਹੈ, ਅਤੇ ਜਦੋਂ ਮੁੱਖ endpoint ਅਸਫਲ ਹੁੰਦਾ ਹੈ ਤਾਂ ਕਿਸੇ ਦੂਜੇ ਪ੍ਰੋਵਾਈਡਰ ਜਾਂ cached ਨਤੀਜੇ ਦੀ ਵਰਤੋਂ ਕਰਦਾ ਹੈ। ਜਦੋਂ ਕੁਝ ਵੀ ਕੰਮ ਨਹੀਂ ਕਰਦਾ, ਤਾਂ ਇਹ ਕਿਸੇ ਭੁਗਤਾਨ ਕਰਦੇ ਗਾਹਕ ਨੂੰ ਬੇਤੁਕੀ ਜਾਣਕਾਰੀ ਦੇਣ ਦੀ ਬਜਾਏ ਇੱਕ ਮਨੁੱਖੀ ਓਪਰੇਟਰ ਕੋਲ ਮਾਮਲਾ ਭੇਜ ਦਿੰਦਾ ਹੈ।

ਸਿਸਟਮ ਦੀ ਭਰੋਸੇਯੋਗਤਾ ਯਕੀਨੀ ਬਣਾਉਣਾ। ਪ੍ਰੋ

This explains a phenomenon that confuses many product teams. Two companies can start with the exact same foundation model—same weights, same context window, same training cutoff—and ship experiences that feel worlds apart. One feels brittle, slow, and weirdly forgetful. The other feels snappy, consistent, and trustworthy.

The difference is never the model itself. It is the system wrapped around it. One team treated the model as the entire product. The other treated it as one component inside a disciplined architecture. The harness is where that discipline lives.

The Shift from Prompts to Architecture

Early AI development put prompt engineering front and center. Tweaking wording, adding examples, and layering in role-play instructions could dramatically improve output quality. That skill still matters, but it has hit diminishing returns as a competitive moat. You cannot prompt your way out of a missing retry policy or a tangled data pipeline that leaks private context into a public-facing response.

The real shift happening right now is a move toward software architecture. Engineers are designing state machines, defining strict interfaces between the model layer and application logic, and treating non-determinism as a first-class engineering concern. They are asking distributed systems questions: How does state persist across a multi-turn conversation? What happens when a downstream tool is unavailable? How do we test a system whose core component is probabilistic? These are the questions that separate a toy from a tool.

Building for Production: Observability and Control

If you are serious about shipping, the harness demands two qualities above all: observability and orchestration.

Observability means you can see what the model received, what it returned, and how long each step took. It means tracing an agent’s decision loop across fourteen tool calls and spotting exactly where it started looping or drifting off mission. Without that visibility, debugging an AI system is like fixing a car engine in the dark.

Orchestration means your business logic stays separate from your model interaction layer. It means versioning prompts the way you version code, so a new deployment does not silently change behavior. It means deliberately testing failure modes—killing an API mid-request, feeding malformed tool results, simulating a context window overflow—to see if the harness keeps the system upright. Frameworks come and go, and whether you adopt an off-the-shelf orchestration library or build your own, the discipline matters more than the brand name.

The Real Takeaway

Foundation models will keep improving. They will get faster, cheaper, and more capable. But a more powerful engine does not fix a broken chassis. The teams that win over the next few years will not be the ones with the fanciest model access. They will be the ones who built a harness that is reliable, observable, and well-orchestrated. They will swap models without rewriting their applications. They will control costs because the harness governs every token. They will sleep through the night because their systems fail gracefully.

Stop obsessing over the model in isolation. Start obsessing over the system that runs it. The future belongs to engineers who build smarter systems around smart models.


This article draws on ideas originally discussed by Abdulaziz Zos in "Beyond The Model".

For more discussions on AI engineering and system design, check out the GyaanSetu learning community.