ਜ਼ਿਆਦਾਤਰ Node.js ਟਿਊਟੋਰਿਅਲ ਐਰਰ ਹੈਂਡਲਿੰਗ (error handling) ਨੂੰ ਇੱਕ ਸਹਾਇਕ ਚੀਜ਼ ਵਜੋਂ ਮੰਨਦੇ ਹਨ। ਤੁਸੀਂ ਇੱਕ ਰੂਟ ਹੈਂਡਲਰ ਨੂੰ try/catch ਬਲਾਕ ਵਿੱਚ ਲਪੇਟਦੇ ਹੋ, ਸਟੈਕ ਟ੍ਰੇਸ (stack trace) ਨੂੰ ਲੌਗ ਕਰਦੇ ਹੋ, ਅਤੇ 500 ਰਿਟਰਨ ਕਰਦੇ ਹੋ। ਇਹ ਸੋਚ ਇਸ ਲਈ ਟਿਕੀ ਹੋਈ ਹੈ ਕਿਉਂਕਿ HTTP ਰਿਕੁਐਸਟ ਦੇ ਦੂਜੇ ਪਾਸੇ ਇੱਕ ਅਸਲੀ ਵਿਅਕਤੀ ਉਡੀਕ ਰਿਹਾ ਹੁੰਦਾ ਹੈ। ਬੈਕਗ੍ਰਾਊਂਡ ਜੌਬਸ (Background jobs) ਵੱਖਰੀਆਂ ਹੁੰਦੀਆਂ ਹਨ। ਇੱਕ ਕਿਊ ਸਿਸਟਮ (queue system) ਵਿੱਚ, ਜਵਾਬ ਦੇਣ ਲਈ ਕੋਈ ਬੇਸਬਰ ਕਲਾਇੰਟ ਨਹੀਂ ਹੁੰਦਾ, ਅਤੇ ਨਾ ਹੀ ਕੋਈ ਆਟੋਮੈਟਿਕ ਬ੍ਰਾਊਜ਼ਰ ਰਿਫ੍ਰੈਸ਼ ਹੁੰਦਾ ਹੈ। ਉੱਥੇ ਸਿਰਫ਼ ਇੱਕ ਵਰਕਰ (worker), ਇੱਕ ਪੇਲੋਡ (payload), ਅਤੇ ਚੁੱਪਚਾਪ ਵਧ ਰਿਹਾ ਇੱਕ ਰੀਟ੍ਰਾਈ ਕਾਊਂਟਰ (retry counter) ਹੁੰਦਾ ਹੈ। ਜਦੋਂ ਚੀਜ਼ਾਂ ਗਲਤ ਹੁੰਦੀਆਂ ਹਨ, ਤਾਂ ਉਹ ਪਹਿਲਾਂ ਹੌਲੀ-ਹੌਲੀ ਗਲਤ ਹੁੰਦੀਆਂ ਹਨ, ਫਿਰ ਇੱਕੋ ਵਾਰ ਸਭ ਕੁਝ ਵਿਗੜ ਜਾਂਦਾ ਹੈ। ਇੱਕ ਗਲਤ ਤਰੀਕੇ ਨਾਲ ਸ਼੍ਰੇਣੀਬੱਧ ਕੀਤੀ ਗਈ ਐਰਰ ਪੂਰੀ ਪਾਈਪਲਾਈਨ ਨੂੰ ਰੋਕ ਸਕਦੀ ਹੈ ਜਾਂ ਰਾਤ ਦੇ ਤਿੰਨ ਵਜੇ ਕਿਸੇ ਇੰਜੀਨੀਅਰ ਨੂੰ ਜਗਾ ਸਕਦੀ ਹੈ।
ਇਹ ਅੰਤਰ ਸਧਾਰਨ ਹੈ। ਰਿਕੁਐਸਟ-ਰਿਸਪਾਂਸ ਚੱਕਰ (Request-response cycles) ਤੇਜ਼ੀ ਨਾਲ ਅਤੇ ਸਾਫ਼ ਤੌਰ 'ਤੇ ਫੇਲ ਹੁੰਦੇ ਹਨ। ਕਿਊ ਦੀ ਫੇਲ੍ਹ ਹੋਣ ਦੀ ਪ੍ਰਕਿਰਿਆ ਸ਼ਾਂਤ ਹੁੰਦੀ ਹੈ। ਡਾਟਾਬੇਸ ਕਨੈਕਸ਼ਨ ਟੁੱਟਣ ਤੋਂ ਪਹਿਲਾਂ ਇੱਕ ਵਰਕਰ ਸੈਂਕੜੇ ਜੌਬਸ ਨੂੰ ਪ੍ਰੋਸੈਸ ਕਰ ਸਕਦਾ ਹੈ। ਸਪੱਸ਼ਟ ਹੈਂਡਲਿੰਗ ਨਿਯਮਾਂ ਤੋਂ ਬਿਨਾਂ, ਵਰਕਰ ਤੁਰੰਤ ਰੀਟ੍ਰਾਈ ਕਰਦਾ ਹੈ, ਪਹਿਲਾਂ ਤੋਂ ਹੀ ਸੰਘਰਸ਼ ਕਰ ਰਹੇ ਡਾਟਾਬੇਸ 'ਤੇ ਦਬਾਅ ਪਾਉਂਦਾ ਹੈ, ਅਤੇ ਕ੍ਰੈਸ਼ ਹੋ ਜਾਂਦਾ ਹੈ। ਕਿਉਂਕਿ ਕੋਈ ਵੀ ਵਰਕਰ ਨੂੰ ਸਿੱਧੇ ਤੌਰ 'ਤੇ ਨਹੀਂ ਦੇਖ ਰਿਹਾ ਹੁੰਦਾ, ਮੁਸੀਬਤ ਦਾ ਪਹਿਲਾ ਸੰਕੇਤ ਅਕਸਰ ਕੈਸਕੇਡਿੰਗ ਬੈਕਅੱਪ (cascading backup) ਜਾਂ ਲੌਗ ਫਾਈਲਾਂ ਨਾਲ ਭਰੀ ਹੋਈ ਡਿਸਕ ਹੁੰਦੀ ਹੈ। ਤੁਹਾਨੂੰ ਸਿਰਫ਼ catch ਬਲਾਕਾਂ ਤੋਂ ਵੱਧ ਦੀ ਲੋੜ ਹੈ। ਤੁਹਾਨੂੰ ਇੱਕ ਅਜਿਹੀ ਰਣਨੀਤੀ ਦੀ ਲੋੜ ਹੈ ਜੋ ਵੱਖ-ਵੱਖ ਫੇਲ੍ਹ ਹੋਣ ਵਾਲੀਆਂ ਘਟਨਾਵਾਂ ਨਾਲ ਵੱਖਰੇ ਤਰੀਕੇ ਨਾਲ ਨਿਬਟੇ ਅਤੇ ਇੱਕ ਖ਼ਰਾਬ ਜੌਬ ਤੋਂ ਪੂਰੇ ਸਿਸਟਮ ਦੀ ਰੱਖਿਆ ਕਰੇ।
ਫੇਲ੍ਹ ਹੋਣ ਦੀਆਂ ਦੋ ਕਿਸਮਾਂ
ਹਰ ਐਰਰ ਨੂੰ ਦੋ ਸ਼੍ਰੇਣੀਆਂ ਵਿੱਚ ਵੰਡ ਕੇ ਸ਼ੁਰੂਆਤ ਕਰੋ।
ਰੀਟ੍ਰਾਈਏਬਲ (Retryable) ਐਰਰ ਅਸਥਾਈ ਹੁੰਦੇ ਹਨ। ਕਿਸੇ Third-party API ਦਾ ਨੈੱਟਵਰਕ ਟਾਈਮਆਊਟ, 429 ਰੇਟ-ਲਿਮਿਟ ਰਿਸਪਾਂਸ, ਜਾਂ ਕਿਸੇ ਟੈਂਪਰੇਰੀ ਡਾਟਾਬੇਸ ਰੈਪਲੀਕਾ ਦਾ ਆਪਣੇ ਪ੍ਰਾਇਮਰੀ ਤੋਂ ਪਿੱਛੇ ਰਹਿ ਜਾਣਾ। ਇਹ ਦਬਾਅ ਦੇ ਲੱਛਣ ਹਨ, ਬੱਗ (bugs) ਨਹੀਂ। ਸਿਸਟਮ ਤਿਰਤੀ ਸਕਿੰਟਾਂ ਵਿੱਚ ਆਪਣੇ ਆਪ ਠੀਕ ਹੋ ਸਕਦਾ ਹੈ। ਰੀਟ੍ਰਾਈਏਬਲ ਜੌਬਸ ਨੂੰ ਦੁਬਾਰਾ ਕੋਸ਼ਿਸ਼ ਕਰਨ ਦਾ ਮੌਕਾ ਮਿਲਣਾ ਚਾਹੀਦਾ ਹੈ, ਪਰ ਸਿਰਫ਼ ਨਿਯੰਤਰਿਤ ਸ਼ਰਤਾਂ ਦੇ ਅਧੀਨ।
ਪਰਮਾਨੈਂਟ (Permanent) ਐਰਰ ਗਲਤੀਆਂ ਹਨ। ਪੇਲੋਡ ਵਿੱਚ ਗਲਤ JSON, ਗੁੰਮ ਹੋਈ ਯੂਜ਼ਰ ID, ਜਾਂ ਸਟੋਰੇਜ ਵਿੱਚ ਮੌਜੂਦ ਨਾ ਹੋਣ ਵਾਲੀ ਕੋਈ ਜ਼ਰੂਰੀ ਫਾਈਲ। ਇਹ ਸੌਵੀਂ ਕੋਸ਼ਿਸ਼ 'ਤੇ ਵੀ ਉਵੇਂ ਹੀ ਫੇਲ ਹੋਣਗੇ ਜਿਵੇਂ ਉਹ ਪਹਿਲੀ ਕੋਸ਼ਿਸ਼ 'ਤੇ ਹੋਏ ਸਨ। ਉਹਨਾਂ ਨੂੰ ਦੁਬਾਰਾ ਕੋਸ਼ਿਸ਼ ਕਰਨ ਨਾਲ CPU ਸਾਈਕਲ ਖ਼ਰਾਬ ਹੁੰਦੇ ਹਨ, ਕਿਊ ਸਲਾਟਾਂ ਦੀ ਬਰਬਾਦੀ ਹੁੰਦੀ ਹੈ, ਅਤੇ ਇੱਕ ਅਜਿਹਾ ਟੌਕਸਿਕ ਬੈਕ-ਪ੍ਰੈਸ਼ਰ (toxic back-pressure) ਪੈਦਾ ਹੁੰਦਾ ਹੈ ਜੋ ਸਹੀ ਜੌਬਸ ਨੂੰ ਦੇਰੀ ਕਰਦਾ ਹੈ। ਪਰਮਾਨੈਂਟ ਫੇਲ੍ਹ ਹੋਣ ਲਈ ਇੱਕੋ ਇੱਕ ਲਾਭਦਾਇਕ ਜਗ੍ਹਾ ਲੌਗ, ਇੱਕ ਅਲਰਟ, ਜਾਂ ਡੈੱਡ-ਲੈਟਰ ਕਿਊ (dead-letter queue) ਹੈ। ਇਹ ਰੀਟ੍ਰਾਈ ਲੂਪ (retry loop) ਦਾ ਹਿੱਸਾ ਨਹੀਂ ਹੋਣਾ ਚਾਹੀਦਾ।
ਇੱਕ ਡੈਸੀਜ਼ਨ ਇੰਜਣ (Decision Engine) ਬਣਾਓ
ਤੁਰੰਤ ਸ਼੍ਰੇਣੀਬੱਧ ਕਰੋ। ਇਸ ਫੈਸਲੇ ਨੂੰ ਕਿਊ ਫਰੇਮਵਰਕ (queue framework) 'ਤੇ ਨਾ ਛੱਡੋ। ਜਿਸ ਪਲ ਤੁਸੀਂ ਐਰਰ ਨੂੰ ਫੜਦੇ ਹੋ, ਉਸੇ ਪਲ ਉਸਦੀ ਕਿਸਮਤ ਦਾ ਫੈਸਲਾ ਕਰੋ।
ਅਭਿਆਸ ਵਿੱਚ, ਇਸਦਾ ਮਤਲਬ ਹੈ ਕਸਟਮ ਐਰਰ ਕਲਾਸਾਂ (custom error classes) ਜਾਂ ਵੈਪਰ ਫੰਕਸ਼ਨਾਂ (wrapper functions) ਨੂੰ ਬਣਾਉਣਾ ਜੋ ਐਰਰ ਨੂੰ ਅੱਗੇ ਭੇਜਣ ਤੋਂ ਪਹਿਲਾਂ ਉਸਦੀ ਜਾਂਚ ਕਰਦੇ ਹਨ। ਜੇਕਰ ਕੋਈ ਡਾਟਾਬੇਸ ਡਰਾਈਵਰ 'connection reset' ਦਿੰਦਾ ਹੈ, ਤਾਂ ਤੁਹਾਡੇ ਹੈਂਡਲਰ ਨੂੰ ਇਸਨੂੰ ਰੀਟ੍ਰਾਈਏਬਲ ਵਜੋਂ ਟੈਗ ਕਰਨਾ ਚਾਹੀਦਾ ਹੈ। ਜੇਕਰ ਕੋਈ ਪੇਲੋਡ ਵੈਲੀਡੇਟਰ 'schema mismatch' ਦਿੰਦਾ ਹੈ, ਤਾਂ ਇਸਨੂੰ ਪਰਮਾਨੈਂਟ ਵਜੋਂ ਟੈਗ ਕਰੋ। ਬਹੁਤ ਸਾਰੇ ਜੌਬ ਪ੍ਰੋਸੈਸਰ ਡਿਫੌਲਟ ਰੂਪ ਵਿੱਚ ਸਭ ਕੁਝ ਰੀਟ੍ਰਾਈ ਕਰਦੇ ਹਨ, ਜੋ ਕਿ ਸਭ ਤੋਂ ਮਹਿੰਗਾ ਫੈਸਲਾ ਹੋ ਸਕਦਾ ਹੈ। ਪਰਮਾਨੈਂਟ ਜੌਬਸ ਨੂੰ ਤੁਰੰਤ ਰੱਦ ਕਰ ਦਿਓ। ਜਾਂ ਤਾਂ ਉਹਨਾਂ ਨੂੰ ਡ੍ਰੌਪ ਕਰ ਦਿਓ ਜਾਂ ਉਹਨਾਂ ਨੂੰ ਡੈੱਡ-ਲੈਟਰ ਕਿਊ (dead-letter queue) ਵੱਲ ਰੀਰੂਟ ਕਰ ਦਿਓ ਜਿੱਥੇ ਉਹ ਮੁੱਖ ਪਾਈਪਲਾਈਨ ਨੂੰ ਖ਼ਰਾਬ ਨਾ ਕਰ ਸਕਣ। ਇਹ ਇੱਕ ਆਦਤ ਕਿਸੇ ਵੀ ਇਨਫਰਾਸਟ੍ਰਕਚਰ
Design every job as if it will run twice, because it might. A worker can fail halfway through processing, get rescheduled, and execute again. If your job charges a customer, sends an email, or increments an inventory count, a naive retry creates duplicates.
The fix is idempotency. Before you perform a side effect, check whether it already happened. Use a unique identifier from the job payload as an idempotency key. Store that key in a short-lived cache or a database table with a uniqueness constraint. If the key exists, skip the work and return success. This turns retries from a risk into a harmless no-op. It takes a few extra lines of code, but it saves you from explaining to finance why revenue doubled overnight.
Protect the Process
Unbounded promise rejections and stray exceptions can kill a Node.js process without warning. In a worker, that means dropped jobs and an orchestrator scrambling to restart the container.
Register global handlers for unhandledRejection and uncaughtException. Their job is not to rescue the application. It is to perform the minimum cleanup necessary, then exit. Let Docker, Kubernetes, or systemd restart the worker with a clean memory state. Limping along after a global handler fires invites memory leaks and corrupted state. A fast, clean death is safer than a slow zombie process that processes jobs incorrectly. Trust your orchestrator to bring you back; do not try to outsmart a corrupted runtime.
Respect the Signal
Workers get shut down during deployments, scaling events, and node rotations. If your process dies the instant it receives SIGTERM, you abort whatever job is currently in flight. That job may never finish, and its retry counter might not even have incremented yet.
Listen for SIGTERM and SIGINT. When a signal arrives, stop pulling new jobs from the queue. Finish the current job if you can. Set a hard timeout, perhaps thirty seconds, after which you exit regardless. This graceful shutdown respects the queue and avoids false failures. Your deployment pipeline should treat a worker that exits cleanly as healthy, while a crashed worker should trigger an alert.
The Real Takeaway
Reliable queue handling is not about catching every error. It is about making deliberate decisions for each failure mode. Retry the transient ones with patience. Bury the permanent ones quickly. Shield your workers from stampedes, protect your data with idempotency keys, and let dying processes exit cleanly. When every failure has a defined path, three in the morning becomes just another hour. Your pipeline keeps moving, and your team keeps sleeping.
