Graceful shutdown: the exit sequence decides between 502s and clean responses
Exiting immediately on a termination signal cuts in-flight requests. The order is stop accepting, drain what is running, release resources, then exit.
Sporadic 502s during a rolling deploy almost always mean the shutdown order is wrong. The process exits the moment it gets a termination signal, in-flight requests are severed, and the load balancer records the failure.
The correct four steps
- Stop accepting new connections: deregister from health checks or close the listening socket
- Wait for in-flight requests: with a ceiling
- Release resources: database connections, queues, temp files
- Exit
Steps 1 and 2 must be separate, and that is the crux. Deregister first, then drain, with an observation window between them, because the load balancer needs one or two health check cycles to actually stop forwarding.
A structure that works
let shuttingDown = false;
const inflight = new Set<Promise<unknown>>();
process.on('SIGTERM', async () => {
shuttingDown = true;
server.close(); // 1. stop accepting
await sleep(3000); // let the LB deregister us
await Promise.race([ // 2. drain, with a deadline
Promise.allSettled([...inflight]),
sleep(25000),
]);
await db.end(); // 3. release
process.exit(0); // 4. exit
});
// health check reflects the state
app.get('/healthz', (req, res) =>
res.status(shuttingDown ? 503 : 200).send('ok'));
Three order mistakes
| Mistake | Result |
|---|---|
process.exit() before closing |
in-flight requests severed, 502s |
| exiting without deregistering | the LB still routes to you, same 502s |
| waiting without a timeout | one stuck request blocks deploy forever |
The third is the nastiest: the drain must have a ceiling. Past it, exit anyway and let the load balancer turn failures into retries.
One extra step in containers
The container must forward the signal to the real process. If the entrypoint is a shell script without exec, the signal goes to the shell, the app never sees it, and it is SIGKILLed at the timeout.
CMD ["node", "server.js"] # exec form, no shell wrapper
On Kubernetes, terminationGracePeriodSeconds must be larger than your drain timeout, or SIGKILL arrives before draining finishes.
How to verify
Do not trust logs. Hammer the service during a deploy and count non-2xx:
while true; do curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/; done
A rolling restart should produce 200 only. Any 502 means one of the four steps is missing.
Shutdown is a procedure, not a line of exit. Stop accepting, drain, release, exit. Skip one and you leave 502s behind on every deploy.

Comments
…