GitLab Workhorse
How we reduced 502 errors by caring about PID 1 in Kubernetes

How we reduced 502 errors by caring about PID 1 in Kubernetes

5/17/2022 · Steve Azzopardi

What this post added

This post details the debugging and resolution of intermittent 502 errors impacting the GitLab Pages service, which were traced back to issues with GitLab Workhorse's graceful termination in Kubernetes. The investigation revealed that Workhorse was not cleanly shutting down upon receiving a SIGTERM signal, leading to prolonged periods where it continued to serve 502 errors while the Puma/webservice container was already terminating. This was exacerbated by Kubernetes' default 30-second termination grace period. The solution involved identifying that the Workhorse process was not correctly handling the SIGTERM signal, and ensuring it did so to allow for a faster exit, thus reducing the window for 502 errors. This improved the reliability of the GitLab Pages service during pod lifecycle events.

Read the original post ↗