
5/17/2022 · Steve Azzopardi
What this post added
This post details the debugging and resolution of intermittent 502 errors impacting the GitLab Pages service, which were traced back to issues with GitLab Workhorse's graceful termination in Kubernetes. The investigation revealed that Workhorse was not cleanly shutting down upon receiving a SIGTERM signal, leading to prolonged periods where it continued to serve 502 errors while the Puma/webservice container was already terminating. This was exacerbated by Kubernetes' default 30-second termination grace period. The solution involved identifying that the Workhorse process was not correctly handling the SIGTERM signal, and ensuring it did so to allow for a faster exit, thus reducing the window for 502 errors. This improved the reliability of the GitLab Pages service during pod lifecycle events.