Work

/

About
Play
Blog
Home/Blog/DevOps
DevOps
•
Updated Oct 7, 2026
•
5 min read

Debugging OOMKilled Pods: Exit Code 137 on a Small Cluster

A full glass container of green blocks with an extra gold block above its capacity
MS
Muhammad Saad

Shopify engineer. I build storefronts, two published Shopify apps, and the infrastructure behind a 70,000+ product store, and I write here about what that work teaches me.

Follow on LinkedInSee my workBook a call
Found this useful? Share it.

Keep reading

Ivory clockwork driving green task blocks along a brass track
DevOps
•
5 min read

systemd Timers vs Cron for Business-Critical Jobs

Choose systemd timers or cron for supplier jobs, with a runnable timer example and checks for missed runs, overlapping work, and business completion.

Oct 7, 2026
Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes
DevOps
•
4 min read

Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes

Moving a production Node and Postgres app from Vercel and Supabase to a single-node k3s cluster on Hetzner: data layer, cluster, backups, CI/CD and cost.

Oct 1, 2026
Postgres recovery workflow: retrieve an archive, restore a database, verify the application
DevOps
•
6 min read

Postgres Backups on Kubernetes: Run a Restore Drill

Test Postgres backups on Kubernetes with an isolated restore drill, an off-server archive, application checks, and a record of actual recovery time.

Oct 1, 2026
Back to the blog

© 2026 Muhammad Saad • Colophon

Connect with me on LinkedIn

Elsewhere

  • Github
  • Testimonials
  • CV
  • LinkedIn

Contact

  • Book a call
  • Email
On this page
  1. Does exit code 137 prove a memory problem?
  2. What should you capture before changing anything?
  3. Is the container limit wrong, or is the node under pressure?
  4. What does my Dahlia setup add to the decision?
  5. How would I choose and verify a fix?
  6. Where should you start?

Debugging OOMKilled pods starts with identifying which container stopped and what the node was doing at that time. Before raising a memory limit, I would preserve the termination evidence and check whether the proposed increase fits the whole server.

Quick answer

Inspect the affected container's current and previous termination states; exit code 137 alone does not establish the cause. Separate an OOM kill from Pod eviction, then compare container demand with node capacity. Change the limit, workload, or placement according to that evidence, and repeat the triggering workload before calling the problem fixed.

Does exit code 137 prove a memory problem?

I would treat the exit code as a starting point, then read the termination reason and timestamps. Kubernetes' memory allocation walkthrough demonstrates an over-limit container with reason: OOMKilled and exitCode: 137. That combination is stronger evidence than the number alone.

There are other reasons for a forced kill. The official Pod termination flow explains that processes still running after the termination grace period receive SIGKILL. Check whether the failure coincided with a rollout or deletion before diagnosing memory exhaustion from an exit code.

Also distinguish the container's state from the Pod's displayed status. Kubernetes documents CrashLoopBackOff as restart backoff, rather than a specific failure cause. I would record the underlying termination reason instead of making the backoff message the diagnosis.

My original migration article used “out of memory” as shorthand for exit code 137. The more useful operational rule is to confirm the reason and investigate the scope before changing resources.

What should you capture before changing anything?

Start with the Pod name, namespace, container name, node, and time of failure. Save the Pod description and full status while they are available. Read both the current state and lastState: a container may already have restarted when you inspect it.

Run the following in Bash with kubectl configured for the affected cluster and permission to read Pods and nodes. Set the variables to existing object names; these commands inspect resources and do not restart them. Store the output in an access-controlled incident record because Pod descriptions can include application configuration.

Capture Pod status and inspect its assigned node
set -euo pipefail
: "${POD_NAMESPACE:?Set the affected namespace}"
: "${POD_NAME:?Set the affected Pod name}"

kubectl -n "$POD_NAMESPACE" describe pod "$POD_NAME"
kubectl -n "$POD_NAMESPACE" get pod "$POD_NAME" -o yaml

POD_NODE=$(kubectl -n "$POD_NAMESPACE" get pod "$POD_NAME" \
  -o jsonpath='{.spec.nodeName}')
if [ -n "$POD_NODE" ]; then
  kubectl describe node "$POD_NODE"
fi

In the Pod output, match the container name to its requests, limits, termination reason, restart count, and image. Inspect init-container status too if startup failed. In the node description, look at conditions, resource allocation, and events alongside the failure timestamp.

I would also retrieve the relevant application and host logs through the team's existing logging tools. Missing evidence should remain marked missing. A healthy replacement container does not tell me what happened to the previous process.

Is the container limit wrong, or is the node under pressure?

These questions lead to different changes. Kubernetes schedules using memory requests; a memory limit constrains a container's consumption. Raising the request alone does not raise its limit, and raising the limit alone does not reserve additional capacity for scheduling.

The node-pressure eviction documentation describes a separate path: the kubelet can terminate Pods to reclaim scarce node resources. If memory runs out before reclamation succeeds, the kernel's OOM killer can act. An OOM-killed container therefore needs node context as well as its configured limit.

Evidence to collect before selecting a change
ObservationWhat I would check nextChange to evaluate
Container reports OOMKilledIts limit, demand around the failure, and node evidenceReduce demand or test a justified limit increase
Pod reports EvictedIts message and the node's resource-pressure eventsAddress the named resource pressure and workload placement
Exit 137 appears during terminationRollout timeline and shutdown behaviorInvestigate termination before changing memory
Failure follows a new imageSame workload against the previous imageRoll back or fix a demonstrated regression
Failures align with a scheduled jobActual job start and finish times across the nodeTest reduced concurrency or a different schedule

This table is a triage plan, not a set of automatic diagnoses. A time correlation tells me where to investigate; it does not prove that a job or deployment caused the failure.

What does my Dahlia setup add to the decision?

In my Dahlia migration from Vercel to k3s, I put the Node application and Postgres on one Hetzner server. The app's documented memory request was 192Mi, with a 768Mi limit. Its HPA allowed 2 to 6 replicas at 70% CPU.

Those are configuration values from that project, not a recommended memory profile for another application. I chose one image and one Deployment because one Express app served the widget, admin, and API. Its caches rebuilt in memory on boot, so startup belongs in any proposed memory test for this design.

The same cluster ran Postgres as a StatefulSet, a profile sweep every 30 minutes, a backup at 03:30, and a catalog refresh at 04:00. Before giving each app replica more memory, I would inventory those workloads together. I would include their actual durations rather than assume different start times prevent overlap.

My migration record does not establish an OOM incident, a measured memory peak, or a tested replacement limit. The procedure here is what I recommend for diagnosing that architecture. I would not turn the documented 768Mi setting into a claim that more memory solved a production failure.

How would I choose and verify a fix?

First, write down the proposed explanation in a testable form. For example: “Cache rebuilding during startup exceeds the app container's current allowance.” That is a hypothesis to test with the same image and representative data in an isolated environment, not an incident reported from Dahlia.

For a justified limit increase, I would prepare a node budget listing app replicas, Postgres, cluster services, and jobs that may overlap. Record configured requests and limits separately from observed demand. Leave explicit capacity for the host and uncertainty instead of treating the app's allowance as the server's entire budget.

If evidence points to excessive work, I would test smaller batches or lower concurrency before selecting a permanent resource setting. If a new image introduced the behavior, compare it with the previous version under the same conditions. My deployment pipeline used commit-SHA image tags so the version under investigation could be identified precisely.

I would accept a change only after the relevant workload finishes, restart evidence is reviewed, and neighboring services still work. For the Dahlia widget, API, and operator CMS, that means checking the affected application paths as well as the container. Record the image, data set, concurrency, observation period, and remaining uncertainty with the result.

Where should you start?

Capture one affected Pod's status and its node description before editing resource values. Fill in the evidence table, identify the missing measurements, and choose one explanation to test. Keep the original configuration beside the proposed change so another engineer can review the decision.

On a shared app-and-database server, I would also review the Postgres backup restore drill before planning disruptive recovery work. A successful restart is only the first observation. Close the investigation when the triggering workload and the services sharing its node have passed the agreed checks.