Debugging OOMKilled pods starts with identifying which container stopped and what the node was doing at that time. Before raising a memory limit, I would preserve the termination evidence and check whether the proposed increase fits the whole server.
Does exit code 137 prove a memory problem?
I would treat the exit code as a starting point, then read the termination reason and timestamps. Kubernetes' memory allocation walkthrough demonstrates an over-limit container with reason: OOMKilled and exitCode: 137. That combination is stronger evidence than the number alone.
There are other reasons for a forced kill. The official Pod termination flow explains that processes still running after the termination grace period receive SIGKILL. Check whether the failure coincided with a rollout or deletion before diagnosing memory exhaustion from an exit code.
Also distinguish the container's state from the Pod's displayed status. Kubernetes documents CrashLoopBackOff as restart backoff, rather than a specific failure cause. I would record the underlying termination reason instead of making the backoff message the diagnosis.
My original migration article used “out of memory” as shorthand for exit code 137. The more useful operational rule is to confirm the reason and investigate the scope before changing resources.
What should you capture before changing anything?
Start with the Pod name, namespace, container name, node, and time of failure. Save the Pod description and full status while they are available. Read both the current state and lastState: a container may already have restarted when you inspect it.
Run the following in Bash with kubectl configured for the affected cluster and permission to read Pods and nodes. Set the variables to existing object names; these commands inspect resources and do not restart them. Store the output in an access-controlled incident record because Pod descriptions can include application configuration.
set -euo pipefail
: "${POD_NAMESPACE:?Set the affected namespace}"
: "${POD_NAME:?Set the affected Pod name}"
kubectl -n "$POD_NAMESPACE" describe pod "$POD_NAME"
kubectl -n "$POD_NAMESPACE" get pod "$POD_NAME" -o yaml
POD_NODE=$(kubectl -n "$POD_NAMESPACE" get pod "$POD_NAME" \
-o jsonpath='{.spec.nodeName}')
if [ -n "$POD_NODE" ]; then
kubectl describe node "$POD_NODE"
fiIn the Pod output, match the container name to its requests, limits, termination reason, restart count, and image. Inspect init-container status too if startup failed. In the node description, look at conditions, resource allocation, and events alongside the failure timestamp.
I would also retrieve the relevant application and host logs through the team's existing logging tools. Missing evidence should remain marked missing. A healthy replacement container does not tell me what happened to the previous process.
Is the container limit wrong, or is the node under pressure?
These questions lead to different changes. Kubernetes schedules using memory requests; a memory limit constrains a container's consumption. Raising the request alone does not raise its limit, and raising the limit alone does not reserve additional capacity for scheduling.
The node-pressure eviction documentation describes a separate path: the kubelet can terminate Pods to reclaim scarce node resources. If memory runs out before reclamation succeeds, the kernel's OOM killer can act. An OOM-killed container therefore needs node context as well as its configured limit.
| Observation | What I would check next | Change to evaluate |
|---|---|---|
| Container reports OOMKilled | Its limit, demand around the failure, and node evidence | Reduce demand or test a justified limit increase |
| Pod reports Evicted | Its message and the node's resource-pressure events | Address the named resource pressure and workload placement |
| Exit 137 appears during termination | Rollout timeline and shutdown behavior | Investigate termination before changing memory |
| Failure follows a new image | Same workload against the previous image | Roll back or fix a demonstrated regression |
| Failures align with a scheduled job | Actual job start and finish times across the node | Test reduced concurrency or a different schedule |
This table is a triage plan, not a set of automatic diagnoses. A time correlation tells me where to investigate; it does not prove that a job or deployment caused the failure.
What does my Dahlia setup add to the decision?
In my Dahlia migration from Vercel to k3s, I put the Node application and Postgres on one Hetzner server. The app's documented memory request was 192Mi, with a 768Mi limit. Its HPA allowed 2 to 6 replicas at 70% CPU.
Those are configuration values from that project, not a recommended memory profile for another application. I chose one image and one Deployment because one Express app served the widget, admin, and API. Its caches rebuilt in memory on boot, so startup belongs in any proposed memory test for this design.
The same cluster ran Postgres as a StatefulSet, a profile sweep every 30 minutes, a backup at 03:30, and a catalog refresh at 04:00. Before giving each app replica more memory, I would inventory those workloads together. I would include their actual durations rather than assume different start times prevent overlap.
My migration record does not establish an OOM incident, a measured memory peak, or a tested replacement limit. The procedure here is what I recommend for diagnosing that architecture. I would not turn the documented 768Mi setting into a claim that more memory solved a production failure.
How would I choose and verify a fix?
First, write down the proposed explanation in a testable form. For example: “Cache rebuilding during startup exceeds the app container's current allowance.” That is a hypothesis to test with the same image and representative data in an isolated environment, not an incident reported from Dahlia.
For a justified limit increase, I would prepare a node budget listing app replicas, Postgres, cluster services, and jobs that may overlap. Record configured requests and limits separately from observed demand. Leave explicit capacity for the host and uncertainty instead of treating the app's allowance as the server's entire budget.
If evidence points to excessive work, I would test smaller batches or lower concurrency before selecting a permanent resource setting. If a new image introduced the behavior, compare it with the previous version under the same conditions. My deployment pipeline used commit-SHA image tags so the version under investigation could be identified precisely.
I would accept a change only after the relevant workload finishes, restart evidence is reviewed, and neighboring services still work. For the Dahlia widget, API, and operator CMS, that means checking the affected application paths as well as the container. Record the image, data set, concurrency, observation period, and remaining uncertainty with the result.
Where should you start?
Capture one affected Pod's status and its node description before editing resource values. Fill in the evidence table, identify the missing measurements, and choose one explanation to test. Keep the original configuration beside the proposed change so another engineer can review the decision.
On a shared app-and-database server, I would also review the Postgres backup restore drill before planning disruptive recovery work. A successful restart is only the first observation. Close the investigation when the triggering workload and the services sharing its node have passed the agreed checks.

