I moved Dahlia, an AI sales concierge running on live Shopify stores, from Vercel and Supabase to a k3s cluster on a Hetzner server. This is what it took, what I would do the same way again, and what surprised me.
First, Be Honest About Why
Dahlia did not need Kubernetes. It is one Node and Express process talking to a few external APIs, and it ran fine on Vercel. What k3s bought was owning the data, a separate namespace per client store, self-healing and autoscaling, and a hosting bill of about €8 a month instead of two platform subscriptions.
It was also a good candidate. Its caches are rebuilt in memory on boot, so replicas just work. One Express app serves the widget, the admin and the API, so it is one image and one Deployment. And the widget finds its server from its own script URL, so cutting over is one script tag on the Shopify theme.
The Data Layer Is the Real Migration
The app talked to Supabase through its hosted HTTP layer, not to Postgres directly. Running Postgres in a pod does not answer those calls. Rather than self-host the whole Supabase stack, I replaced 23 query call sites across five tables with plain SQL through the pg driver, then ran Postgres as a StatefulSet on a persistent volume.
Prepare the Server Before Installing Anything
ufw default deny incoming
ufw default allow outgoing
ufw allow 22/tcp && ufw allow 80/tcp && ufw allow 443/tcp
ufw allow 6443/tcp # kube API, for CI deploys
ufw enable
curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="\
--disable traefik \
--disable servicelb \
--write-kubeconfig-mode 644" sh -Firewall before k3s, because k3s opens ports itself. Add swap on small servers; k3s and Postgres both need the headroom. Disable Traefik in favour of ingress-nginx, which is what most documentation assumes, and disable the bundled load balancer, since on one server ingress-nginx can bind ports 80 and 443 directly.
My bootstrap script refuses to run if ports 80 or 443 are taken or there is less than about 2.2 GB of free memory. That check exists because the first server it was pointed at turned out to be a live host serving five domains. Installing there would have taken them all down.
The Cluster Layout
| Piece | Kubernetes object | Notes |
|---|---|---|
| App | Deployment + HPA | 2 to 6 replicas at 70% CPU; requests 100m and 192Mi, limit 768Mi |
| Database | StatefulSet + PVC | One replica; the disk mounts on one node only |
| TLS | ingress-nginx + cert-manager | Let's Encrypt certificates renewed automatically |
| Jobs | CronJobs | Profile sweep every 30 minutes, backup at 03:30, catalog refresh at 04:00 |
| Clients | Kustomize overlays | Shared base, dev and prod overlays, one namespace per client store |
Backups Are Now Your Job
Supabase handled backups quietly. Now a CronJob dumps the database daily and keeps 14 days, but on the same server, so it survives a bad migration and not a dead disk. Copy the dumps off the box, and do a restore drill into the dev namespace at least once. An untested backup is not a backup.
The Postgres backup restore drill turns that recommendation into an acceptance sheet: retrieve an off-server archive, restore in isolation, and verify the application before claiming a recovery time.
CI/CD by Commit SHA
- GitHub Actions runs the checks, builds the image with Buildx and pushes it to GitHub Container Registry.
- Images are tagged by commit SHA, so "what is deployed?" and "roll back to what?" both have exact answers.
- Pull requests build the image but never publish it.
- The deploy job is skipped until the cluster credentials exist, so the pipeline could be built before the server was.
When Something Breaks
| Symptom | Usually means |
|---|---|
| ImagePullBackOff | The registry package is private, or the tag does not exist |
| CrashLoopBackOff | The app exits on boot; read the previous container's logs |
| Exit code 137 | Out of memory; raise the limit or find the leak |
| Pod stuck Pending | Nothing can schedule it, usually a volume that cannot bind |
| Ready check fails, health check passes | The database is unreachable, which is the probes working as designed |
Was It Worth It?
For learning and for owning the data, yes. Not everything needs a cluster, though: the operations platform for a 70,000-product store runs happily as systemd services on one server. The trade is clear: cheaper hosting and full control in exchange for owning uptime and backups. If you would rather someone else handle that for your app, book a call, or see how Dahlia works in the case study.