Work

/

About
Play
Blog
Home/Blog/DevOps
DevOps
•
Updated Oct 1, 2026
•
4 min read

Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes

Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes
MS
Muhammad Saad

Shopify engineer. I build storefronts, two published Shopify apps, and the infrastructure behind a 70,000+ product store, and I write here about what that work teaches me.

Follow on LinkedInSee my workBook a call
Found this useful? Share it.

Keep reading

Postgres recovery workflow: retrieve an archive, restore a database, verify the application
DevOps
•
6 min read

Postgres Backups on Kubernetes: Run a Restore Drill

Test Postgres backups on Kubernetes with an isolated restore drill, an off-server archive, application checks, and a record of actual recovery time.

Oct 1, 2026
Running a 70,000-Product Shopify Store on Seven Suppliers: Stock, Pricing and Orders
Shopify
•
3 min read

Running a 70,000-Product Shopify Store on Seven Suppliers: Stock, Pricing and Orders

How a 70,000-product Shopify store syncs stock, prices and orders across seven suppliers without spreadsheets: dry runs, feed-health holds and safe repricing.

Oct 1, 2026
Shopify Collection Sorting for Large Catalogs: Rank by What Actually Sells
Shopify
•
4 min read

Shopify Collection Sorting for Large Catalogs: Rank by What Actually Sells

Rank Shopify collections on revenue, margin, stock and click-through instead of one sort order, and re-sort large catalogs hourly without API limits.

Oct 1, 2026
Back to the blog

© 2026 Muhammad Saad • Colophon

Connect with me on LinkedIn

Elsewhere

  • Github
  • Testimonials
  • CV
  • LinkedIn

Contact

  • Book a call
  • Email
On this page
  1. First, Be Honest About Why
  2. The Data Layer Is the Real Migration
  3. Prepare the Server Before Installing Anything
  4. The Cluster Layout
  5. Backups Are Now Your Job
  6. CI/CD by Commit SHA
  7. When Something Breaks
  8. Was It Worth It?

I moved Dahlia, an AI sales concierge running on live Shopify stores, from Vercel and Supabase to a k3s cluster on a Hetzner server. This is what it took, what I would do the same way again, and what surprised me.

Quick answer

Make the app stateless, replace hosted database calls with plain Postgres, firewall the server before installing k3s, disable Traefik and the service load balancer, use ingress-nginx with cert-manager, run Postgres as a StatefulSet, take backups off the box, and deploy by commit SHA from CI.

First, Be Honest About Why

Dahlia did not need Kubernetes. It is one Node and Express process talking to a few external APIs, and it ran fine on Vercel. What k3s bought was owning the data, a separate namespace per client store, self-healing and autoscaling, and a hosting bill of about €8 a month instead of two platform subscriptions.

It was also a good candidate. Its caches are rebuilt in memory on boot, so replicas just work. One Express app serves the widget, the admin and the API, so it is one image and one Deployment. And the widget finds its server from its own script URL, so cutting over is one script tag on the Shopify theme.

The Data Layer Is the Real Migration

The app talked to Supabase through its hosted HTTP layer, not to Postgres directly. Running Postgres in a pod does not answer those calls. Rather than self-host the whole Supabase stack, I replaced 23 query call sites across five tables with plain SQL through the pg driver, then ran Postgres as a StatefulSet on a persistent volume.

Prepare the Server Before Installing Anything

Firewall first, then k3s without Traefik or the service load balancer
ufw default deny incoming
ufw default allow outgoing
ufw allow 22/tcp && ufw allow 80/tcp && ufw allow 443/tcp
ufw allow 6443/tcp   # kube API, for CI deploys
ufw enable

curl -sfL https://get.k3s.io | INSTALL_K3S_EXEC="\
  --disable traefik \
  --disable servicelb \
  --write-kubeconfig-mode 644" sh -

Firewall before k3s, because k3s opens ports itself. Add swap on small servers; k3s and Postgres both need the headroom. Disable Traefik in favour of ingress-nginx, which is what most documentation assumes, and disable the bundled load balancer, since on one server ingress-nginx can bind ports 80 and 443 directly.

My bootstrap script refuses to run if ports 80 or 443 are taken or there is less than about 2.2 GB of free memory. That check exists because the first server it was pointed at turned out to be a live host serving five domains. Installing there would have taken them all down.

The Cluster Layout

What runs where
PieceKubernetes objectNotes
AppDeployment + HPA2 to 6 replicas at 70% CPU; requests 100m and 192Mi, limit 768Mi
DatabaseStatefulSet + PVCOne replica; the disk mounts on one node only
TLSingress-nginx + cert-managerLet's Encrypt certificates renewed automatically
JobsCronJobsProfile sweep every 30 minutes, backup at 03:30, catalog refresh at 04:00
ClientsKustomize overlaysShared base, dev and prod overlays, one namespace per client store

Backups Are Now Your Job

Supabase handled backups quietly. Now a CronJob dumps the database daily and keeps 14 days, but on the same server, so it survives a bad migration and not a dead disk. Copy the dumps off the box, and do a restore drill into the dev namespace at least once. An untested backup is not a backup.

The Postgres backup restore drill turns that recommendation into an acceptance sheet: retrieve an off-server archive, restore in isolation, and verify the application before claiming a recovery time.

CI/CD by Commit SHA

  • GitHub Actions runs the checks, builds the image with Buildx and pushes it to GitHub Container Registry.
  • Images are tagged by commit SHA, so "what is deployed?" and "roll back to what?" both have exact answers.
  • Pull requests build the image but never publish it.
  • The deploy job is skipped until the cluster credentials exist, so the pipeline could be built before the server was.

When Something Breaks

Symptoms and what they usually mean
SymptomUsually means
ImagePullBackOffThe registry package is private, or the tag does not exist
CrashLoopBackOffThe app exits on boot; read the previous container's logs
Exit code 137Out of memory; raise the limit or find the leak
Pod stuck PendingNothing can schedule it, usually a volume that cannot bind
Ready check fails, health check passesThe database is unreachable, which is the probes working as designed

Was It Worth It?

For learning and for owning the data, yes. Not everything needs a cluster, though: the operations platform for a 70,000-product store runs happily as systemd services on one server. The trade is clear: cheaper hosting and full control in exchange for owning uptime and backups. If you would rather someone else handle that for your app, book a call, or see how Dahlia works in the case study.