Work

/

About
Play
Blog
Home/Blog/DevOps
DevOps
•
Updated Oct 7, 2026
•
5 min read

systemd Timers vs Cron for Business-Critical Jobs

Ivory clockwork driving green task blocks along a brass track
MS
Muhammad Saad

Shopify engineer. I build storefronts, two published Shopify apps, and the infrastructure behind a 70,000+ product store, and I write here about what that work teaches me.

Follow on LinkedInSee my workBook a call
Found this useful? Share it.

Keep reading

A full glass container of green blocks with an extra gold block above its capacity
DevOps
•
5 min read

Debugging OOMKilled Pods: Exit Code 137 on a Small Cluster

Diagnose OOMKilled pods before raising memory limits: preserve container evidence, check node pressure, and test changes against a shared k3s budget.

Oct 7, 2026
Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes
DevOps
•
4 min read

Moving a Node App from Vercel to k3s on Hetzner: What It Actually Takes

Moving a production Node and Postgres app from Vercel and Supabase to a single-node k3s cluster on Hetzner: data layer, cluster, backups, CI/CD and cost.

Oct 1, 2026
Postgres recovery workflow: retrieve an archive, restore a database, verify the application
DevOps
•
6 min read

Postgres Backups on Kubernetes: Run a Restore Drill

Test Postgres backups on Kubernetes with an isolated restore drill, an off-server archive, application checks, and a record of actual recovery time.

Oct 1, 2026
Back to the blog

© 2026 Muhammad Saad • Colophon

Connect with me on LinkedIn

Elsewhere

  • Github
  • Testimonials
  • CV
  • LinkedIn

Contact

  • Book a call
  • Email
On this page
  1. What my store operations platform actually uses
  2. When would I keep cron or choose a timer?
  3. How do you try a timer without touching store data?
  4. What does Persistent actually guarantee?
  5. How would I prove that the scheduled job worked?
  6. Where should you start?

Choosing systemd timers versus cron for a business job starts with what should happen when the scheduled time is missed. I use systemd timers in BC Supply Ops, where recurring automation supports catalog, stock, pricing, and supplier operations. My selection checklist asks how a job catches up, avoids competing writes, and proves that its intended work finished.

Quick answer

On a Linux host using systemd, choose a timer when you want the schedule attached to a named service and need explicit catch-up behavior. Use an OnCalendar timer with Persistent=true when a missed trigger should run after reactivation. Keep completion records and write safeguards in the application: a timer trigger is not proof that a supplier update succeeded.

What my store operations platform actually uses

The BC Supply Ops case study documents a 70,000+ product store, 7 suppliers, and 10+ systemd services and timers. The platform runs on a single VPS behind nginx and TLS. Recurring jobs run through timers; operators also trigger guarded actions.

Those actions have dry-run plans, write flags, paced workers, and audit logs. I would preserve that separation when changing a schedule: the operating system decides when to start, while the application decides whether a write is allowed. Changing cron syntax into a timer file must not remove those decisions.

The same case study also documents separate product-feed rebuilds on cron cadences. I am not claiming every recurring task was migrated to systemd. The examples below are a proposed starting point for a new job, not copies of the production unit files.

When would I keep cron or choose a timer?

Debian's crontab reference describes commands selected by matching time and date fields. For an existing cron job, I would first inspect the surrounding implementation. It may already have locking, failure reporting, and an application record of unfinished work.

I would keep that job if its behavior is understood and meets the business requirement. A different scheduler alone does not make a price update safer. Before migrating, write down what the current command does when run manually and what evidence an operator uses to judge it.

Questions to settle before changing a recurring job
RequirementWhat I would recordAcceptance check
Missed executionCatch up, skip, or require reviewDisable scheduling across a due time in a test environment
Slow executionMaximum acceptable runtimeHold a test worker open across the next trigger
Competing writersEvery scheduled and manual entry pointTry the operator action while scheduled work is active
Partial completionDurable record of completed itemsInterrupt a test batch and inspect its remaining work
FreshnessLatest acceptable source timestampSupply an old snapshot and expect a hold
OwnershipPerson responsible for each failure stateConfirm the operator can find the failed run and its next step

I would favor a systemd timer for a new host-based job when a named service and explicit missed-trigger policy fit the operating model. For a cluster workload, I would evaluate its existing job mechanism instead; my Dahlia hosting walkthrough describes a different deployment environment.

How do you try a timer without touching store data?

Here is a harmless demonstration for a Linux machine with systemd. Save the files under /etc/systemd/system/ using an administrator account. The service only prints a message; it has no store credentials and performs no business work.

business-job-demo.service — a harmless one-shot worker
[Unit]
Description=Demonstrate a scheduled business job

[Service]
Type=oneshot
ExecStart=/usr/bin/echo Timer demonstration completed
TimeoutStartSec=30s
business-job-demo.timer — a daily catch-up policy
[Unit]
Description=Run the business job demonstration daily

[Timer]
OnCalendar=*-*-* 09:00:00 UTC
Persistent=true
Unit=business-job-demo.service

[Install]
WantedBy=timers.target

The systemd service reference describes Type=oneshot and TimeoutStartSec. The timeout here is a demonstration setting, not a recommended limit for supplier imports. Before replacing the echo command, choose a dedicated service account, explicit executable paths, and the directories that the real worker needs.

On that Linux machine, these commands validate the files and exercise the service. Check that /usr/bin/echo exists before starting.

Validate, run, and inspect the demonstration
sudo systemd-analyze verify /etc/systemd/system/business-job-demo.service /etc/systemd/system/business-job-demo.timer
systemd-analyze calendar '*-*-* 09:00:00 UTC'
sudo systemctl daemon-reload
sudo systemctl start business-job-demo.service
sudo journalctl -u business-job-demo.service -n 20 --no-pager
sudo systemctl enable --now business-job-demo.timer
systemctl list-timers --all business-job-demo.timer
# Remove the recurring demonstration when finished:
sudo systemctl disable --now business-job-demo.timer

I would inspect the displayed next trigger before leaving a real schedule enabled. Use the business's intended timezone and check the host's clock. Keep the demo separate from existing job names so it cannot accidentally replace a working service.

What does Persistent actually guarantee?

The systemd timer reference explains that Persistent=true records the last trigger and catches up when an inactive calendar timer missed an activation. It does not replay a separate job for every missed business date. It also does not track whether the application completed successfully.

That distinction changes how I would design the worker. A current-stock reconciliation should fetch and validate a current snapshot after downtime. A purchase-order dispatcher should inspect its durable ledger before considering anything already attempted.

I would not make either decision from the timer timestamp. The supplier feed hold workflow gives the stock example: a job can start on time and still need to reject its input. Give that outcome a visible status rather than treating “process ran” as “inventory is current.”

The timer reference also specifies that an already active target unit is left running instead of being started again. That protection covers that unit. I would still use an application lock or equivalent coordination wherever a portal action, another service, or another host can reach the same writer.

How would I prove that the scheduled job worked?

I would keep separate timestamps for attempted execution and successful business completion. The run record should include the input version, operation type, outcome, and a pointer to detailed logs. For held work, include the reason and the person or process expected to release it.

My proposed acceptance exercise would include a successful run, an unavailable source, an incomplete source, and an interruption after partial progress. Repeat the interruption test with a manual trigger to expose competing entry points. These are recommended tests, not reported production incidents or results from BC Supply Ops.

I would also choose a business deadline for alerting. The useful signal is that a required reconciliation has not completed by its deadline, including a run that was held or never started. An exit status and the newest completed ledger entry should be available together when investigating.

Where should you start?

Pick one recurring job and write its missed-run policy before changing its scheduler. Try the harmless unit pair, then exercise the real worker in a test environment with writes disabled. Keep a record of what happens after downtime and interruption.

Use the multi-supplier operations overview to identify the surrounding write guards. A migration is ready when another operator can identify the last completed task, explain outstanding work, and follow a documented recovery step.