Skip to content
Go back

Temporal vs Cron: Durable Home Lab Jobs

By KingPin 15 min read
Temporal vs Cron: Durable Home Lab Jobs
Contents

Full example: Clone the working files at github.com/KingPin/sumguy-examples/devops/temporal-homelab-workflows

The 2 AM Backup That Died on Step 3

Your backup script has five steps: snapshot, compress, upload, verify, notify. At 2 AM the upload hits a DNS hiccup and the script exits. Steps 4 and 5 never run. You find out nine days later when you need the restore.

Cron did its job. It started the script. Everything after that is your problem, and a bash script has no memory of where it was.

Temporal is a workflow engine that remembers. You write the five steps as code, and the server records each completed step in a history. If the worker dies during step 3, a new worker replays the history, skips steps 1 and 2, and picks up at step 3. Retries, timeouts, and sleeps are built in.

Short version: for a single rsync on a timer, Temporal is a forklift moving a couch. For a multi-step job where a half-finished run is worse than no run, it earns its keep. This article shows both sides and the exact files to run.

What Cron Actually Gives You

Cron gives you one thing: a process starts at a time. Here is the typical setup.

0 2 * * * /opt/scripts/nightly-backup.sh >> /var/log/backup.log 2>&1

Inside that script you get whatever you wrote by hand. A set -e at the top means the first failure kills the run. Without it, a failed upload sails straight into a “verify” step that verifies nothing. Retries mean a for loop with sleep. Timeouts mean wrapping commands in timeout. Resuming from step 3 means writing marker files and checking them at the top of every step.

You can build all of that. People do, and the result is a 300-line script that only its author understands. A systemd timer improves the scheduling half: Persistent=true catches missed runs, journald keeps the logs, and OnFailure= can fire an alert unit. The systemd timers vs cron comparison covers that ground. It still doesn’t fix the multi-step problem. A service unit is one process, and when that process dies, the state dies with it.

Silent failure is the other gap. If you have ever discovered a dead job by accident, cron silent failures has the checklist, and a dead-man’s switch like Healthchecks closes most of it. That combo (timer plus Healthchecks plus a tidy script) is the honest baseline you should compare Temporal against.

What Temporal Gives You Instead

Temporal splits your job into two kinds of code.

A worker is a process you run that executes both. The server holds the history and hands tasks to workers. When a worker dies, tasks wait on the server until a worker comes back. Nothing is lost, because the server, not the worker, holds the state.

That gives you four things a bash script has to fake:

  1. Resume after a crash. Kill the worker mid-run, restart it, and the workflow continues from the last completed step.
  2. Retry policies as data. Backoff, max attempts, and per-step timeouts are arguments, not shell loops.
  3. Durable timers. workflow.sleep for six hours survives worker restarts and server restarts.
  4. A history you can read. The web UI shows every step, every attempt, every error message, with timestamps.

Running Temporal on One Box

There are two honest ways to run it at home.

Option 1: the dev server (start here)

The Temporal CLI bundles a whole server. One command, no containers.

Terminal window
temporal server start-dev --db-filename ./temporal.db

The frontend listens on port 7233 and the web UI on 8233. The --db-filename flag matters: without it, workflow history lives in memory and disappears when the process stops. With it, history lives in a SQLite file. Temporal labels this a development server, so treat it as a way to learn and to run low-stakes jobs, not as something to build a company on. For a home lab, plenty of people will never need more.

Option 2: Postgres, server, and UI in Compose

If you want the real server topology, use containers. The old temporalio/docker-compose repo is archived (read-only since January 2026). Its README points to the samples-server repo, which holds the current compose files. Those files are labeled for local development and testing, and Temporal points production users at Helm charts on Kubernetes. On a single home lab box, a trimmed Postgres-only stack is a reasonable middle ground.

The file below is that trimmed stack, modeled on the samples-server Postgres file. I pinned versions current as of late September 2026: server 1.32.0, UI 2.54.1.

docker-compose.yml
# Self-hosted Temporal for one box: Postgres + server + web UI.
# Modeled on temporalio/samples-server (compose/docker-compose-postgres.yml).
services:
postgres:
image: postgres:16
environment:
POSTGRES_USER: temporal
POSTGRES_PASSWORD: temporal
volumes:
- pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U temporal"]
interval: 5s
timeout: 5s
retries: 30
# One-shot: create both databases and apply the schemas, then exit.
schema-setup:
image: temporalio/admin-tools:1.32.0
depends_on:
postgres:
condition: service_healthy
entrypoint: ["/bin/sh", "-c"]
command:
- |
set -e
T="temporal-sql-tool --plugin postgres12 --ep postgres -u temporal -p 5432"
export SQL_PASSWORD=temporal
for db in temporal temporal_visibility; do
$$T --db $$db create
$$T --db $$db setup-schema -v 0.0
done
$$T --db temporal update-schema -d /etc/temporal/schema/postgresql/v12/temporal/versioned
$$T --db temporal_visibility update-schema -d /etc/temporal/schema/postgresql/v12/visibility/versioned
restart: "no"
temporal:
image: temporalio/server:1.32.0
depends_on:
schema-setup:
condition: service_completed_successfully
environment:
DB: postgres12
DB_PORT: "5432"
POSTGRES_USER: temporal
POSTGRES_PWD: temporal
POSTGRES_SEEDS: postgres
BIND_ON_IP: 0.0.0.0
DYNAMIC_CONFIG_FILE_PATH: config/dynamicconfig/development-sql.yaml
ports:
- "127.0.0.1:7233:7233"
volumes:
- ./dynamicconfig:/etc/temporal/config/dynamicconfig
healthcheck:
test: ["CMD", "nc", "-z", "localhost", "7233"]
interval: 5s
timeout: 3s
start_period: 20s
retries: 30
restart: unless-stopped
# One-shot: the server image does not create the "default" namespace itself.
create-namespace:
image: temporalio/admin-tools:1.32.0
depends_on:
temporal:
condition: service_healthy
environment:
TEMPORAL_ADDRESS: temporal:7233
entrypoint: ["/bin/sh", "-c"]
command:
- temporal operator namespace describe -n default || temporal operator namespace create -n default
restart: "no"
ui:
image: temporalio/ui:2.54.1
depends_on:
temporal:
condition: service_healthy
environment:
TEMPORAL_ADDRESS: temporal:7233
ports:
- "127.0.0.1:8080:8080"
restart: unless-stopped
volumes:
pgdata:

A few things to know about this file:

dynamicconfig/development-sql.yaml
limit.maxIDLength:
- value: 255
constraints: {}

Bring it up and wait for the healthy status.

Terminal window
docker compose up -d
docker compose ps -a

The UI is at http://localhost:8080. The server’s gRPC frontend is on localhost:7233.

The Nightly Backup as a Workflow

Now the five steps. Activities first. They are safe fakes: they write under /tmp and sleep. The upload activity fails on its first two attempts on purpose, so you can watch retries without breaking a network.

activities.py
import asyncio
import gzip
import shutil
from pathlib import Path
from temporalio import activity
WORK = Path("/tmp/temporal-homelab")
@activity.defn
async def snapshot(target: str) -> str:
"""Pretend to snapshot a dataset by writing a file."""
WORK.mkdir(exist_ok=True)
path = WORK / f"{target}.snap"
path.write_text("pretend this is 500 GB of photos\n" * 1000)
await asyncio.sleep(2)
activity.logger.info("snapshot done: %s", path)
return str(path)
@activity.defn
async def compress(snap: str) -> str:
out = snap + ".gz"
with open(snap, "rb") as src, gzip.open(out, "wb") as dst:
shutil.copyfileobj(src, dst)
await asyncio.sleep(2)
return out
@activity.defn
async def upload(archive: str) -> str:
"""Fake upload. Fails on attempts 1 and 2 to show retries."""
attempt = activity.info().attempt
await asyncio.sleep(1)
if attempt < 3:
raise ConnectionError(f"remote host unreachable (attempt {attempt})")
remote = WORK / "remote" / Path(archive).name
remote.parent.mkdir(exist_ok=True)
shutil.copy(archive, remote)
return str(remote)
@activity.defn
async def verify(remote: str) -> None:
with gzip.open(remote, "rb") as f:
if not f.read(16).startswith(b"pretend"):
raise ValueError("archive content is wrong")
@activity.defn
async def notify(message: str) -> None:
print(f"NOTIFY: {message}", flush=True)

activity.info().attempt is the current attempt number, starting at 1. In real life you would replace the body of each activity with a restic call, an rclone copy, or a curl to your notification service.

The workflow wires them together.

workflows.py
from dataclasses import dataclass
from datetime import timedelta
from temporalio import workflow
from temporalio.common import RetryPolicy
with workflow.unsafe.imports_passed_through():
from activities import compress, notify, snapshot, upload, verify
@dataclass
class BackupInput:
target: str = "photos"
settle_seconds: int = 10
@workflow.defn
class NightlyBackup:
@workflow.run
async def run(self, inp: BackupInput) -> str:
snap = await workflow.execute_activity(
snapshot, inp.target, start_to_close_timeout=timedelta(minutes=5)
)
archive = await workflow.execute_activity(
compress, snap, start_to_close_timeout=timedelta(minutes=30)
)
# Flaky network step: retry with backoff, give up after 5 tries.
remote = await workflow.execute_activity(
upload,
archive,
start_to_close_timeout=timedelta(minutes=10),
retry_policy=RetryPolicy(
initial_interval=timedelta(seconds=2),
backoff_coefficient=2.0,
maximum_attempts=5,
),
)
# Durable timer: survives worker restarts and server restarts.
await workflow.sleep(timedelta(seconds=inp.settle_seconds))
await workflow.execute_activity(
verify, remote, start_to_close_timeout=timedelta(minutes=10)
)
msg = f"backup of {inp.target} verified at {remote}"
await workflow.execute_activity(
notify, msg, start_to_close_timeout=timedelta(seconds=30)
)
return msg

Read that top to bottom. It looks like a bash script with better manners, and that is the appeal. Three details are worth knowing.

workflow.unsafe.imports_passed_through() tells the Python SDK’s sandbox not to reload your activity module inside the workflow. Without it you get slow startup and odd import errors. Activities are imported here only as references.

Every step has a start_to_close_timeout. Temporal requires you to set either that or schedule_to_close_timeout for each activity. It caps how long one attempt may run. Set it to something generous: a compress step on a big dataset can take an hour.

The retry policy on upload says: first retry after 2 seconds, double the wait each time, give up after 5 attempts. Activities that don’t set a policy get the default, which retries forever with backoff. That default is friendly for flaky APIs and dangerous for a step that will never succeed, so set maximum_attempts when a step has a real failure mode.

The sleep between upload and verify is a durable timer. Ten seconds is for the demo. In real life it might be an hour, to let a remote object store settle, or a day, to wait for a cold-storage restore test. The timer lives on the server, so a worker restart does not reset it.

The worker and the starter

The worker registers the workflow and activities on a task queue and waits.

worker.py
import asyncio
from temporalio.client import Client
from temporalio.worker import Worker
from activities import compress, notify, snapshot, upload, verify
from workflows import NightlyBackup
TASK_QUEUE = "homelab"
async def main() -> None:
client = await Client.connect("localhost:7233")
worker = Worker(
client,
task_queue=TASK_QUEUE,
workflows=[NightlyBackup],
activities=[snapshot, compress, upload, verify, notify],
)
print("worker up, waiting for tasks")
await worker.run()
if __name__ == "__main__":
asyncio.run(main())

The starter kicks off one run and waits for the result.

starter.py
import asyncio
import sys
import time
from temporalio.client import Client
from workflows import BackupInput, NightlyBackup
async def main() -> None:
target = sys.argv[1] if len(sys.argv) > 1 else "photos"
client = await Client.connect("localhost:7233")
result = await client.execute_workflow(
NightlyBackup.run,
BackupInput(target=target),
id=f"nightly-backup-{target}-{int(time.time())}",
task_queue="homelab",
)
print(f"result: {result}")
if __name__ == "__main__":
asyncio.run(main())

The pyproject.toml pins the SDK. Version 1.33.0 needs Python 3.10 or newer.

pyproject.toml
[project]
name = "temporal-homelab-workflows"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = ["temporalio==1.33.0"]

Watching It Survive a Crash

Run the worker in one terminal and the starter in another.

Terminal window
uv run python worker.py
uv run python starter.py photos

The upload fails twice, the retries back off, and the third attempt succeeds. The starter prints:

result: backup of photos verified at /tmp/temporal-homelab/remote/photos.snap.gz

Now the fun part. Start another run, and press Ctrl+C on the worker while the upload is retrying. The workflow shows as Running in the UI, and it just waits. Start the worker again. In my test the worker was down for about 45 seconds, and the run finished right after the restart, with steps 1 and 2 not repeated. The event history shows the timer starting and firing ten seconds later, then verify and notify completing. Compare that to your bash script, which would have been dead at the first exit 1.

That resume is why you run a workflow engine. Everything else is convenience.

Replacing the Crontab Line

Temporal has its own scheduler. A Schedule starts a workflow on a cron expression, and the server keeps it running whether or not your worker is up.

Terminal window
temporal schedule create \
--schedule-id nightly-photos \
--cron "0 2 * * *" \
--workflow-id nightly-backup-photos \
--type NightlyBackup \
--task-queue homelab \
--input '{"target":"photos","settle_seconds":10}'

The cron expression is evaluated in UTC unless you set a timezone. The --input JSON maps to the BackupInput dataclass fields. If the worker is down at 2 AM, the workflow task waits on the queue, and the run happens when the worker returns.

When Temporal Earns Its Keep

Use it when the job matches at least two of these:

Good home lab candidates: nightly backup with a restore test, a media pipeline (download, transcode, tag, move, refresh library), a certificate rotation across several hosts, or a staged update of a Docker fleet with health checks between hosts.

When It Is a Forklift Moving a Couch

Skip it when:

Temporal replaces script complexity with infrastructure complexity. Below three steps, the script wins. Above that, the workflow code stays readable while the script turns into spaghetti, and the infrastructure cost is fixed.

There is also a learning cost. Determinism rules are real: no random, no datetime.now(), and no network calls inside workflow code. Put anything nondeterministic in an activity. The sandbox catches many mistakes, but you will hit one eventually, and your 2 AM self will thank you for reading the docs on that before it happens.

The SumGuy Take

Start with the systemd timer plus Healthchecks. It handles most home lab jobs without breaking a sweat. When you catch yourself writing marker files to remember which step you were on, or a retry loop with three nested if blocks, that is the signal. Run temporal server start-dev --db-filename ./temporal.db, port one job, and see whether the UI makes you feel smarter or just busier. If it helps, move to the Compose stack. If it doesn’t, you lost an evening and gained a better understanding of your own script.

Common Questions

Does Temporal need Kubernetes to self-host?

No. Temporal runs from a single CLI command (temporal server start-dev) or from Docker Compose with Postgres, server, and UI containers. Temporal’s own compose files are labeled for development and testing, and its production guidance points to Kubernetes with Helm charts. A single home lab box can use Compose fine.

Does Temporal lose workflow history if the dev server restarts?

Only if you skip the flag. temporal server start-dev keeps history in memory unless you pass --db-filename with a file path. With that flag, the dev server stores history in a SQLite file, so workflows survive restarts. The Compose stack in this article stores history in Postgres on a named volume.

How do I upgrade a self-hosted Temporal server?

Upgrade Temporal one minor version at a time, never skipping one. Back up the Postgres volume, bump the server, admin-tools, and UI tags to the next minor version, then run docker compose up -d. The example’s schema-setup container applies migrations with temporal-sql-tool before the new server starts. Read each version’s release notes, then repeat.

Which Python version does the Temporal SDK need?

Python 3.10 or newer. The temporalio package at version 1.33.0 declares that minimum on PyPI. The example project pins temporalio==1.33.0 and sets requires-python = ">=3.10" in its pyproject.toml, and uv picks a matching interpreter for you.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Next Post
VLAN Trunking Gotchas Nobody Warns You

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts