Full example: Clone the working files at github.com/KingPin/sumguy-examples/devops/temporal-homelab-workflows
The 2 AM Backup That Died on Step 3
Your backup script has five steps: snapshot, compress, upload, verify, notify. At 2 AM the upload hits a DNS hiccup and the script exits. Steps 4 and 5 never run. You find out nine days later when you need the restore.
Cron did its job. It started the script. Everything after that is your problem, and a bash script has no memory of where it was.
Temporal is a workflow engine that remembers. You write the five steps as code, and the server records each completed step in a history. If the worker dies during step 3, a new worker replays the history, skips steps 1 and 2, and picks up at step 3. Retries, timeouts, and sleeps are built in.
Short version: for a single rsync on a timer, Temporal is a forklift moving a couch. For a multi-step job where a half-finished run is worse than no run, it earns its keep. This article shows both sides and the exact files to run.
What Cron Actually Gives You
Cron gives you one thing: a process starts at a time. Here is the typical setup.
0 2 * * * /opt/scripts/nightly-backup.sh >> /var/log/backup.log 2>&1Inside that script you get whatever you wrote by hand. A set -e at the top means the first failure kills the run. Without it, a failed upload sails straight into a “verify” step that verifies nothing. Retries mean a for loop with sleep. Timeouts mean wrapping commands in timeout. Resuming from step 3 means writing marker files and checking them at the top of every step.
You can build all of that. People do, and the result is a 300-line script that only its author understands. A systemd timer improves the scheduling half: Persistent=true catches missed runs, journald keeps the logs, and OnFailure= can fire an alert unit. The systemd timers vs cron comparison covers that ground. It still doesn’t fix the multi-step problem. A service unit is one process, and when that process dies, the state dies with it.
Silent failure is the other gap. If you have ever discovered a dead job by accident, cron silent failures has the checklist, and a dead-man’s switch like Healthchecks closes most of it. That combo (timer plus Healthchecks plus a tidy script) is the honest baseline you should compare Temporal against.
What Temporal Gives You Instead
Temporal splits your job into two kinds of code.
- Workflows are the orchestration: the order of steps, the retry rules, the sleeps. Workflow code must be deterministic, because the server replays it from history to rebuild state.
- Activities are the actual work: run
restic, call an API, write a file. Activities can do anything and can fail, and Temporal retries them per your policy.
A worker is a process you run that executes both. The server holds the history and hands tasks to workers. When a worker dies, tasks wait on the server until a worker comes back. Nothing is lost, because the server, not the worker, holds the state.
That gives you four things a bash script has to fake:
- Resume after a crash. Kill the worker mid-run, restart it, and the workflow continues from the last completed step.
- Retry policies as data. Backoff, max attempts, and per-step timeouts are arguments, not shell loops.
- Durable timers.
workflow.sleepfor six hours survives worker restarts and server restarts. - A history you can read. The web UI shows every step, every attempt, every error message, with timestamps.
Running Temporal on One Box
There are two honest ways to run it at home.
Option 1: the dev server (start here)
The Temporal CLI bundles a whole server. One command, no containers.
temporal server start-dev --db-filename ./temporal.dbThe frontend listens on port 7233 and the web UI on 8233. The --db-filename flag matters: without it, workflow history lives in memory and disappears when the process stops. With it, history lives in a SQLite file. Temporal labels this a development server, so treat it as a way to learn and to run low-stakes jobs, not as something to build a company on. For a home lab, plenty of people will never need more.
Option 2: Postgres, server, and UI in Compose
If you want the real server topology, use containers. The old temporalio/docker-compose repo is archived (read-only since January 2026). Its README points to the samples-server repo, which holds the current compose files. Those files are labeled for local development and testing, and Temporal points production users at Helm charts on Kubernetes. On a single home lab box, a trimmed Postgres-only stack is a reasonable middle ground.
The file below is that trimmed stack, modeled on the samples-server Postgres file. I pinned versions current as of late September 2026: server 1.32.0, UI 2.54.1.
# Self-hosted Temporal for one box: Postgres + server + web UI.# Modeled on temporalio/samples-server (compose/docker-compose-postgres.yml).services: postgres: image: postgres:16 environment: POSTGRES_USER: temporal POSTGRES_PASSWORD: temporal volumes: - pgdata:/var/lib/postgresql/data healthcheck: test: ["CMD-SHELL", "pg_isready -U temporal"] interval: 5s timeout: 5s retries: 30
# One-shot: create both databases and apply the schemas, then exit. schema-setup: image: temporalio/admin-tools:1.32.0 depends_on: postgres: condition: service_healthy entrypoint: ["/bin/sh", "-c"] command: - | set -e T="temporal-sql-tool --plugin postgres12 --ep postgres -u temporal -p 5432" export SQL_PASSWORD=temporal for db in temporal temporal_visibility; do $$T --db $$db create $$T --db $$db setup-schema -v 0.0 done $$T --db temporal update-schema -d /etc/temporal/schema/postgresql/v12/temporal/versioned $$T --db temporal_visibility update-schema -d /etc/temporal/schema/postgresql/v12/visibility/versioned restart: "no"
temporal: image: temporalio/server:1.32.0 depends_on: schema-setup: condition: service_completed_successfully environment: DB: postgres12 DB_PORT: "5432" POSTGRES_USER: temporal POSTGRES_PWD: temporal POSTGRES_SEEDS: postgres BIND_ON_IP: 0.0.0.0 DYNAMIC_CONFIG_FILE_PATH: config/dynamicconfig/development-sql.yaml ports: - "127.0.0.1:7233:7233" volumes: - ./dynamicconfig:/etc/temporal/config/dynamicconfig healthcheck: test: ["CMD", "nc", "-z", "localhost", "7233"] interval: 5s timeout: 3s start_period: 20s retries: 30 restart: unless-stopped
# One-shot: the server image does not create the "default" namespace itself. create-namespace: image: temporalio/admin-tools:1.32.0 depends_on: temporal: condition: service_healthy environment: TEMPORAL_ADDRESS: temporal:7233 entrypoint: ["/bin/sh", "-c"] command: - temporal operator namespace describe -n default || temporal operator namespace create -n default restart: "no"
ui: image: temporalio/ui:2.54.1 depends_on: temporal: condition: service_healthy environment: TEMPORAL_ADDRESS: temporal:7233 ports: - "127.0.0.1:8080:8080" restart: unless-stopped
volumes: pgdata:A few things to know about this file:
schema-setupis a one-shot container that creates both databases and applies the schema withtemporal-sql-tool. It exits when done.create-namespaceexists because the plain server image does not create thedefaultnamespace for you.- Ports are bound to
127.0.0.1. Put a reverse proxy with auth in front of the UI before you expose it, because the UI has no login in this setup. - The server needs a dynamic config file, or it exits at startup. This one is tiny:
limit.maxIDLength: - value: 255 constraints: {}Bring it up and wait for the healthy status.
docker compose up -ddocker compose ps -aThe UI is at http://localhost:8080. The server’s gRPC frontend is on localhost:7233.
The Nightly Backup as a Workflow
Now the five steps. Activities first. They are safe fakes: they write under /tmp and sleep. The upload activity fails on its first two attempts on purpose, so you can watch retries without breaking a network.
import asyncioimport gzipimport shutilfrom pathlib import Path
from temporalio import activity
WORK = Path("/tmp/temporal-homelab")
@activity.defnasync def snapshot(target: str) -> str: """Pretend to snapshot a dataset by writing a file.""" WORK.mkdir(exist_ok=True) path = WORK / f"{target}.snap" path.write_text("pretend this is 500 GB of photos\n" * 1000) await asyncio.sleep(2) activity.logger.info("snapshot done: %s", path) return str(path)
@activity.defnasync def compress(snap: str) -> str: out = snap + ".gz" with open(snap, "rb") as src, gzip.open(out, "wb") as dst: shutil.copyfileobj(src, dst) await asyncio.sleep(2) return out
@activity.defnasync def upload(archive: str) -> str: """Fake upload. Fails on attempts 1 and 2 to show retries.""" attempt = activity.info().attempt await asyncio.sleep(1) if attempt < 3: raise ConnectionError(f"remote host unreachable (attempt {attempt})") remote = WORK / "remote" / Path(archive).name remote.parent.mkdir(exist_ok=True) shutil.copy(archive, remote) return str(remote)
@activity.defnasync def verify(remote: str) -> None: with gzip.open(remote, "rb") as f: if not f.read(16).startswith(b"pretend"): raise ValueError("archive content is wrong")
@activity.defnasync def notify(message: str) -> None: print(f"NOTIFY: {message}", flush=True)activity.info().attempt is the current attempt number, starting at 1. In real life you would replace the body of each activity with a restic call, an rclone copy, or a curl to your notification service.
The workflow wires them together.
from dataclasses import dataclassfrom datetime import timedelta
from temporalio import workflowfrom temporalio.common import RetryPolicy
with workflow.unsafe.imports_passed_through(): from activities import compress, notify, snapshot, upload, verify
@dataclassclass BackupInput: target: str = "photos" settle_seconds: int = 10
@workflow.defnclass NightlyBackup: @workflow.run async def run(self, inp: BackupInput) -> str: snap = await workflow.execute_activity( snapshot, inp.target, start_to_close_timeout=timedelta(minutes=5) ) archive = await workflow.execute_activity( compress, snap, start_to_close_timeout=timedelta(minutes=30) ) # Flaky network step: retry with backoff, give up after 5 tries. remote = await workflow.execute_activity( upload, archive, start_to_close_timeout=timedelta(minutes=10), retry_policy=RetryPolicy( initial_interval=timedelta(seconds=2), backoff_coefficient=2.0, maximum_attempts=5, ), ) # Durable timer: survives worker restarts and server restarts. await workflow.sleep(timedelta(seconds=inp.settle_seconds)) await workflow.execute_activity( verify, remote, start_to_close_timeout=timedelta(minutes=10) ) msg = f"backup of {inp.target} verified at {remote}" await workflow.execute_activity( notify, msg, start_to_close_timeout=timedelta(seconds=30) ) return msgRead that top to bottom. It looks like a bash script with better manners, and that is the appeal. Three details are worth knowing.
workflow.unsafe.imports_passed_through() tells the Python SDK’s sandbox not to reload your activity module inside the workflow. Without it you get slow startup and odd import errors. Activities are imported here only as references.
Every step has a start_to_close_timeout. Temporal requires you to set either that or schedule_to_close_timeout for each activity. It caps how long one attempt may run. Set it to something generous: a compress step on a big dataset can take an hour.
The retry policy on upload says: first retry after 2 seconds, double the wait each time, give up after 5 attempts. Activities that don’t set a policy get the default, which retries forever with backoff. That default is friendly for flaky APIs and dangerous for a step that will never succeed, so set maximum_attempts when a step has a real failure mode.
The sleep between upload and verify is a durable timer. Ten seconds is for the demo. In real life it might be an hour, to let a remote object store settle, or a day, to wait for a cold-storage restore test. The timer lives on the server, so a worker restart does not reset it.
The worker and the starter
The worker registers the workflow and activities on a task queue and waits.
import asyncio
from temporalio.client import Clientfrom temporalio.worker import Worker
from activities import compress, notify, snapshot, upload, verifyfrom workflows import NightlyBackup
TASK_QUEUE = "homelab"
async def main() -> None: client = await Client.connect("localhost:7233") worker = Worker( client, task_queue=TASK_QUEUE, workflows=[NightlyBackup], activities=[snapshot, compress, upload, verify, notify], ) print("worker up, waiting for tasks") await worker.run()
if __name__ == "__main__": asyncio.run(main())The starter kicks off one run and waits for the result.
import asyncioimport sysimport time
from temporalio.client import Client
from workflows import BackupInput, NightlyBackup
async def main() -> None: target = sys.argv[1] if len(sys.argv) > 1 else "photos" client = await Client.connect("localhost:7233") result = await client.execute_workflow( NightlyBackup.run, BackupInput(target=target), id=f"nightly-backup-{target}-{int(time.time())}", task_queue="homelab", ) print(f"result: {result}")
if __name__ == "__main__": asyncio.run(main())The pyproject.toml pins the SDK. Version 1.33.0 needs Python 3.10 or newer.
[project]name = "temporal-homelab-workflows"version = "0.1.0"requires-python = ">=3.10"dependencies = ["temporalio==1.33.0"]Watching It Survive a Crash
Run the worker in one terminal and the starter in another.
uv run python worker.pyuv run python starter.py photosThe upload fails twice, the retries back off, and the third attempt succeeds. The starter prints:
result: backup of photos verified at /tmp/temporal-homelab/remote/photos.snap.gzNow the fun part. Start another run, and press Ctrl+C on the worker while the upload is retrying. The workflow shows as Running in the UI, and it just waits. Start the worker again. In my test the worker was down for about 45 seconds, and the run finished right after the restart, with steps 1 and 2 not repeated. The event history shows the timer starting and firing ten seconds later, then verify and notify completing. Compare that to your bash script, which would have been dead at the first exit 1.
That resume is why you run a workflow engine. Everything else is convenience.
Replacing the Crontab Line
Temporal has its own scheduler. A Schedule starts a workflow on a cron expression, and the server keeps it running whether or not your worker is up.
temporal schedule create \ --schedule-id nightly-photos \ --cron "0 2 * * *" \ --workflow-id nightly-backup-photos \ --type NightlyBackup \ --task-queue homelab \ --input '{"target":"photos","settle_seconds":10}'The cron expression is evaluated in UTC unless you set a timezone. The --input JSON maps to the BackupInput dataclass fields. If the worker is down at 2 AM, the workflow task waits on the queue, and the run happens when the worker returns.
When Temporal Earns Its Keep
Use it when the job matches at least two of these:
- Three or more steps where a half-finished run leaves a mess (backup, then prune, then verify).
- Long waits between steps: hours or days, such as “restore test the backup 24 hours later.”
- Flaky dependencies that need different retry rules per step.
- You want to see what happened without grepping logs, including which attempt failed and why.
- Several jobs share the same shape, so one worker and one UI beat five bespoke scripts.
Good home lab candidates: nightly backup with a restore test, a media pipeline (download, transcode, tag, move, refresh library), a certificate rotation across several hosts, or a staged update of a Docker fleet with health checks between hosts.
When It Is a Forklift Moving a Couch
Skip it when:
- The job is one command.
restic backupon a systemd timer withOnFailure=and a Healthchecks ping is 15 lines and zero extra containers. - The job is idempotent and cheap to rerun. If starting over costs nothing, resume-from-step-3 buys you nothing.
- You don’t already run a database you trust, and you don’t want to babysit one more service. Compose gives you a Postgres, a server, and a UI to patch. That is four containers for something a timer does with zero.
- Nobody would notice a missed night. Temporal removes silent failure from the step level, but the worker process itself can still be down. You need monitoring on the worker either way.
Temporal replaces script complexity with infrastructure complexity. Below three steps, the script wins. Above that, the workflow code stays readable while the script turns into spaghetti, and the infrastructure cost is fixed.
There is also a learning cost. Determinism rules are real: no random, no datetime.now(), and no network calls inside workflow code. Put anything nondeterministic in an activity. The sandbox catches many mistakes, but you will hit one eventually, and your 2 AM self will thank you for reading the docs on that before it happens.
The SumGuy Take
Start with the systemd timer plus Healthchecks. It handles most home lab jobs without breaking a sweat. When you catch yourself writing marker files to remember which step you were on, or a retry loop with three nested if blocks, that is the signal. Run temporal server start-dev --db-filename ./temporal.db, port one job, and see whether the UI makes you feel smarter or just busier. If it helps, move to the Compose stack. If it doesn’t, you lost an evening and gained a better understanding of your own script.
Common Questions
Does Temporal need Kubernetes to self-host?
No. Temporal runs from a single CLI command (temporal server start-dev) or from Docker Compose with Postgres, server, and UI containers. Temporal’s own compose files are labeled for development and testing, and its production guidance points to Kubernetes with Helm charts. A single home lab box can use Compose fine.
Does Temporal lose workflow history if the dev server restarts?
Only if you skip the flag. temporal server start-dev keeps history in memory unless you pass --db-filename with a file path. With that flag, the dev server stores history in a SQLite file, so workflows survive restarts. The Compose stack in this article stores history in Postgres on a named volume.
How do I upgrade a self-hosted Temporal server?
Upgrade Temporal one minor version at a time, never skipping one. Back up the Postgres volume, bump the server, admin-tools, and UI tags to the next minor version, then run docker compose up -d. The example’s schema-setup container applies migrations with temporal-sql-tool before the new server starts. Read each version’s release notes, then repeat.
Which Python version does the Temporal SDK need?
Python 3.10 or newer. The temporalio package at version 1.33.0 declares that minimum on PyPI. The example project pins temporalio==1.33.0 and sets requires-python = ">=3.10" in its pyproject.toml, and uv picks a matching interpreter for you.