Server
Infrastructure-as-code and service definitions for luke-else.co.uk's self-hosted server estate: three Hetzner Cloud VPS instances provisioned with OpenTofu, each bootstrapped and deployed with Ansible, running a set of Docker Compose stacks behind Traefik.
Contents
- Architecture
- Repository layout
- Prerequisites
- Control panel (
control.sh) - Provisioning the infrastructure
- Bootstrapping and deploying with Ansible
- Service inventory
- First-time setup
- Development container
- Security notes
Architecture
Three servers, one shared private network:
architecture-beta
group cloud(cloud)[Hetzner]
group network(cloud)[network] in cloud
service dev(mdi:server)[dev] in network
service prod(mdi:server)[prod] in network
service prodfirewall(mdi:firewall)[firewall] in cloud
service vpn(mdi:server)[vpn] in cloud
service vpnfirewall(mdi:firewall)[firewall] in cloud
service gateway(mdi:web)[gateway] in cloud
service backups(mdi:bucket)[S3 Backups]
dev:L -- R:prod
prod:B -- T:prodfirewall
vpn:B -- T:vpnfirewall
prodfirewall: L -- R: gateway
vpnfirewall: B -- T: gateway
dev:T -- B:backups
prod:T -- B:backups
vpn:R -- L:backups
gateway represents the public internet, not a provisioned resource. backups is the S3-compatible bucket every host's backup-docker-compose.yml pushes to and restores from - not a Hetzner resource, see Persistent data lives in named Docker volumes, backed up to S3.
| Server | Purpose | Network |
|---|---|---|
dev |
Gitea, CI runner, dev-facing Traefik | Private network only (no public firewall exposure beyond CI/CD) |
prod |
Public-facing websites, Bitwarden, RustDesk, status page, prod Traefik | Private network + public firewall |
vpn |
WireGuard (wg-easy) + its own Traefik | Not attached to the private network — kept isolated so a compromised VPN endpoint can't pivot to dev/prod |
Each host's stateful data lives in local named Docker volumes, backed up nightly to a shared S3 bucket and restored automatically on first boot - see below. No host has a Hetzner Cloud Volume attached any more.
dev and prod share a private Hetzner network (10.0.1.0/24 by default) so CI/CD on dev can reach deployment targets on prod without exposing that traffic publicly. vpn is deliberately kept off this network. Each server has its own Hetzner Cloud Firewall (see infra/modules/<host>/main.tf) that only opens the ports actually used by the compose stacks running on it, plus SSH restricted to var.allowed_ssh_source_ips.
Repository layout
.
├── infra/ # OpenTofu (Terraform-compatible) config — provisions the 3 servers, network, firewalls, DNS
│ ├── main.tf # root module: wires network + dev/prod/vpn/dns modules together
│ ├── variables.tf # shared inputs (sizes, locations, IP ranges, SSH key names)
│ ├── outputs.tf # pass-through outputs from each host module
│ ├── ssh.tf # looks up each SSH key already uploaded to Hetzner Cloud
│ ├── versions.tf # provider requirements
│ ├── terraform.tfvars.example
│ └── modules/
│ ├── network/ # shared private network + subnet (used by dev and prod)
│ ├── dev/ # dev server + firewall, labeled role=dev
│ ├── prod/ # prod server + firewall, labeled role=prod
│ ├── vpn/ # vpn server + firewall (no private network), labeled role=vpn
│ └── dns/ # Hetzner DNS zones + A records for every domain in var.dns_zones
├── ansible/ # Ansible: bootstraps each server and deploys/starts services/<host>/ onto it
│ ├── inventory/hcloud.yml # dynamic inventory — queries the Hetzner API, groups by the role label above
│ ├── roles/
│ │ ├── bootstrap/ # Docker, deploy user, sshd hardening, unattended-upgrades
│ │ └── deploy/ # copies services/<host>/, renders .env (S3 backup credentials) + Runners/docker-compose.yml
│ └── playbooks/ # bootstrap.yml, deploy.yml, spinup.yml, spindown.yml, site.yml,
│ └── group_vars/ # deploy_user, service_group, backup_s3_*, etc. - must sit next to
│ # the playbooks for Ansible to auto-load it (see ansible/README.md)
├── services/ # Docker Compose stacks, grouped by which server they run on
│ ├── dev/ # Gitea + CI runner + Traefik + backup/restore (git.luke-else.co.uk, cicd.luke-else.co.uk)
│ ├── prod/ # Websites, Bitwarden, RustDesk, status page + Traefik + backup/restore
│ ├── vpn/ # WireGuard (wg-easy) + Traefik + backup/restore
│ └── todo.md # Outstanding manual setup/hardening tasks
├── docs/
│ └── architecture.md # Source of the architecture diagram above
├── .devcontainer/ # Git submodule: shared devcontainer for working on this repo (OpenTofu tooling)
└── assets/
Each of services/dev, services/prod, services/vpn follows the same convention: one *-docker-compose.yml (or subdirectory, e.g. Runners/) per logical service, a backup-docker-compose.yml + restore.sh pair for that host's S3 backup/restore (see below), plus a spinup.sh / spindown.sh pair that brings up or tears down every stack on that host in the right order. infra/modules/ and ansible/roles/deploy both mirror this same dev/prod/vpn split, so a given host's cloud resources, bootstrap/deploy logic, and compose stacks are easy to find side by side.
OpenTofu and Ansible have a clean split: OpenTofu only ever provisions cloud resources (servers, network, firewalls, DNS) and never touches anything over SSH. Everything from "the server exists" onward — installing Docker, creating the deploy user, hardening SSH, copying services/<host>/, and running spinup.sh/spindown.sh — is Ansible's job. See ansible/README.md for the full rundown.
Prerequisites
- A Hetzner Cloud project and API token
- One or more SSH keys uploaded to that project (Console → Security → SSH Keys) - all are installed on every server - plus the private key matching one of them available locally (Ansible uses it to bootstrap and deploy — see below)
- OpenTofu
>= 1.6.0 - Ansible
>= 2.15and thehetzner.hcloudcollection (ansible-galaxy collection install -r ansible/requirements.yml) - Ownership of the domains in
var.dns_zonesat whatever registrar they're bought through, so you can point their NS records at Hetzner (see Managing DNS — the zones and records themselves are created for you) - An S3-compatible bucket (e.g. Hetzner Object Storage, AWS S3, or any other S3-compatible provider) and an access key/secret pair with read/write access to it - not provisioned by OpenTofu (the
hcloudprovider has no Object Storage resource), so create this yourself and pass it to Ansible asBACKUP_S3_BUCKET/BACKUP_S3_ACCESS_KEY_ID/BACKUP_S3_SECRET_ACCESS_KEY(andBACKUP_S3_ENDPOINTfor anything other than real AWS S3) - see Persistent data lives in named Docker volumes, backed up to S3
Docker + the Compose plugin, the non-root deploy user, and SSH hardening no longer need doing by hand — Ansible's bootstrap role handles all of that (see below).
Control panel (control.sh)
If you'd rather not remember the exact command order, control.sh at the repo root is an interactive menu that steps you through the whole pre-setup and bring-up. Run it from anywhere:
./control.sh
It offers a Check prerequisites option (tools, HCLOUD_TOKEN/DEPLOY_USER/BACKUP_S3_* env vars, terraform.tfvars, SSH key, the hetzner.hcloud collection), a Configuration section, individual OpenTofu (init/plan/apply/output/destroy) and Ansible (bootstrap/deploy/spinup/spindown per role) actions, and a Guided full setup that runs everything in the firewall-imposed order — vpn first, pause for you to connect the VPN, then dev/prod. It's only a wrapper around the same tofu and ansible-playbook commands documented below, so you can always fall back to running them by hand.
The Configuration section covers every variable both tools need:
- Set variables prompts for
HCLOUD_TOKEN,DEPLOY_USER(the non-root user Ansible creates — defaults todeploy),ANSIBLE_SSH_PRIVATE_KEY_FILE, and theBACKUP_S3_*credentials, one at a time (leave blank to keep the current value, secrets are input-masked). They're saved to.control.envat the repo root — gitignored, never committed — which is loaded automatically the next time you runcontrol.sh, so you only enter them once. This is scoped tocontrol.shitself, not your regular shell. - Copy terraform.tfvars.example -> terraform.tfvars creates
infra/terraform.tfvarsfor you and prompts forssh_key_names(the one required field with no default) so you don't leave the example's placeholder key names in place by accident. Everything else in that file (location, server types,dns_zones, etc.) still has sensible defaults you can edit by hand afterwards - see Provisioning the infrastructure below.
Provisioning the infrastructure (infra/)
cd infra
export HCLOUD_TOKEN=your-hetzner-api-token # never commit this
cp terraform.tfvars.example terraform.tfvars
$EDITOR terraform.tfvars # set ssh_key_names at minimum
tofu init
tofu plan
tofu apply
This creates, via module.network / module.dev / module.prod / module.vpn / module.dns in infra/main.tf:
hcloud_network+ subnet, shared bydevandprod(modules/network)- one
hcloud_server+hcloud_firewallper host, scoped to the ports each host actually uses, each server labeledrole = "dev"/"prod"/"vpn"for Ansible's dynamic inventory (modules/dev,modules/prod,modules/vpn) - one Hetzner DNS zone per domain in
var.dns_zones, plus every A record in Service inventory (modules/dns— see Managing DNS)
No servers have an attached volume — persistent data lives in named Docker volumes backed up to S3 instead, see below.
Useful outputs: tofu output dev_ipv4, tofu output prod_ipv4, tofu output vpn_ipv4, tofu output dns_nameservers.
terraform.tfvars and any *.tfvars file are gitignored — never commit real values there. Defaults for server sizes, locations, and IP ranges live in infra/variables.tf and are passed down into the modules from infra/main.tf; override them per-environment via terraform.tfvars.
OpenTofu never connects to the servers over SSH — no provisioners, no bootstrap.sh, no copying services/ — that's all Ansible now. See below.
Managing DNS
module.dns (in infra/modules/dns) creates one Hetzner DNS zone per domain in var.dns_zones (default: luke-else.co.uk, divine-couture.co.uk, snexo.co.uk) and an hcloud_zone_rrset per record in var.records — kept in sync with local.dns_records (A records, one per hostname currently referenced by a Traefik Host() rule anywhere under services/, pointed at whichever of dev/prod/vpn actually serves it) and local.mail_records (see below) in infra/main.tf. Zones have prevent_destroy = true, matching the volumes — losing one deletes every record in it.
Creating the zone doesn't make Hetzner authoritative for the domain by itself: you still need to point that domain's NS records at Hetzner's nameservers at whichever registrar it's registered through. Run tofu output dns_nameservers after applying to get the exact nameservers per domain, and set those as the domain's NS records at the registrar. Propagation can take a while depending on the registrar and the domain's previous NS TTL.
Adding a new subdomain: add an entry to local.dns_records in infra/main.tf (zone, name, and the target module.<host>.ipv4) and re-run tofu apply — don't hand-create records in the Hetzner console, they'll drift from state. Note name = "@" is Hetzner's convention for a zone's apex record (e.g. bare snexo.co.uk), not an empty string.
Only A records for IPv4 are managed for hosts — none of the modules currently track servers' IPv6 addresses, so AAAA records aren't generated even though the firewalls already allow IPv6 traffic.
var.records isn't limited to A records: each entry takes an optional type (A by default, or MX/TXT/CNAME/etc.), and entries sharing the same zone/name/type are merged into a single Hetzner RRSet with multiple values (needed for e.g. two TXT strings under the same name) - see infra/modules/dns/variables.tf for the exact value format each type expects (MX embeds the priority in the value string, TXT values need escaped double quotes).
Mail records (Microsoft 365)
local.mail_records in infra/main.tf tracks the DNS records Microsoft 365 needs for luke-else.co.uk - the MX record, the SPF and Google-site-verification TXT records (merged into one @ TXT RRSet), and the autodiscover CNAME. These carried over from the domain's DNS provider prior to Hetzner and were never applied to the Hetzner zone, so they're genuinely new records here (not something to tofu import - nothing to collide with).
Not yet included: a DMARC policy (_dmarc TXT) and DKIM signing (two selector1._domainkey/selector2._domainkey CNAMEs). DKIM's targets are generated per-tenant by Microsoft 365 (Defender portal → Email & collaboration → Policies & rules → DKIM) and can't be guessed - enable DKIM there first, then add the two CNAMEs it gives you to local.mail_records.
infra/modules/dns/moved.tf re-points every pre-existing A record at its new resource address, since generalizing the module to support non-A record types changed how they're keyed internally - without it, tofu apply would destroy and recreate every A record for no functional reason. It's a one-time migration aid: run tofu plan first to confirm the only real changes are the new mail records being created, then the file can be deleted after a successful apply.
Bootstrapping and deploying with Ansible (ansible/)
Once tofu apply has created the servers, ansible/ takes over everything else: installing Docker, creating the deploy user, hardening SSH, copying services/<host>/ to each server, rendering .env (S3 backup credentials) and, for dev, Runners/docker-compose.yml, and running spinup.sh/spindown.sh. Full detail lives in ansible/README.md; the short version:
cd ansible
ansible-galaxy collection install -r requirements.yml
export HCLOUD_TOKEN=your-hetzner-api-token # never commit this
# vpn first - its firewall accepts SSH from var.allowed_ssh_source_ips directly
ansible-playbook playbooks/site.yml -l role_vpn
# SSH to vpn as deploy and connect to the WireGuard tunnel it just started, then:
ansible-playbook playbooks/site.yml -l role_dev,role_prod
Hosts are discovered dynamically from the Hetzner API (ansible/inventory/hcloud.yml), grouped into role_dev/role_prod/role_vpn by the role label OpenTofu sets on each server — there's no static inventory file to keep in sync, and nothing here reads Terraform state.
Persistent data lives in named Docker volumes, backed up to S3
Every stateful service's compose file mounts a plain named Docker volume rather than a bind mount to a Hetzner volume - gitea_data:/data, bitwarden_data:/data/, letsencrypt_data:/letsencrypt, and so on across services/dev/*.yml, services/prod/*.yml, and services/vpn/*.yml. Each volume is declared with an explicit name: in every compose file that mounts it (the same pattern already used for the shared proxy network), so Compose creates it once and every stack that references it reuses the same underlying volume regardless of which compose file happens to run first.
Each host's backup-docker-compose.yml defines two services against those same named volumes:
backup—offen/docker-volume-backup, running on a nightly cron (0 3 * * *), tars up every mounted volume and uploads it tos3://$AWS_S3_BUCKET_NAME/$AWS_S3_PATH/<host>-<timestamp>.tar.gz, keeping 14 days of history (BACKUP_RETENTION_DAYS). Containers labeleddocker-volume-backup.stop-during-backup=true(Gitea, Bitwarden, RustDesk, Uptime Kuma, snexo, wg-easy) are stopped for the duration of the backup so their data is captured in a consistent state, then restarted automatically. Started byspinup.shalongside everything else.restore— a one-shot container (amazon/aws-clirunningrestore.sh) thatspinup.shruns before anything else starts. It checks each of the host's named volumes; any that are already empty are populated from the most recent object under that host's S3 prefix, any that already contain data are left completely untouched. On a fresh server with no prior backup (or a bucket that doesn't exist yet), it's a no-op - services just start with empty volumes as before.
AWS_S3_BUCKET_NAME, AWS_S3_PATH (set to the host name, e.g. dev), AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and AWS_ENDPOINT are rendered into services/<host>/.env by Ansible's deploy role from BACKUP_S3_BUCKET/BACKUP_S3_ACCESS_KEY_ID/BACKUP_S3_SECRET_ACCESS_KEY/BACKUP_S3_ENDPOINT in your shell environment (see Prerequisites) — docker compose auto-loads .env from its working directory. spinup.sh on every host refuses to start anything if .env is missing.
The Gitea Actions runners are the one exception: their /data is a disposable local cache (runner identity re-registers with Gitea on every spinup.sh run anyway), so it's a plain named volume per runner (gitea_runner_N_data, baked directly into the rendered Runners/docker-compose.yml) that's deliberately left out of the backup/restore stack.
To trigger a backup or restore manually rather than waiting for cron or a restart:
cd services/<host>
docker compose -f backup-docker-compose.yml exec backup backup # force an immediate backup
docker compose -f backup-docker-compose.yml run --rm restore # restore any currently-empty volumes
Gitea Actions runners
dev always runs three Gitea Actions runner containers, configured automatically on deploy. Re-running ansible-playbook playbooks/deploy.yml -l role_dev renders roles/deploy/templates/runners-docker-compose.yml.j2 straight onto the server as services/dev/Runners/docker-compose.yml — three runner-N services, each with its own container name and /data volume so their registrations don't collide. That file is generated on the remote host, so edit the template rather than hand-editing the rendered file (the count lives in a {% raw %}{% set runner_count = 3 %}{% endraw %} line at the top).
Rendering it doesn't start anything by itself — re-run ansible-playbook playbooks/spinup.yml -l role_dev afterwards to apply a template change.
Registration tokens are handled automatically, not baked into the generated file: services/dev/spinup.sh waits for Gitea to come up, runs gitea actions generate-runner-token inside the Gitea container, and writes the result to Runners/.env, which Compose loads automatically. There's no manual admin-UI step for this anymore.
First deploy: bootstrapping order matters
dev and prod's firewalls only accept SSH from vpn's public IP (see Security notes), but vpn's own WireGuard service isn't running until you deploy it — so on a from-scratch estate, Ansible can't reach dev/prod yet. Bring it up in this order:
ansible-playbook playbooks/site.yml -l role_vpn— bootstraps, deploys, and startsvpnonly; its firewall allows SSH fromvar.allowed_ssh_source_ipsdirectly.- Visit
https://vpn.luke-else.co.ukand complete wg-easy's first-run setup wizard (admin username/password, then add a client peer and download its config/QR code) - see Setting up WireGuard (wg-easy). Import that config into a WireGuard client and connect. ansible-playbook playbooks/site.yml -l role_dev,role_prod— now that your machine is tunneled throughvpn, its NATed egress IP matches the firewall rule anddev/prodbecome reachable.
If step 3 is run before you're connected to the VPN, Ansible will simply fail to connect — re-run it once connected.
Deploying the services (services/)
After ansible-playbook playbooks/deploy.yml has copied services/<host>/ to /home/deploy/services/<host> on the matching server, SSH in as deploy and run the matching script from inside that directory — or just use ansible-playbook playbooks/spinup.yml / spindown.yml, which do exactly this remotely (see ansible/README.md):
# on dev
./spinup.sh # Traefik → Gitea → (waits, generates a runner token) → Runners → Watchtower
./spindown.sh # reverse order, then prunes images/volumes
# on prod
./spinup.sh # Traefik, → Watchtower → status → websites → Bitwarden → RustDesk
./spindown.sh
# on vpn
./spinup.sh # Traefik → wg-easy → Watchtower
./spindown.sh
Setting up WireGuard (wg-easy)
vpn runs wg-easy - WireGuard plus a web UI for managing peers - fronted by Traefik like everything else. spinup.sh only starts the container; wg-easy itself has no automated first-run setup here (see services/vpn/vpn-docker-compose.yml), so once it's up:
- Visit
https://vpn.luke-else.co.ukand complete the setup wizard - pick an admin username and a strong password (this account can view/rotate every client's private key, and it's reachable from the public internet). Enabling 2FA afterwards (wg-easy supports it) is worth doing since this isn't just a status page like the old OpenVPN setup was. - Set the host clients should connect to (
vpn.luke-else.co.uk) and the port (51820, matching the firewall rule ininfra/modules/vpn/main.tf) when prompted. - Add a client per device from the UI, then download its config file or scan the QR code directly into the WireGuard app.
wg-easy's own state (admin account, peers, keys) lives in the wg_easy_data named volume, backed up to S3 like everything else - see Persistent data lives in named Docker volumes, backed up to S3. Client configs themselves are only ever downloaded through the UI, not stored anywhere else.
If you ever need to force a re-sync without going through Ansible (e.g. testing a local edit before committing it), a manual scp -r services/prod deploy@<prod_ipv4>:~/services still works fine - ansible-playbook playbooks/deploy.yml will just overwrite it again next time it's run.
Each stack can also be managed individually with plain Compose, e.g.:
cd services/prod
docker compose -f bitwarden-docker-compose.yml up -d
docker compose -f bitwarden-docker-compose.yml down
All three hosts run Watchtower polling every 60s with cleanup enabled, so images are kept current automatically once deployed — the spinup scripts only need to be re-run after adding/removing a service or changing compose files.
Every public-facing service is fronted by its host's own Traefik instance, terminating TLS via Let's Encrypt (tlschallenge, port 80/443). Each stack joins an external proxy Docker network and opts in via traefik.enable=true labels rather than publishing ports directly (RustDesk and the Gitea SSH port are the deliberate exceptions, since they aren't HTTP).
Service inventory
dev
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.cicd.luke-else.co.uk |
| Gitea | gitea-docker-compose.yml |
git.luke-else.co.uk (HTTP), SSH on 222 |
| Gitea Actions runner(s) | Runners/docker-compose.yml (generated — see Scaling Gitea Actions runners) |
N/A |
| Watchtower | watchtower-docker-compose.yml |
— |
| Backup/restore | backup-docker-compose.yml |
— |
prod
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.luke-else.co.uk |
| Websites | web-docker-compose.yml |
luke-else.co.uk, dev.luke-else.co.uk, metarius.luke-else.co.uk, www.divine-couture.co.uk, snexo.co.uk |
| Status page (Uptime Kuma) | status-docker-compose.yml |
status.luke-else.co.uk |
| Bitwarden (Vaultwarden) | bitwarden-docker-compose.yml |
bitwarden.luke-else.co.uk |
| RustDesk relay (hbbs/hbbr) | rd-docker-compose.yml |
rd.luke-else.co.uk, ports 21115-21119 |
| Watchtower | watchtower-docker-compose.yml |
— |
| Backup/restore | backup-docker-compose.yml |
— |
vpn
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.vpn.luke-else.co.uk |
| WireGuard (wg-easy) | vpn-docker-compose.yml |
vpn.luke-else.co.uk, UDP 51820 |
| Watchtower | watchtower-docker-compose.yml |
— |
| Backup/restore | backup-docker-compose.yml |
— |
First-time setup
A few things need manual attention before a stack is fully live — tracked in services/todo.md, summarized here:
- General host hardening: non-root user, Docker, and unattended-upgrades are now handled automatically by Ansible's
bootstraprole; UFW is the remaining manual item inservices/todo.md(the Hetzner Cloud Firewalls already allowlist per-host ports — see Security notes).
Development container
.devcontainer is a git submodule providing a ready-to-use OpenTofu development environment (VS Code + OpenTofu/Docker/Mermaid extensions). After cloning:
git submodule update --init --recursive
Then reopen the repo in VS Code with the Dev Containers extension.
Security notes
- Real secrets (
HCLOUD_TOKEN,BACKUP_S3_ACCESS_KEY_ID/BACKUP_S3_SECRET_ACCESS_KEY,*.tfvars, your SSH private key) must never be committed — see.gitignore. The S3 credentials end up in each host'sservices/<host>/.env(rendered by Ansible, mode0600, gitignored) - anyone with root on a host can read them, same trust boundary as everything elsedeploycan already do there. - SSH to
devandprodis restricted to thevpnserver's own public IP — you must be tunneled into the VPN to reach them over SSH. SSH tovpnitself is gated byvar.allowed_ssh_source_ips; narrow this from the default0.0.0.0/0once you know your own egress IP(s). vpnis intentionally excluded from the private network so that a compromised VPN endpoint cannot reachdevorproddirectly over it — SSH access still works because the VPN's egress traffic is NATed through its own public IP.- Firewalls are allowlists scoped per host in
infra/modules/<host>/main.tf— only ports actually used by that host's compose stacks (plus the cross-host private network range) are open. - Every host's Ansible bootstrap role disables SSH password authentication, restricts root login to key-only, and creates a separate sudo user (
deploy_user, seeansible/playbooks/group_vars/all.yml) for day-to-day access. - DNS zones (
modules/dns) carry both Hetzner's owndelete_protectionand Terraform'sprevent_destroy— losing a zone takes every record in it with it, including fordivine-couture.co.ukandsnexo.co.uk, not justluke-else.co.uk.
