20 KiB
Server
Infrastructure-as-code and service definitions for luke-else.co.uk's self-hosted server estate: three Hetzner Cloud VPS instances provisioned with OpenTofu, each bootstrapped and deployed with Ansible, running a set of Docker Compose stacks behind Traefik.
Contents
- Architecture
- Repository layout
- Prerequisites
- Provisioning the infrastructure
- Bootstrapping and deploying with Ansible
- Service inventory
- First-time setup
- Development container
- Security notes
Architecture
Three servers, one shared private network:
architecture-beta
group cloud(cloud)[Hetzner]
group network(cloud)[network] in cloud
service disk1(mdi:disk)[Storage] in cloud
service disk2(mdi:disk)[Storage] in cloud
service dev(mdi:server)[dev] in network
service prod(mdi:server)[prod] in network
service prodfirewall(mdi:firewall)[firewall] in cloud
service vpn(mdi:server)[vpn] in cloud
service vpnfirewall(mdi:firewall)[firewall] in cloud
service gateway(mdi:web)[gateway] in cloud
dev:L -- R:prod
disk1:B -- T:prod
disk2:B -- T:dev
prod:B -- T:prodfirewall
vpn:B -- T:vpnfirewall
prodfirewall: L -- R: gateway
vpnfirewall: B -- T: gateway
gateway represents the public internet, not a provisioned resource.
| Server | Purpose | Network | Volume |
|---|---|---|---|
dev |
Gitea, CI runner, dev-facing Traefik | Private network only (no public firewall exposure beyond CI/CD) | dev-storage |
prod |
Public-facing websites, Bitwarden, RustDesk, status page, prod Traefik | Private network + public firewall | prod-storage |
vpn |
OpenVPN + its own Traefik | Not attached to the private network — kept isolated so a compromised VPN endpoint can't pivot to dev/prod |
none |
dev and prod share a private Hetzner network (10.0.1.0/24 by default) so CI/CD on dev can reach deployment targets on prod without exposing that traffic publicly. vpn is deliberately kept off this network. Each server has its own Hetzner Cloud Firewall (see infra/modules/<host>/main.tf) that only opens the ports actually used by the compose stacks running on it, plus SSH restricted to var.allowed_ssh_source_ips.
Repository layout
.
├── infra/ # OpenTofu (Terraform-compatible) config — provisions the 3 servers, network, firewalls, volumes, DNS
│ ├── main.tf # root module: wires network + dev/prod/vpn/dns modules together
│ ├── variables.tf # shared inputs (sizes, locations, IP ranges, SSH key names)
│ ├── outputs.tf # pass-through outputs from each host module
│ ├── ssh.tf # looks up each SSH key already uploaded to Hetzner Cloud
│ ├── versions.tf # provider requirements
│ ├── terraform.tfvars.example
│ └── modules/
│ ├── network/ # shared private network + subnet (used by dev and prod)
│ ├── dev/ # dev server + firewall + volume, labeled role=dev
│ ├── prod/ # prod server + firewall + volume, labeled role=prod
│ ├── vpn/ # vpn server + firewall (no private network, no volume), labeled role=vpn
│ └── dns/ # Hetzner DNS zones + A records for every domain in var.dns_zones
├── ansible/ # Ansible: bootstraps each server and deploys/starts services/<host>/ onto it
│ ├── inventory/hcloud.yml # dynamic inventory — queries the Hetzner API, groups by the role label above
│ ├── group_vars/ # deploy_user, dev_runner_count, etc.
│ ├── roles/
│ │ ├── bootstrap/ # Docker, deploy user, sshd hardening, unattended-upgrades
│ │ └── deploy/ # copies services/<host>/, renders .env + Runners/docker-compose.yml
│ └── playbooks/ # bootstrap.yml, deploy.yml, spinup.yml, spindown.yml, site.yml
├── services/ # Docker Compose stacks, grouped by which server they run on
│ ├── dev/ # Gitea + CI runner + Traefik (git.luke-else.co.uk, cicd.luke-else.co.uk)
│ ├── prod/ # Websites, Bitwarden, RustDesk, status page + Traefik
│ ├── vpn/ # OpenVPN + Traefik
│ └── todo.md # Outstanding manual setup/hardening tasks
├── docs/
│ └── architecture.md # Source of the architecture diagram above
├── .devcontainer/ # Git submodule: shared devcontainer for working on this repo (OpenTofu tooling)
└── assets/
Each of services/dev, services/prod, services/vpn follows the same convention: one *-docker-compose.yml (or subdirectory, e.g. Runners/) per logical service, plus a spinup.sh / spindown.sh pair that brings up or tears down every stack on that host in the right order. infra/modules/ and ansible/roles/deploy both mirror this same dev/prod/vpn split, so a given host's cloud resources, bootstrap/deploy logic, and compose stacks are easy to find side by side.
OpenTofu and Ansible have a clean split: OpenTofu only ever provisions cloud resources (servers, network, firewalls, volumes, DNS) and never touches anything over SSH. Everything from "the server exists" onward — installing Docker, creating the deploy user, hardening SSH, copying services/<host>/, and running spinup.sh/spindown.sh — is Ansible's job. See ansible/README.md for the full rundown.
Prerequisites
- A Hetzner Cloud project and API token
- One or more SSH keys uploaded to that project (Console → Security → SSH Keys) - all are installed on every server - plus the private key matching one of them available locally (Ansible uses it to bootstrap and deploy — see below)
- OpenTofu
>= 1.6.0 - Ansible
>= 2.15and thehetzner.hcloudcollection (ansible-galaxy collection install -r ansible/requirements.yml) - Ownership of the domains in
var.dns_zonesat whatever registrar they're bought through, so you can point their NS records at Hetzner (see Managing DNS — the zones and records themselves are created for you)
Docker + the Compose plugin, the non-root deploy user, and SSH hardening no longer need doing by hand — Ansible's bootstrap role handles all of that (see below).
Provisioning the infrastructure (infra/)
cd infra
export HCLOUD_TOKEN=your-hetzner-api-token # never commit this
cp terraform.tfvars.example terraform.tfvars
$EDITOR terraform.tfvars # set ssh_key_names at minimum
tofu init
tofu plan
tofu apply
This creates, via module.network / module.dev / module.prod / module.vpn / module.dns in infra/main.tf:
hcloud_network+ subnet, shared bydevandprod(modules/network)- one
hcloud_server+hcloud_firewallper host, scoped to the ports each host actually uses, each server labeledrole = "dev"/"prod"/"vpn"for Ansible's dynamic inventory (modules/dev,modules/prod,modules/vpn) - one
hcloud_volumeeach fordevandprod(modules/dev,modules/prod;vpnhas none) - one Hetzner DNS zone per domain in
var.dns_zones, plus every A record in Service inventory (modules/dns— see Managing DNS)
Useful outputs: tofu output dev_ipv4, tofu output prod_ipv4, tofu output vpn_ipv4, tofu output dns_nameservers, tofu output dev_data_dir, tofu output prod_data_dir (the last two are informational only now — Ansible looks the data directory up itself, see below).
terraform.tfvars and any *.tfvars file are gitignored — never commit real values there. Defaults for server sizes, locations, and IP ranges live in infra/variables.tf and are passed down into the modules from infra/main.tf; override them per-environment via terraform.tfvars.
OpenTofu never connects to the servers over SSH — no provisioners, no bootstrap.sh, no copying services/ — that's all Ansible now. See below.
Managing DNS
module.dns (in infra/modules/dns) creates one Hetzner DNS zone per domain in var.dns_zones (default: luke-else.co.uk, divine-couture.co.uk, snexo.co.uk) and an hcloud_zone_rrset A record for every hostname currently referenced by a Traefik Host() rule anywhere under services/ — kept in sync with local.dns_records in infra/main.tf, pointed at whichever of dev/prod/vpn actually serves it. Zones have prevent_destroy = true, matching the volumes — losing one deletes every record in it.
Creating the zone doesn't make Hetzner authoritative for the domain by itself: you still need to point that domain's NS records at Hetzner's nameservers at whichever registrar it's registered through. Run tofu output dns_nameservers after applying to get the exact nameservers per domain, and set those as the domain's NS records at the registrar. Propagation can take a while depending on the registrar and the domain's previous NS TTL.
Adding a new subdomain: add an entry to local.dns_records in infra/main.tf (zone, name, and the target module.<host>.ipv4) and re-run tofu apply — don't hand-create records in the Hetzner console, they'll drift from state. Note name = "@" is Hetzner's convention for a zone's apex record (e.g. bare snexo.co.uk), not an empty string.
Only A records for IPv4 are managed here — none of the modules currently track servers' IPv6 addresses, so AAAA records aren't generated even though the firewalls already allow IPv6 traffic.
Bootstrapping and deploying with Ansible (ansible/)
Once tofu apply has created the servers, ansible/ takes over everything else: installing Docker, creating the deploy user, hardening SSH, copying services/<host>/ to each server, rendering the files OpenTofu used to generate (.env's DATA_DIR, dev's Runners/docker-compose.yml), and running spinup.sh/spindown.sh. Full detail lives in ansible/README.md; the short version:
cd ansible
ansible-galaxy collection install -r requirements.yml
export HCLOUD_TOKEN=your-hetzner-api-token # never commit this
# vpn first - its firewall accepts SSH from var.allowed_ssh_source_ips directly
ansible-playbook playbooks/site.yml -l role_vpn
# SSH to vpn as deploy and connect to the OpenVPN it just started, then:
ansible-playbook playbooks/site.yml -l role_dev,role_prod
Hosts are discovered dynamically from the Hetzner API (ansible/inventory/hcloud.yml), grouped into role_dev/role_prod/role_vpn by the role label OpenTofu sets on each server — there's no static inventory file to keep in sync, and nothing here reads Terraform state.
Persistent data lives on the volumes, not the server disk
dev and prod's compose files bind-mount container data to paths on the dev-storage / prod-storage volumes rather than the server's local disk, so it survives the server being destroyed and recreated: ${DATA_DIR}/gitea:/data, ${DATA_DIR}/bitwarden/:/data/, and so on across services/dev/*.yml and services/prod/*.yml. vpn has no volume, so it's untouched - see the architecture table above.
DATA_DIR resolves deterministically: Hetzner always automounts a volume at /mnt/HC_Volume_<volume-id>. Ansible's deploy role looks the volume up by name (hcloud_volume_info for dev-storage/prod-storage) directly against the Hetzner API and renders that path into services/<host>/.env as DATA_DIR=... on the remote host — docker compose auto-loads .env from its working directory for variable substitution. Nothing is written locally; re-run ansible-playbook playbooks/deploy.yml if a volume is ever destroyed and recreated (new id). spinup.sh on both hosts refuses to start anything if .env is missing, rather than silently falling back to a relative path.
The Gitea Actions runners are the one exception to the .env mechanism: their /data path is baked directly into the rendered Runners/docker-compose.yml, because Runners/.env is already reserved for the registration token and gets overwritten on every deploy.
Scaling Gitea Actions runners
The number of Gitea Actions runner containers on dev is set by dev_runner_count in ansible/group_vars/role_dev.yml (default 3). Re-running ansible-playbook playbooks/deploy.yml -l role_dev renders roles/deploy/templates/runners-docker-compose.yml.j2 straight onto the server as services/dev/Runners/docker-compose.yml — one runner-N service per count, each with its own container name and /data volume so their registrations don't collide. That file is generated on the remote host: change dev_runner_count and redeploy rather than hand-editing it.
Rendering it doesn't start anything by itself — re-run ansible-playbook playbooks/spinup.yml -l role_dev afterwards to apply a count change.
Registration tokens are handled automatically, not baked into the generated file: services/dev/spinup.sh waits for Gitea to come up, runs gitea actions generate-runner-token inside the Gitea container, and writes the result to Runners/.env, which Compose loads automatically. There's no manual admin-UI step for this anymore.
First deploy: bootstrapping order matters
dev and prod's firewalls only accept SSH from vpn's public IP (see Security notes), but vpn's own OpenVPN service isn't running until you deploy it — so on a from-scratch estate, Ansible can't reach dev/prod yet. Bring it up in this order:
ansible-playbook playbooks/site.yml -l role_vpn— bootstraps, deploys, and startsvpnonly; its firewall allows SSH fromvar.allowed_ssh_source_ipsdirectly.- SSH in as
deployand connect to the OpenVPN service you just started with an OpenVPN client. ansible-playbook playbooks/site.yml -l role_dev,role_prod— now that your machine is tunneled throughvpn, its NATed egress IP matches the firewall rule anddev/prodbecome reachable.
If step 3 is run before you're connected to the VPN, Ansible will simply fail to connect — re-run it once connected.
Deploying the services (services/)
After ansible-playbook playbooks/deploy.yml has copied services/<host>/ to /home/deploy/services/<host> on the matching server, SSH in as deploy and run the matching script from inside that directory — or just use ansible-playbook playbooks/spinup.yml / spindown.yml, which do exactly this remotely (see ansible/README.md):
# on dev
./spinup.sh # Traefik → Gitea → (waits, generates a runner token) → Runners → Watchtower
./spindown.sh # reverse order, then prunes images/volumes
# on prod
./spinup.sh # Traefik, → Watchtower → status → websites → Bitwarden → RustDesk
./spindown.sh
# on vpn
./spinup.sh # Traefik → OpenVPN → Watchtower
./spindown.sh
If you ever need to force a re-sync without going through Ansible (e.g. testing a local edit before committing it), a manual scp -r services/prod deploy@<prod_ipv4>:~/services still works fine - ansible-playbook playbooks/deploy.yml will just overwrite it again next time it's run.
Each stack can also be managed individually with plain Compose, e.g.:
cd services/prod
docker compose -f bitwarden-docker-compose.yml up -d
docker compose -f bitwarden-docker-compose.yml down
All three hosts run Watchtower polling every 60s with cleanup enabled, so images are kept current automatically once deployed — the spinup scripts only need to be re-run after adding/removing a service or changing compose files.
Every public-facing service is fronted by its host's own Traefik instance, terminating TLS via Let's Encrypt (tlschallenge, port 80/443). Each stack joins an external proxy Docker network and opts in via traefik.enable=true labels rather than publishing ports directly (RustDesk and the Gitea SSH port are the deliberate exceptions, since they aren't HTTP).
Service inventory
dev
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.cicd.luke-else.co.uk |
| Gitea | gitea-docker-compose.yml |
git.luke-else.co.uk (HTTP), SSH on 222 |
| Gitea Actions runner(s) | Runners/docker-compose.yml (generated — see Scaling Gitea Actions runners) |
N/A |
| Watchtower | watchtower-docker-compose.yml |
— |
prod
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.luke-else.co.uk |
| Websites | web-docker-compose.yml |
luke-else.co.uk, dev.luke-else.co.uk, metarius.luke-else.co.uk, www.divine-couture.co.uk, snexo.co.uk |
| Status page (Uptime Kuma) | status-docker-compose.yml |
status.luke-else.co.uk |
| Bitwarden (Vaultwarden) | bitwarden-docker-compose.yml |
bitwarden.luke-else.co.uk |
| RustDesk relay (hbbs/hbbr) | rd-docker-compose.yml |
rd.luke-else.co.uk, ports 21115-21119 |
| Watchtower | watchtower-docker-compose.yml |
— |
vpn
| Service | Compose file | Domain(s) |
|---|---|---|
| Traefik | traefik-docker-compose.yml |
traefik.vpn.luke-else.co.uk |
| OpenVPN (Dockovpn) | vpn-docker-compose.yml |
vpn.luke-else.co.uk, UDP 1194 |
| Watchtower | watchtower-docker-compose.yml |
— |
First-time setup
A few things need manual attention before a stack is fully live — tracked in services/todo.md, summarized here:
- General host hardening: non-root user, Docker, and unattended-upgrades are now handled automatically by Ansible's
bootstraprole; UFW is the remaining manual item inservices/todo.md(the Hetzner Cloud Firewalls already allowlist per-host ports — see Security notes).
Development container
.devcontainer is a git submodule providing a ready-to-use OpenTofu development environment (VS Code + OpenTofu/Docker/Mermaid extensions). After cloning:
git submodule update --init --recursive
Then reopen the repo in VS Code with the Dev Containers extension.
Security notes
- Real secrets (
HCLOUD_TOKEN,*.tfvars, your SSH private key) must never be committed — see.gitignore. - SSH to
devandprodis restricted to thevpnserver's own public IP — you must be tunneled into the VPN to reach them over SSH. SSH tovpnitself is gated byvar.allowed_ssh_source_ips; narrow this from the default0.0.0.0/0once you know your own egress IP(s). vpnis intentionally excluded from the private network so that a compromised VPN endpoint cannot reachdevorproddirectly over it — SSH access still works because the VPN's egress traffic is NATed through its own public IP.- Firewalls are allowlists scoped per host in
infra/modules/<host>/main.tf— only ports actually used by that host's compose stacks (plus the cross-host private network range) are open. - Every host's Ansible bootstrap role disables SSH password authentication, restricts root login to key-only, and creates a separate sudo user (
deploy_user, seeansible/group_vars/all.yml) for day-to-day access. - DNS zones (
modules/dns) carry both Hetzner's owndelete_protectionand Terraform'sprevent_destroy— losing a zone takes every record in it with it, including fordivine-couture.co.ukandsnexo.co.uk, not justluke-else.co.uk.
