# Deployment pattern — generic guide (tailnet-first VPS, everything-as-files)

**What this is:** James's preferred deployment pattern, written as a reusable guide for ANY project on ANY new server. Reader: future James, a future Claude session, or any devops person. It assumes: Ubuntu LTS server, a Cloudflare account (API tokens obtainable), a Tailscale account, SSH key access, a git repo per project on the dev machine. It does NOT cover installing Ubuntu itself.

**Proven twice:**

- **2026-08-19, ai-new-business** on a Germany VPS — four apps, tailnet-gated, host Postgres. That repo's `deploy/` folder is the living exemplar; `91_QUICK_KNOWLEDGE_BASE_INFO_HUB.md` documents that specific instance.
- **2026-08-21/22, james.in.th** on a second, INHERITED VPS — one Astro SSR container, ingress migrated out of a Nginx Proxy Manager GUI into a Caddyfile, and the first PUBLIC (Cloudflare-proxied) host under this pattern. `james-in-th/deploy/` + its `deploy/SERVER-SETUP.md` document that instance. Sections §2 and §5 exist because of it; day two added §5.5 (the firewall change that broke a consumer nobody had inventoried), the tailnet-addressing block, and five of §6's gotchas.

---

## 0. The six principles (why this pattern)

1. **Infrastructure is files in the project repo** (Dockerfile, compose, Caddyfile, deploy script). No PaaS GUI, no hand-configured server state. Reason: James's workforce is AI sessions — anything that is a file can be read, fixed, and redeployed by any session; anything in a GUI cannot. A dead server is rebuilt from the repo + one secrets file.
2. **Tailnet-first, public by explicit decision.** New deployments start reachable only through Tailscale: DNS points at the server's Tailscale IP (100.64.0.0/10 is not publicly routable) and the firewall blocks public traffic anyway. Real domains + real HTTPS certs still work (DNS-01 needs no public reachability). Going public later is one DNS change + one firewall rule — a James decision, never a default.
3. **Build on the server, not the laptop.** Dev machines are often arm64 (Apple Silicon); servers are x86. `git push` moves code; the server bakes Docker images. Dev work never touches Docker — local dev stays `next dev`/hot-reload on the laptop.
4. **Postgres native on the host, apps in containers.** One Postgres serves all projects on the box, each project with its own database + own non-superuser role. Simpler ops, shared backups, no volume-managed DB state.
5. **Secrets live in exactly one file per machine, never in git.** Server: `~/apps/<project>/.env` (0600), written over SSH stdin. Laptop: the repo's gitignored `.env`. Docker builds that need secrets get them as BuildKit secret mounts, never as image layers.
6. **Only the front door listens on a public interface.** One Caddy owns 80/443; every app binds `127.0.0.1:<port>`. Upstreams are loopback, never the host's own public IP — a proxy that forwards to the box's public address sends packets out of the container and back in through the host firewall, which is how a firewall change breaks working sites without touching the apps (§5.3). Loopback upstreams need no firewall rule and are unreachable from off-box by construction.

## 1. Prepare a new server (once per server)

Order matters — firewall before anything else listens.

1. **Access:** confirm `ssh <user>@<ip>` with key auth and passwordless sudo.
2. **Tailscale:** install (`curl -fsSL https://tailscale.com/install.sh | sh`), `sudo tailscale up` (James authenticates via the printed URL). Note the 100.x address — it is the server's real address from now on.
3. **Firewall (ufw):** `allow OpenSSH` FIRST, then `allow in on tailscale0`, then `default deny incoming`, `default allow outgoing`, `ufw --force enable`. Verify from outside: only 22 answers publicly. On a box that already runs containers or an inherited proxy, read §5 and §2.1 BEFORE enabling anything — ufw does not filter what you think it filters, and enabling it can break sites it never touches.
4. **Docker:** install Docker Engine + Compose plugin; `usermod -aG docker <user>`.
5. **Postgres (host, not Docker):** install the current PG major from PGDG. Then LOCK IT DOWN before creating anything: `pg_hba.conf` = local peer + scram for `127.0.0.1/32`, `100.64.0.0/10` (tailnet), `172.16.0.0/12` (Docker); never `0.0.0.0/0`. Any client that legitimately lives off the tailnet gets its own `/32`, never a wider net. `listen_addresses='*'` is fine BECAUSE ufw blocks public 5432; also `ufw allow from 172.16.0.0/12 to any port 5432 proto tcp` so containers get through. Reload.
   - **Narrowing an EXISTING pg_hba is `SELECT pg_reload_conf()`, not a restart** — no session is dropped. Copy the file to `pg_hba.conf.<timestamp>` beside itself first, and do not relax ufw because pg_hba is now tight: they are independent layers and each catches what the other misses.
   - **Verify each allowed source by reaching the PASSWORD PROMPT.** A prompt proves a rule matched; `no pg_hba.conf entry for host …` proves none did. You do not need working credentials to prove the rule, and a successful login does not tell you WHICH rule let you in.
6. **Backups:** a `~/backups/backup-<name>.sh` running `sudo -u postgres pg_dump <db> | gzip` per database, nightly via crontab, `find -mtime +14 -delete` rotation. (Dump as the postgres superuser: row-level-security databases cannot be dumped by their owner role.)
7. **Ingress (once per server):** one Caddy container serves ALL projects on the box. Build it with the Cloudflare DNS plugin (2-line Dockerfile: `FROM caddy:2-builder` + `xcaddy build --with github.com/caddy-dns/cloudflare`, copy binary into `FROM caddy:2`). Run with `network_mode: host`, mount a Caddyfile, persist `/data` in a volume. It lives in whichever repo is "primary" on that server.

## 2. Adopt a server that already runs things (only when inheriting a box)

A box with someone else's work on it is not a fresh box. Nothing here is optional.

1. **Baseline before you change anything.** Capture, in one file: every hostname the box serves with its status code and `time_total`, then `docker ps`, `ss -tlnp`, `sudo ufw status verbose`, `sudo iptables -S`, `sudo iptables -t nat -S`, `crontab -l`. Without a before-state your next change is indistinguishable from a fault that was already there — on 2026-08-21 a session enabled ufw without first testing two staging sites and then could not prove whether it had broken them.
2. **Inventory what listens AND which system owns it.** `ss -tlnp` plus `docker ps`, never one alone: host-native services are ufw's business, Docker-published ports are not (§5).
3. **Audit it as if it were hostile.** A box that looks freshly provisioned can still be wide open: the 2026-08-21 box ran Postgres with `listen_addresses='*'` and a `pg_hba` rule allowing `0.0.0.0/0`. Password-protected is not closed. Prove openness from OFF the box, close it, prove it again.
4. **Move GUI ingress into files.** Read every proxy host out of the GUI (NPM: Hosts → Proxy Hosts), write one Caddyfile block per hostname, and change every upstream to `127.0.0.1`. 80/443 have exactly one owner, so the old proxy must stop before Caddy starts: write the complete Caddyfile first, swap in one step, then test every hostname — not just the one you came for.
5. **Broken inherited hostnames are James's call, not yours.** Two of the four here were dead on arrival (nothing listening on the port; DNS pointing at a different box's tailnet address). Carrying a dead hostname in the Caddyfile costs nothing and preserves the knowledge; dropping it loses it silently. Ask, and record the answer in a comment beside the block.
6. **Sunset in this order:** new ingress proven on every hostname → old proxy stopped but not deleted → its containers, images and volumes removed only on James's explicit, current word (§7 q5).

## 3. Onboard a project (repeat per project)

1. **Code path:** on the server `git init --bare ~/repos/<project>.git` and set `git symbolic-ref HEAD refs/heads/main` inside it (clones fail confusingly otherwise if the default branch is main). `git clone ~/repos/<project>.git ~/apps/<project>/repo`. On the laptop: `git remote add vps ssh://<user>@<ip>/home/<user>/repos/<project>.git`. Deploy = `git push vps main` + run the deploy script on the server. (GitHub can be added later as a second remote; scan history for secrets before ANY remote — even a private one.)
2. **Secrets file:** `~/apps/<project>/.env`, mode 0600, written via `ssh 'cat >> …'` with values piped from the laptop (secrets never on a command line — they end up in shell history and `ps`). Generate passwords server-side with `openssl rand -hex 24`.
3. **Database:** as `sudo -u postgres psql`: `CREATE ROLE <project>_admin LOGIN PASSWORD '…'` (+`CREATEROLE` only if the project provisions its own sub-roles), `CREATE DATABASE <project> OWNER <project>_admin`. From containers the host is `172.17.0.1`. Add the database to the backup script.
4. **`deploy/` folder in the project repo:**
   - `Dockerfile` — bake runtime + deps + compiled app into the image (`FROM node:<LTS>-slim`, install the package manager, `COPY . .`, install with a frozen lockfile, build). If the build needs env/DB (Next.js prerendering does), build with `network: host` and mount the env file as a BuildKit secret.
   - `compose.yaml` — one service per long-running process, `restart: unless-stopped`, `env_file: [.env]`, ports bound to `127.0.0.1:<port>` ONLY (Caddy is the only front door). An `ops` service (same image, `profiles: [ops]`, secrets file mounted at the app's expected `.env` path) for migrations/seeds.
   - `deploy.sh` — copy `~/apps/<project>/.env` into `deploy/.env` (gitignored) → `docker compose --profile ops build` → run migrations via the ops service → `docker compose up -d`.
   - Add `.env` to the repo's `.dockerignore` (a `COPY . .` must never be able to bake a secrets file into a layer).
5. **Hostname:** one block in the server's Caddyfile: `<host> { tls { dns cloudflare {env.<TOKEN_VAR>} }  reverse_proxy 127.0.0.1:<port> }`, restart Caddy. DNS: A records (apex and/or wildcard) → the server's TAILSCALE IP, DNS-only (grey cloud — a 100.x address cannot be proxied). The Cloudflare token needs Zone-DNS-edit for that zone, and lives in the ingress project's server `.env`.
6. **Verify like an outsider:** `curl https://<host>` from a tailnet device (200 + valid cert); one real end-to-end action against the deployed app (not just a status page); `nc -z <public-ip> <port>` from outside confirms nothing new is publicly open; `sudo ufw status`. **Anything with client-side behaviour needs a real browser, not curl** — both production bugs on james.in.th's first day (a build-time gate that failed open under SSR, a `<script>` chunk the bundler dropped) returned HTTP 200 with plausible-looking HTML, and neither was visible to anything that does not execute JavaScript and click.

## 4. Going public later (per project, James's call only)

Repoint the A record(s) from the Tailscale IP to the public IP; `sudo ufw allow 443` (and 80 for redirects). Nothing else changes — same certs, same Caddy, same containers. Before flipping: rate-limit/abuse-harden any endpoint that costs money per request (LLM calls especially), and decide whether Cloudflare proxying (orange cloud) is wanted.

**Proxied (orange cloud), as james.in.th now is.** Set the record to Proxied and the zone's SSL/TLS mode to **Full (strict)**: strict requires a publicly-trusted certificate at the origin, which DNS-01 already issued, so Caddy needs no change and renewals still need no inbound port. Two consequences to design for: the origin now sees Cloudflare's addresses, so the visitor's IP must be read from `CF-Connecting-IP` FIRST and the leading `X-Forwarded-For` entry only as a fallback (anything reading the socket address reports a Cloudflare edge — and behind a proxied hostname the first `X-Forwarded-For` entry can be an edge address too, so the two headers are not interchangeable); and origin 443 can optionally be narrowed to Cloudflare's published ranges once you are sure nothing else needs it.

**Repointing a hostname that is already live elsewhere** is the same single record change, with two rules: lower the TTL well before, and leave the old host serving until the new one is verified on every stable path (a Keybase proof, a PGP key, linked PDFs — anything a third party fetches by exact URL).

## 5. Firewalls on a Docker host — five facts that are not intuitive

1. **ufw filters INPUT; Docker-published ports never traverse it.** Docker writes its own rules into `FORWARD` and the `nat` table, so `ufw deny 9200` on a published container port is accepted, appears in `ufw status`, and does nothing at all. The mechanism that works is the `DOCKER-USER` chain, which Docker evaluates before its own rules: `iptables -I DOCKER-USER -p tcp --dport 9200 -s <allowed-ip> -j ACCEPT` followed by a matching `-j DROP`, made permanent with `iptables-persistent` (otherwise it is gone at the next reboot). A box that already carries DOCKER-USER rules — the 2026-08-21 box did, for Elasticsearch and Kibana: one source IP allowed, everything else dropped — is CORRECT. Do not "fix" such rules into ufw; that silently opens the port.
2. **Host-native services are the opposite case,** which is exactly why §1.5 can leave Postgres on `listen_addresses='*'`: a host process's port sits in INPUT, where ufw really does protect it. The same Postgres in a container with a published port would be exposed to the internet.
3. **A proxy that forwards to the host's PUBLIC IP re-enters the firewall.** Enabling ufw with `default deny incoming` silently took two live sites down while both apps stayed perfectly healthy — 0.016 s answers on loopback throughout — because the packets died between the proxy container and the app: out of the container, back into the host's INPUT chain, dropped. GUI proxies invite this; NPM in particular, because `127.0.0.1` genuinely does not work from inside its container, so people write the box's public IP into the forward host. Stopgap: `ufw allow from 172.16.0.0/12 to any port <app-port> proto tcp`. Cure: loopback upstreams (principle 6), after which no app port needs a firewall rule at all.
4. **`ufw allow in on tailscale0` matches the INTERFACE, not the identity.** A client being a member of the tailnet buys it nothing if it dials the server's PUBLIC IP: those packets arrive on the public interface and meet `default deny` like anyone else's. The client must dial the tailnet address — the 100.x, the tailnet IPv6, or the MagicDNS name. "Allow the tailnet" is therefore not something you can do to a client from the server side; it is something the client's configuration has to do.
5. **Closing a port means enumerating its CONSUMERS, not its listeners.** `ss -tlnp`, `docker ps` and `ufw status` inventory what the box SERVES; every client of that port lives on some other machine and is invisible to all three, so an audit built only from them will close a port that something depends on and report success. On 2026-08-21 enabling ufw closed public 5432 on the james.in.th box; seven unattended backup jobs on James's laptop had been dumping databases across it and began failing at ~16:20 that day with `timeout expired`. Nothing alerted, and the breakage surfaced a day later by accident, while answering an unrelated question. **Before closing a port, take a consumer census from the SERVICE's own records, not the box's:** for Postgres, `SELECT DISTINCT client_addr FROM pg_stat_activity` plus the distinct source addresses in the connection log over the preceding days (turn `log_connections=on` on first if it is off — the census needs history, not a snapshot); for HTTP, the access log. Then repoint and re-verify each consumer, or the port stays open. Scheduled clients (cron, launchd, systemd timers) are the ones that never complain: budget for the fact that their failure is silent and their absence is only noticed when someone needs a backup.

**Tailnet addresses: three forms, and which to use.** Each node holds a stable IPv4 (`100.64.0.0/10`), a stable IPv6 (`fd7a:115c:a1e0::/48`) and a MagicDNS name `<host>.<tailnet>.ts.net`. All three reach the same machine and all three survive reboots, reconnects and network changes — they differ in what breaks them and in who is reading them.

- **IPv4 `100.x` — unattended jobs, config files, firewall rules, `pg_hba` CIDRs.** Nothing has to resolve it, so no DNS setting on the client can break it. Default choice for anything that runs while you are asleep.
- **MagicDNS name — anything a human types, and anything that must outlive the node being rebuilt.** It is the only form that normally survives a rebuild, because it follows the hostname. Its own weakness is different: it resolves only while the client uses Tailscale's DNS, so it is the wrong choice for a scheduled job on a machine whose DNS you do not control.
- **IPv6 `fd7a:…` — rarely by hand.** Worth knowing for an IPv6-only client or a v6 ACL; otherwise the v4 address is shorter and every tool takes it.

**What breaks them.** Rebooting, reconnecting, moving between networks: none of these change any of the three. Deleting the node from the tailnet and re-adding it is the one thing that does, and it hits the forms differently — the IPv4 and IPv6 are reallocated, so both change, while the MagicDNS name comes back, because it is built from the hostname. The exception: if the old node record is still in the tailnet, Tailscale avoids the clash by handing the new node `<host>-1`. So the name is the durable one, with a caveat; the addresses are not.

**Moving a client onto the tailnet costs it something.** A public IP is reachable from any network; a `100.x` address is reachable only while Tailscale is running on that client. A scheduled job that used to run from anywhere now runs only when its machine is on the tailnet.

**Enabling a firewall over SSH — arm a dead-man's switch first.** `sudo bash -c 'nohup sh -c "sleep 300; ufw --force disable" >/dev/null 2>&1 & echo $! > /tmp/ufw-dms.pid'`, then enable, then verify from a SECOND connection and from off the box, then cancel with `sudo kill "$(cat /tmp/ufw-dms.pid)"`. Cancel by PID, never by pattern: `pkill -f "sleep 300"` also matches the shell that is running it, because the pattern is in that shell's own command line — so it kills your session and the safety timer together, the worst of both outcomes.

**`ufw status` states an intention; only an outside probe states a fact.** `nc -z <public-ip> <port>` and `curl https://<host>/` from another machine, against every hostname on the box, before and after any firewall change.

## 6. Gotchas (each cost real debugging time — read before your first deploy)

- **`docker compose build` silently skips services behind `profiles:`.** Build with `--profile <name>` or your ops/migrate container runs a stale image.
- **Fresh databases find ordering bugs that dev DBs hide.** A dev DB that "always had" a role/extension/table masks migration-order assumptions. The first deploy IS the first true test of migrations-from-zero; expect one fix.
- **Some scripts require a `.env` FILE and ignore injected env vars.** Compose `env_file` sets process env; a script that walks the disk for `.env` still fails. Fix: mount the secrets file at the expected path in the ops container only.
- **Row-level security with FORCE binds the table OWNER too:** owner-role `pg_dump` fails; back up as the postgres superuser via peer-auth sudo.
- **`.pgpass` is keyed by the host STRING, not the address it resolves to.** Repointing a client from `<public-ip>` to a tailnet IP or a MagicDNS name makes libpq stop matching the existing line, and the failure is `fe_sendauth: no password supplied` — which reads as a credentials or auth-method bug and is neither. Add a line per host form the client may use, and remember the fix belongs on the CLIENT, not in `pg_hba`.
- **Bare repos default HEAD to `master`;** set the symbolic-ref to main or clones check out nothing.
- **Auth origin allowlists must learn the deployed domain.** Cookie/origin-checking auth (Better-Auth etc.) configured for localhost rejects the real domain with opaque 403s. Make trusted origins env-driven (e.g. a `PUBLIC_ORIGIN_SUFFIX`) so dev needs nothing and prod sets one var.
- **`NEXT_PUBLIC_*` vars are baked at BUILD time** — they must be present during the image build (the BuildKit secret mount covers this), not just at runtime.
- **Two builds from one codebase: compile the switch in, do not hide it at runtime.** One env var (`BUILD_MODE=ssr`) chooses the adapter and output dir; a `vite.define` replaces a flag with a boolean LITERAL so the server-only markup folds away and is absent from the static output. Hiding it with CSS would still ship it on a static host.
- **Astro's `build.server` resolves RELATIVE to `outDir`.** Setting `outDir: './dist-ssr'` and `build.server: './dist-ssr/server'` together produces `dist-ssr/dist-ssr/server/`. Set `outDir` alone and let the defaults follow it.
- **`output: 'server'` ignores `getStaticPaths()`,** warns that it did, and suggests `export const prerender = true`. Do NOT take that suggestion on a route that exists to render per request — it freezes the route and defeats the entire build. Validate `Astro.params` at request time against the same map the static build enumerates.
- **`Astro.rewrite('/404')` renders the 404 page with status 200.** Wrap it — `new Response(await (await Astro.rewrite('/404')).text(), { status: 404 })` — or you have shipped a soft 404 that crawlers happily index.
- **Component frontmatter that reads the filesystem is build-time in a STATIC build and PER REQUEST in SSR** — inside a runtime image that carries no `public/` directory at all. A "build-time gate" written that way does not merely stop working, it fails OPEN: james.in.th shipped `<audio>` elements for files that were not there and an empty sound manifest that silently unbound 27 handlers, while every page still returned 200. Resolve such inventories ONCE in `astro.config.mjs` and pass them through `vite.define`, the same mechanism as the build-mode flag — then both modes agree by construction.
- **Astro hoists a component's `<script>` blocks and merges them into ONE chunk.** `{condition && <script>…}` therefore does not exclude that script; it can drop the whole chunk, taking unrelated scripts with it. Here it removed the click-sound engine from all six locked routes while leaving it working on `/`. Gate such a script at RUNTIME, inside the script (an early return on a DOM attribute), never with a conditional the bundler is free to reinterpret.
- **A container HEALTHCHECK must probe the cheapest thing that proves the process serves.** Pointing a 30-second check at the app's real page is 2,880 renders a day, and if rendering calls third-party APIs each render is billable, rate-limited traffic that no visitor asked for — james.in.th earned an HTTP 429 from a news API purely by asking itself whether it was alive. Probe a static asset the same process serves (`/favicon.ico`).
- **A 200 does not prove the intended image is running.** A stale image answers identically. A deploy script must assert on CONTENT unique to the build it meant to ship (`curl -s http://127.0.0.1:<port>/ | grep -q '<marker>'`) and say so loudly when the marker is missing.
- **A gitignored file the RUNTIME needs is a hole in "rebuild from the repo + one secrets file".** james.in.th's audio is deliberately out of git, so it exists only in the server's checkout, must be re-uploaded after a fresh clone, and must NOT be excluded by `.dockerignore` or the image ships without it. Either commit such assets or write the out-of-band list into the server notes beside the secrets file — principle 1 only holds for what is actually in the repo.
- **Plain `pnpm install` can drift lockfiles** (re-floating transitive versions). Use `--frozen-lockfile` everywhere outside deliberate upgrades.
- **A brand-new VPS may arrive wide open** (DB listening publicly, no firewall). Audit and harden BEFORE putting anything on it — scanners find fresh IPs within hours. The 2026-08-21 box proved the sharper version: a tidy, in-use machine had Postgres on `listen_addresses='*'` with `pg_hba` allowing `0.0.0.0/0` — password-protected and internet-reachable. "Looks maintained" is not evidence; only a probe from outside is.

## 7. Questions a session must ask James before deploying something new

These are HIS decisions; the pattern deliberately does not default them. Everything else in this guide needs no questions.

1. **Which server?** Existing box (which?) or new VPS (then: region/specs — his convention: resource-heavy → Germany, cheap+big).
2. **Which domain/hostname?** Existing zone or a new domain (new → he registers it, puts DNS on Cloudflare, mints a Zone-DNS token).
3. **Tailnet-gated or public?** (Default tailnet; public needs his explicit word + the §4 checklist.)
4. **Is there existing data anywhere that must move?** (Old server/DB → export+import phase; "fresh start" must be his words, never assumed.)
5. **May anything old be turned off?** Sunsetting servers/services is destructive — his explicit, current authorization each time.
6. **Anything unusual about secrets?** (Third-party tokens to mint/rotate, file-based keys like APNs `.p8` needing volume mounts.)

## 8. Daily ops crib

- Deploy: `git push vps main` → `ssh <server> 'cd ~/apps/<project>/repo && git pull && deploy/deploy.sh'` — the script copies the secrets file in, builds the image on the box, and (james.in.th's does) refuses to claim success until content unique to the intended build renders.
- Logs: `docker compose logs -f <service>` · restart: `docker compose restart <service>` · what's running: `docker compose ps`
- One-off tasks: `docker compose --profile ops run --rm ops <command>`
- DB from the laptop: connect to `<tailscale-ip>:5432` (works because pg_hba trusts the tailnet). Give `.pgpass` a line for whichever host string you actually type.
- Before changing anything about 5432 or `pg_hba`: `sudo -u postgres psql -c 'SELECT DISTINCT client_addr FROM pg_stat_activity'` plus the connection log — the consumers live on other machines and no on-box inventory will show them (§5.5).
- Change a hostname or upstream: edit the Caddyfile in the repo → push/pull → `docker compose up -d --force-recreate caddy`. Force-recreate, not restart: a single-file bind mount follows the inode, and `git pull` replaces the file.
- Firewall truth needs two commands, not one: `sudo ufw status verbose` (host services) AND `sudo iptables -S DOCKER-USER` (containers).
- Prove a hostname from off the box: `curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' https://<host>/`.
- Rebuild the world after total server loss: new server → §1 → restore `.env` files + latest dump → §3 per project.
