# deploy/ — the whole production story (spec §9.48 / O12)

Target: the Germany VPS (Ubuntu 24, `james@92.118.206.89`, Tailscale `100.71.64.33`).
Design decisions and rationale: spec §9.48. Strategy discussion: 2026-08-19 session.

## Shape

- One Docker image per app (`deploy/Dockerfile`, `ARG APP`), compose fleet in
  `deploy/compose.yaml`. **No public ports** — apps bind 127.0.0.1 on the box;
  public ingress (Cloudflare Tunnel + a domain) is a separate later step.
- Postgres 18 runs on the HOST (not in compose): database `ai_factory`, admin
  role `factory_admin` (CREATEROLE, non-superuser), app role `ai_kernel_app`
  provisioned by `pnpm db:migrate` itself. pg_hba allows localhost + Tailscale
  (100.64.0.0/10) + Docker (172.16.0.0/12) only; ufw allows SSH + tailnet only.
- Secrets live ONLY in `~/apps/ai-new-business/.env` on the server (0600).
  `deploy.sh` copies it to `deploy/.env` (gitignored) for compose; the build
  mounts it as a BuildKit secret so it never becomes an image layer.
- Code reaches the box by git push to a bare repo (`~/repos/ai-new-business.git`),
  checkout at `~/apps/ai-new-business/repo`. From the Mac: `git push vps main`.

## Deploy (from the Mac)

```
deploy/push-and-deploy.sh            # push main, start the detached server deploy, tail its log
NOTIFY=1 deploy/push-and-deploy.sh   # same + one line to James's n8n webhook
```

Fast pipeline (2026-09-24): `apps/paiduay/docs/DEPLOY-FAST.md` — layered
Dockerfile, skipped in-image type check, gated standalone runtime, detached
`deploy/deploy.sh` (`--all` = the old whole-fleet + migrate behaviour).

## First-time server bootstrap (already done 2026-08-19; recorded for rebuild)

```
git init --bare ~/repos/ai-new-business.git
git clone ~/repos/ai-new-business.git ~/apps/ai-new-business/repo
# secrets file: ~/apps/ai-new-business/.env  (see .env.example for keys;
#   POSTGRES_HOST=172.17.0.1 — the host as seen from containers)
deploy/deploy.sh
# seed demo data (one-time):
docker compose --profile ops run --rm ops pnpm db:seed
docker compose --profile ops run --rm ops pnpm --filter clinic-demo exec tsx scripts/seed-demo-data.ts --fresh
```

## Look at it (§9.49 — live, tailnet-gated)

From any device on James's tailnet:
- **https://6326638.xyz** — the hub (entry point + whole doc corpus, browsable)
- **https://demo.6326638.xyz** · **clinic.** · **barber.** · **cameo.** — the apps

DNS (apex + wildcard) points at the Tailscale IP `100.71.64.33` — publicly
unroutable, so no one off the tailnet can connect; ufw also keeps 80/443 shut.
Certs: real Let's Encrypt via DNS-01 (Caddy + cloudflare plugin, token in the
server `.env`). **Go fully public later:** repoint DNS to `92.118.206.89` +
`sudo ufw allow 443` — but only after O23 (GET hardening) to protect API credit.

Fallback without tailnet: `ssh -L 3300:localhost:3300 james@92.118.206.89`.

## Ops

- Logs: `docker compose logs -f demo` · restart: `docker compose restart clinic`
- Migrate only: `docker compose --profile ops run --rm ops pnpm db:migrate`
- Backups: nightly `pg_dump` of `ai_factory` (03:20) + nightly tarball of the
  paiduay uploads volume (03:40) — see "Backups" and "Restore" below.

## Backups (what is protected, and what is not)

Two nightly jobs, both in `james`'s crontab on the box, both writing to
`~/backups`:

```
20 3 * * * /home/james/backups/backup-ai-factory.sh
40 3 * * * /home/james/apps/ai-new-business/backup-uploads.sh >> /home/james/backups/backup-uploads.log 2>&1
```

| Job | Covers | Retention | Verified |
|---|---|---|---|
| `backup-ai-factory.sh` | `pg_dump` of the whole `ai_factory` database — every schema, so `kernel.records`, `kernel.outbox`, `kernel.hits`, `kernel.attachments` metadata, `kernel_jobs.*` | 14 days (`-mtime +14`) | `gunzip -t` |
| `backup-uploads.sh` (`deploy/backup-uploads.sh`, deployed to `~/apps/ai-new-business/`) | the `ai-factory_paiduay-uploads` volume — sellers' card photos (`/data/uploads`) | 30 days | `gzip -t` + `tar tzf` inside the script |

Notes worth knowing before you rely on them:

- The dump is **plain SQL**, not custom format — restore with `psql`, not
  `pg_restore`. It contains no `CREATE DATABASE` and no `CREATE ROLE`, but does
  `ALTER … OWNER TO factory_admin` and `GRANT … TO ai_kernel_app`: **both roles
  must exist before you restore**, or every ownership line fails.
- A dump is only as new as 03:20. Migrations applied during the day (e.g.
  `0004_outbox.sql`, `0005_hits.sql` on 2026-09-17) are not in that morning's
  file — they land in the next one.
- Uploads retention (30d) is deliberately **shorter** than the 90-day photo
  retention policy, so photos cannot outlive the policy inside archives. An
  erasure request is only fully honoured once the last archive containing the
  photo ages out — up to 30 days.
- **Everything is on one box.** Both backups sit on the same disk as the data
  they protect: they survive `rm -rf`, a bad migration and a redeploy, but not
  loss of the VPS. See "Off-box copy" below.

## Restore

Order matters: **DB first, then uploads, then bring the fleet up.** Restoring
uploads first is harmless; restoring the DB *after* the apps are running is not
(live writes race the restore).

Both halves were rehearsed on 2026-09-17 against the 2026-09-17 archives — the
dump into a throwaway database (`ON_ERROR_STOP=1`, clean: 10 records / 4 tenants
/ 19 parties) and the tarball into a throwaway volume (7 files) — so the commands
below are copied from a run that worked, not from memory.

```bash
# 0. stop everything that writes (do NOT touch the luna/tuaton projects)
cd ~/apps/ai-new-business/repo/deploy
docker compose stop demo clinic barber cameo paiduay

# 1. DATABASE  (ai_factory, plain-SQL gz dump -> psql)
DUMP=~/backups/ai_factory-2026-09-17.sql.gz
gunzip -t "$DUMP"                       # integrity first, always
sudo -u postgres dropdb --if-exists ai_factory
sudo -u postgres createdb ai_factory
# roles the dump assumes but does not create (skip any that already exist):
sudo -u postgres psql -c "CREATE ROLE factory_admin LOGIN CREATEROLE;"
sudo -u postgres psql -c "CREATE ROLE ai_kernel_app LOGIN;"
zcat "$DUMP" | sudo -u postgres psql -v ON_ERROR_STOP=1 -d ai_factory

# 2. UPLOADS  (photos back into the named volume)
ARC=~/backups/paiduay-uploads-2026-09-17.tgz
gzip -t "$ARC" && tar tzf "$ARC" | head
docker volume create ai-factory_paiduay-uploads
docker run --rm \
  -v ai-factory_paiduay-uploads:/dst \
  -v ~/backups:/src:ro \
  alpine:3.20 \
  sh -c 'rm -rf /dst/* && tar xzf /src/paiduay-uploads-2026-09-17.tgz -C /dst'

# 3. UP
docker compose up -d
docker compose ps
```

Restoring into a *fresh* box instead: bootstrap first (see "First-time server
bootstrap"), then run the two restores above **instead of** `pnpm db:seed` —
`deploy/deploy.sh` will run `pnpm db:migrate` over the restored schema, which is
a no-op when the dump is current.

Spot-check after a restore: card count matches
(`sudo -u postgres psql -d ai_factory -Atc "select count(*) from kernel.records"`)
and a card's photo actually loads on https://paiduay.6326638.xyz.

## Off-box copy (NOT yet implemented — recommendation)

Both archives live on `/dev/sda1` next to the database. If the VPS is lost, so
are the backups. MinIO (`MINIO_HOST` in the server `.env`) runs on a *different*
host, `46.247.108.78:9000`, so it genuinely is off-box — but it is a second
single machine James owns, not durable storage.

Cheapest real fix, in preference order:

```bash
# A. Cloudflare R2 (10 GB free, zero egress fees; the Cloudflare account
#    already exists for the DNS-01 certs). One-time:
sudo apt install -y rclone
rclone config create r2 s3 provider=Cloudflare \
  access_key_id=<R2_KEY> secret_access_key=<R2_SECRET> \
  endpoint=https://<ACCOUNT_ID>.r2.cloudflarestorage.com
# then one cron line at 04:10:
10 4 * * * /usr/bin/rclone sync /home/james/backups r2:ai-factory-backups \
  --include 'ai_factory-*.sql.gz' --include 'paiduay-uploads-*.tgz' \
  --max-age 35d >> /home/james/backups/offsite.log 2>&1

# B. Pull to James's Mac (free, but only as reliable as the Mac being on).
#    Runs FROM the Mac, so a lost VPS cannot delete the copies:
rsync -av --include='ai_factory-*.sql.gz' --include='paiduay-uploads-*.tgz' \
  --exclude='*' james@92.118.206.89:/home/james/backups/ ~/backups/ai-factory/

# C. Existing MinIO at 46.247.108.78 — same rclone recipe, s3 provider=Minio.
#    Better than nothing, worse than R2: one more box to lose.
```

PDPA angle: the uploads archive contains sellers' photographs — personal data.
Anywhere it is copied inherits the obligations. R2 lets you pin the bucket to a
jurisdiction and set a 35-day lifecycle rule so the off-box copy ages out on the
same clock as the local one; an unmanaged copy on a laptop does not, and an
erasure request then has to chase it by hand. Encrypt at rest before it leaves
the box (`rclone` crypt remote, or `gpg -c`) if photos are ever more than the
handful of demo cards there today.
