# AI gateway — cost and resilience (WO-1 … WO-7)

**APPROVED by James 2026-09-18, DEFERRED — build when he says so; do not start it unsolicited.**
**Remind him it is pending in every session until it is built** (one line, see §5).

Written by Fable 5.1, 2026-09-18. Every number below was measured from the code; the estimator is
`chars ÷ 3.5`. Line numbers are as of this date, and `packages/core/src/config.ts` is being edited **in
flight** by the G1 session (Gemini adapter + `KERNEL_PROVIDER`) — locate things there by symbol, not line.

---

## 1. The problem, in today's numbers

A single receptionist turn on `frenday.xyz` billed **~21,500 input tokens** on Opus 4.8 (~$0.30/turn).
Here is where every one of them comes from.

| Piece | Where | chars | ~tokens |
|---|---|---:|---:|
| System prompt | `apps/paiduay/content/ai.ts:20`–54 → manifest `prompts["webchat.system"]` (`lib/manifest.ts:130`) | 10,971 | ~2,900–3,150 |
| Policy grounding (4 chunks, **all four always returned**) | `lib/manifest.ts:151`–156 → `packages/modules/scheduled-interactions/src/chat.ts:27`–31 | 7,553 | ~2,160 |
| Tool definitions (4) | `create_booking` 230 · `get_bookings` 56 · `get_companion_profile` 147 · `request_meetup` 537 | 3,397 | ~970 |
| **Stable prefix per provider call** | | **21,921** | **~6,260** |
| `get_companion_profile` result (one provider) | `companion-module.ts` `buildProfile` | ~2,600 | ~730 |

6,260 ≠ 21,500. **The multiplier is the tool loop.** `AnthropicAdapter.chat`
(`packages/core/src/ai/anthropic.ts:42`–88) loops, and line 51 does
`usage.inputTokens += response.usage.input_tokens` — so one *turn* is 2–3 *billed calls*, each
re-sending the whole prefix plus the growing tail (history via `conversationAsChatMessages`, the
profile JSON, the request-meetup result). 3 × 6,260 + tail ≈ 21,500. **We pay for the same 6,260
tokens three times per turn, and nothing is cached.**

Two further findings, both measured:

- **The grounding is static but arrives non-deterministically.** `LexicalRetriever`
  (`ai/retriever.ts:15`–24) does `ORDER BY score DESC` with `k=4` over exactly 4 rows, so the *set* is
  always identical but the *order* changes with the customer's wording. The manifest comment at
  `lib/manifest.ts:24`–32 says this is deliberate (the brain must see the whole policy). Retrieval
  therefore buys nothing and costs cache-stability.
- **A timestamp poisons the prefix.** `chat.ts:37` appends
  `Current date & time: ${new Date().toISOString()}` — milliseconds — into the *single* system string.
  Nothing in `system` can ever cache while that is true.

One correction to the brief: the activity/area tables (489 + 195 tok) and the all-providers table
(6,118 tok) are **never sent to the model** — they are resolved server-side. WO-2 is re-scoped below.

**Harness cost.** `--app paiduay --battery all` = 46 fixtures / **94 turns** (consequence 15, injection 8,
quotes 17, refusals 20, request 10, safety 24). 94 × 21.5k ≈ **2.0M input tokens ≈ $10 on Opus 4.8
input alone.** James's $5 top-up bought half a pass. That is the whole story of this morning.

---

## 2. The work orders

### WO-1 — Prompt caching · ~5 h

**Goal:** cache the 6,260-token prefix so the 2nd and 3rd round of every turn, and every later turn in
the conversation, read it at 0.1× instead of paying 1×.

**Change:**
- `packages/core/src/ai/gateway.ts:17`–20 — `Usage` gains `cacheReadTokens` + `cacheWriteTokens`.
- `gateway.ts:101`–108 — `AdapterChatRequest.system: string` becomes
  `system: Array<{ text: string; cache?: boolean }>` (ordered, provider-agnostic; a Gemini adapter maps
  `cache: true` to its own mechanism or ignores it).
- `packages/core/src/ai/anthropic.ts:34`–50 — render `tools` with `cache_control` on the **last** tool,
  and `system` as text blocks with `cache_control: { type: "ephemeral" }` on the last cached block.
  Sum the three usage fields (see below). Never mutate tool order.
- `packages/modules/scheduled-interactions/src/chat.ts:34`–38 — emit
  `[{prompt, cache:false}, {policy, cache:true}, {timestamp-and-volatile, cache:false}]`. Move the
  clock **after** the breakpoint and round it to the minute.

**API (cited).** `CacheControlEphemeral = { type: 'ephemeral'; ttl?: '5m' | '1h' }` —
anthropic-sdk-typescript `src/resources/messages/messages.ts` (context7). Placement on a system content
block: `tests/api-resources/messages/messages.test.ts`. Render order is `tools → system → messages`; a
breakpoint on the last system block caches tools **and** system (claude-api skill,
`shared/prompt-caching.md`). Max **4** breakpoints per request. Default TTL `5m`; cache **write** 1.25×
(5m) / 2× (1h), **read** 0.1× (same source, § Economics). Total input =
`input_tokens + cache_creation_input_tokens + cache_read_input_tokens` (SDK `Usage` JSDoc; summed in
`src/lib/tools/BetaToolRunner.ts`).

**The trap.** Minimum cacheable prefix is model-dependent and **Haiku 4.5's is 4,096 tokens** — shorter
prefixes silently do not cache, no error (`shared/prompt-caching.md` § API reference). Tools + system
alone is ~4,100 tokens: *razor thin*. **The breakpoint must sit after the policy block (~6,260), never
after `system`.** This is a hard constraint on WO-2.

**Acceptance.** `ai/caching.test.ts` → `describe("cache breakpoints")`:
`it("marks the last tool and the last cached system block, and nothing after it")` — asserts exactly one
`cache_control` in `tools`, one in `system`, the volatile block after it carries none, and the cached
prefix is ≥ 4,096 estimated tokens. `it("renders byte-identical prefixes for two different customer
messages")`. **Live:** two consecutive turns in one conversation; the second must report
`cache_read_input_tokens ≥ 6,000`. Replace the `÷3.5` estimate with `client.messages.countTokens` in the
same session and record the real figure in spec §9.

**Risks:** any upstream invalidator drops the rate to zero with no error — the second unit test guards it.
Caches are workspace-scoped.

### WO-2 — Grounding trim, re-scoped · ~3 h

**Goal:** stop paying for retrieval that retrieves everything, and cap the **uncached** part of a turn at
**≤ 6k input tokens**.

**Change:** `packages/modules/scheduled-interactions/src/chat.ts:27`–31 — when the tenant's grounding
set is small enough to send whole (it is: 4 chunks, 2,160 tok), skip `retriever.search` and render the
manifest's chunks **sorted by `source`**, deterministically. Retrieval stays for tenants whose grounding
outgrows a threshold (`GROUNDING_INLINE_MAX_TOKENS`, config, default 3,000). Second: cap history —
`conversationAsChatMessages` grows without bound; keep the last N turns (`CHAT_HISTORY_TURNS`, default 12)
and count the rest as the tail budget. The companion record already arrives as a tool result for **one**
provider only (~730 tok) — nothing to trim there; do not "optimise" it into the prompt, that would break
the get-profile-before-answering doctrine the prompt depends on.

**Acceptance.** `chat.grounding.test.ts` → `it("inlines a small grounding set in source order and never
calls the retriever")` and `it("falls back to retrieval above the inline threshold")`. **Live:** one turn;
assert `input_tokens + cache_creation` (i.e. the uncached part) ≤ 6,000.

**Risks:** trimming the cached prefix below Haiku's 4,096 floor silently disables WO-1 — the WO-1 test
guards it. Trimming history too hard breaks multi-turn refusal memory (safety battery catches it).

### WO-3 — Per-task model routing · ~4 h

**Goal:** chat, structured extraction and vision stop sharing one model — and each gets the request
shape its model actually accepts.

**Change:** `packages/core/src/config.ts` — add `KERNEL_MODEL_CHAT`, `KERNEL_MODEL_OBJECT`,
`KERNEL_MODEL_VISION`, each `.optional()`, each falling back to `KERNEL_MODEL_DEFAULT`; expose as
`KernelConfig.models: { chat, object, vision, default }`. `packages/core/src/ai/llm.ts:64`–70 —
`resolve(ai)` takes the call kind (`"chat" | "object" | "embed"`, already known at each call site:
`llm.ts:96`, `:133`, `:153`) and picks accordingly; an explicit `AIContext.model` still wins. Provider is
orthogonal — this composes with G1's `KERNEL_PROVIDER` (in flight; do not touch its internals).

**Ship a per-model capability table** in `packages/core/src/ai/models.ts`: min cacheable prefix, thinking
mode, effort support, prices. **Pre-flight blocker to check first:** `anthropic.ts:47` hardcodes
`thinking: { type: "adaptive" }`. Per the claude-api skill's Thinking & Effort table, adaptive is the
4.6+ on-mode; **Haiku 4.5 takes `{ type: "enabled", budget_tokens: N }`** (min 1024, `< max_tokens`).
Commit `86b392e` made Haiku the default, so this line is now either a 400 or an ignored field on every
production chat. Verify against a real call the moment credits exist, and make the thinking block a
function of the model, not a constant.

**Acceptance.** `ai/routing.test.ts` → `it("routes chat/object/vision to their configured models and
falls back to the default")`, `it("shapes thinking per model from the capability table")`. **Live:** one
chat + one slip read (`lib/pro/slip-reader.ts`) with different models configured; two `ai_usage` rows with
two different `model` values.

**Risks:** a per-call model change **invalidates the cache** for that namespace — routing and caching
interact, so route per *purpose*, never per request.

### WO-4 — Error classification · ~3 h

**Goal:** no silent failure, ever. Today `apps/paiduay/app/api/chat/route.ts:70`–76 catches everything,
logs, and returns 503 + `chat.unavailable` — which is the right customer answer but tells James nothing.

**Change:** `packages/core/src/errors.ts` — add `ProviderUnavailableError extends KernelError` with
`{ provider, status, retryable, reason }`. `packages/core/src/ai/anthropic.ts` — wrap both `chat()` and
`object()`: map `Anthropic.BadRequestError` whose message contains `credit balance is too low` →
`reason: "credit"`; `RateLimitError` (429) → `"rate_limit"`; 529 → `"overloaded"`; any 5xx →
`"provider_error"`; `APIConnectionError` → `"network"`. Use the SDK's typed classes, never string-matching
the status. `packages/core/src/ai/llm.ts:103`–116 — on `ProviderUnavailableError` in `chat()`, mirror the
budget-denial shape (`llm.ts:88`–93): return the manifest's honest line (`chat.unavailable` already exists
bilingually — `apps/paiduay/content/copy.ts:352`, wired at `lib/manifest.ts:122`) with
`stopReason: "provider_unavailable"`, zero usage, no `ai_usage` row. `object()`/`embed()` keep throwing.

**One event per hour to James.** In the same branch call `enqueueOutbox` (`packages/core/src/outbox/outbox.ts:125`)
with `eventType: "ai.unavailable"` and
`dedupeKey: \`ai.unavailable:${reason}:${new Date().toISOString().slice(0,13)}\`` — the unique
`(tenant_id, dedupe_key)` + `onConflictDoNothing` at `:138` **is** the once-an-hour guard, no new table,
no new state. It rides the existing n8n webhook.

**Acceptance.** `ai/unavailable.test.ts` (real DB + a `MockAdapter` that throws each class) →
`it("answers the honest unavailable line and meters nothing when the provider refuses")`,
`it("enqueues exactly one ai.unavailable outbox row per hour per reason")` (call it five times, assert
one row). **Live:** point `CLAUDE_API` at a dead key, send a message, see the Thai unavailable line in the
chat and one webhook on James's phone.

**Risks:** over-broad classification could swallow a genuine 400 (bad request shape) as "unavailable" and
hide a bug — classify only the listed statuses and re-throw everything else.

### WO-5 — Cost visibility in baht · ~4 h

**Goal:** turn spend from a mystery into a row on `/admin/numbers` (screen **Snake**). Spec §9.27 already
names the blocker: *"Cost-based budgets (USD) need a per-model price table — the `ai_usage.cost_usd`
column exists and stays unpopulated until then."* This closes it.

**Change:** price table in `ai/models.ts` (WO-3) — `{ inputPerMTok, outputPerMTok, cacheReadMult: 0.1,
cacheWrite5mMult: 1.25, cacheWrite1hMult: 2 }`; seed Haiku 4.5 `$1 / $5`, Opus 4.8 `$5 / $25`, Opus 5
`$5 / $25` (claude-api skill model table, cached 2026-06-24 — re-check on build).
`KERNEL_USD_THB` in config (no default guess — set it from the day's rate, document it).
Migration `packages/core/migrations/0006_ai_usage_cache.sql`: `ALTER TABLE kernel.ai_usage ADD COLUMN
cache_read_tokens integer NOT NULL DEFAULT 0, ADD COLUMN cache_write_tokens integer NOT NULL DEFAULT 0`
(the table's RLS policy and grants from `0001_init.sql:139`–150 / `:233` already cover new columns —
no new policy needed); mirror in `db/schema.ts:275`–290. `ai/llm.ts:188`–197 `meter()` computes and writes
`cost_usd`. Then on `apps/paiduay/app/admin/numbers/page.tsx` add a fourth `iv-stat` tile reading today's
`ai_usage` (Bangkok day): **"วันนี้ใช้ไป ฿…"**, with a `small` line showing `tokens used / daily budget`
(`KERNEL_DAILY_TOKEN_BUDGET`) and the cache-hit percentage.

**Acceptance.** `ai/cost.test.ts` → `it("prices a metered call from the model table including cache read
and write multipliers")` (pure), `it("writes cost_usd on every metered row")` (DB).
`admin-numbers.test.ts` → `it("shows today's baht spend and the budget it is against")`. **Live:** send one
message, reload `/admin/numbers`, the baht figure moves; it also lands free in the trust harness, which
already sums `costUsd` (`packages/trust-harness/src/observe.ts:95`–118) and reports 0 today.

**Risks:** a stale price table lies confidently — stamp the table with the date it was verified and print
that date under the tile.

### WO-6 — Harness spend guard · ~3 h

**Goal:** the harness prints what a run will cost **before** spending it, and refuses to burn a top-up by
accident.

**Change:** `packages/trust-harness/src/run.ts` — `turnCount` already exists at `:169`. Multiply it by a
measured per-turn figure (from WO-1's `countTokens` pass, stored as a constant per app in `apps.ts`),
print `estimate: ~N turns × ~M tok = ~X tok ≈ ฿Y`, and **refuse** above `HARNESS_MAX_TOKENS` (env,
default 300,000) unless `--allow-spend`. The estimate prints on every run, including `--dry-run` (`:174`–184),
which becomes the natural "how much would this cost" command.
Add `--provider` / `--model`: **these are declarations the harness verifies, not settings it applies** —
the harness drives the app over HTTP, so the model lives in the *server's* env.
`run.ts:279` currently reports `process.env.KERNEL_MODEL_DEFAULT ?? …` from the **harness's** env, which is
misleading. Fix it by reading the model back from the `ai_usage` rows `observe.ts` already fetches (add
`model` to its select) and **fail with exit 2 when the observed model ≠ `--model`.**

**Acceptance.** `harness/estimate.test.ts` → `it("estimates from fixture turn counts and refuses above
HARNESS_MAX_TOKENS without --allow-spend")` (pure, exit 2, no HTTP), `it("fails when the observed model is
not the declared one")`. **Live:** `--battery all --dry-run` prints ~94 turns and a baht figure; without
`--allow-spend` a full pass refuses.

**Risks:** an under-estimate gives false comfort — derive it from the *measured* prefix, re-measure after
any prompt change, and print the date the constant was measured.

---

## 3. WO-7 — Per-model prompt and harness tuning · ~8 h + credits

**James's explicit order, 2026-09-18.** The prompt, the grounding and the harness's checks were all tuned
against **Opus 4.8**. Haiku 4.5 and Gemini flash-lite are being swapped in **as-is**, knowingly. Later,
tune properly. This is not optional polish: a cheaper model that leaks the doctrine costs more than it
saves, and the doctrine is the product.

**Procedure.**
1. WO-1…WO-6 first. Without WO-6's guard and WO-5's baht column this run is unaffordable and unmeasurable.
2. Freeze the fixtures. Run **every** battery (94 turns) once per (provider, model) pair:
   `claude-haiku-4-5`, `claude-opus-4-8`, and G1's Gemini model. One pass ≈ ฿— per WO-5's tile; budget it
   before starting and tell James the figure first.
3. Keep **one results table per model** in `docs/trust-reports/per-model-<date>.md`: model, pass / fail /
   untested per battery, doctrine pass rate, tokens/turn, cache-hit rate, ฿/turn, ฿/pass.
4. Where a model leaks (accepts on the companion's behalf, invents a rate, answers Thai in English,
   breaks the age or money rules), **write a prompt variant**, not a global prompt change:
   `content/ai.ts` gains `systemFor(model)` returning the base prompt plus a model-specific addendum.
   The base prompt stays the Opus one; addenda are per-model deltas so a fix for Haiku cannot regress Opus.
5. Re-run only the battery that failed until it passes, then re-run all six to prove nothing else moved.

**Acceptance bar (the decision rule James approved):** ship the model with the **highest doctrine pass
rate**, and among models that tie within one check, the **cheapest ฿/pass**. Absolute floor: **every
`sev: "critical"` safety and injection check passes, zero `untested`** — a cheap model that fails one
money, age, danger or impersonation check is disqualified regardless of price. Record the chosen model and
the table's path in spec §9 and set `KERNEL_MODEL_CHAT` to it.

---

## 4. Build order and parallelism

| Wave | Runs | Why |
|---|---|---|
| A | **WO-4** ∥ **WO-6** | Independent of prompt shape, and neither needs credits to verify (mock adapter / `--dry-run`). First: silent failure and accidental spend are the two live hazards. |
| B | **WO-3** | Its capability table is a dependency of WO-1 (cacheable minimum) and WO-5 (prices), and it carries the Haiku thinking fix. Coordinate with G1 — both touch `config.ts` and `llm.ts:64`–70. |
| C | **WO-1**, then **WO-2** | Same files (`chat.ts:27`–38); WO-2's trim is only safe once WO-1's floor test exists. |
| D | **WO-5** | Needs WO-3's price table and WO-1's cache-token fields. |
| E | **WO-7** | Needs all of the above and a funded key. |

Two sessions max, per `docs/handoff/60-opus-execution-protocol.md`: one on A+B, one waiting for B before
C+D. Never gate two checkouts against the shared Postgres at once.

---

## 5. How to remind James

At the **top** of the first substantive reply in any session that touches AI code, prompts, models,
pricing or the trust harness, one line, no ceremony:

> Standing reminder: the AI gateway cost/resilience work orders (`docs/foundation/work-orders/ai-gateway-cost-and-resilience.md`, WO-1…WO-7) are approved and deferred — say the word and I'll build wave A (error classification + harness spend guard; no credits needed to verify).

Drop the reminder only when spec §9 records the build. If credits are exhausted, lead with WO-4 and WO-6 —
they are the two that are *cheaper to build than to keep not building*.
