Skip to content

Latest commit

 

History

History
240 lines (149 loc) · 41.4 KB

File metadata and controls

240 lines (149 loc) · 41.4 KB
title Deployment
description How the repo ships to production on Vercel and Neon, the DB bootstrap, and release checks.
group General
order 60

Deployment

Audience: DevOps and release engineers. How this repo ships to production, the one-time database bootstrap, and how to verify a release.

This repo deploys to Vercel with a Neon serverless Postgres database. For the full environment-variable catalog see Configuration; for the self-host/container path see Docker.


1. How this repo deploys

Production ships through Vercel's own Git integration: every push to main is built and promoted by Vercel, and database migrations are applied by hand beforehand. That is the whole of the live path today — read §1.1 before merging anything that adds a migration.

The repo also carries two pieces of tooling that automate the safe ordering: a GitHub Actions pipeline (deploy.yml, §1.2) and a local command-line client (drk-deploy, §1.3). The Actions pipeline is written but not configured, and skips itself; the CLI works today.

1.1 The live path: Vercel Git integration + hand-applied migrations

What happens today
Trigger every push to main — a merge is a production deploy
Deployer Vercel's Git integration (deployments appear under the repo's Production environment, created by vercel[bot])
Migrations none — Vercel builds and promotes; it knows nothing about the database
Env read at build + runtime Vercel → Project → Settings → Environment Variables → Production

Because that path cannot migrate, the ordering contract is enforced by a person. It is a standing operator gate, not a suggestion:

A pull request that adds or changes a migration is applied to production FIRST and merged SECOND.

While the PR is still open, run pnpm db:app:migrate against the production DATABASE_URL (the direct/unpooled endpoint — §5), confirm it succeeded, and only then merge. A PR that changes src/db/migrations/better-auth-schema.sql (a better-auth upgrade, a new plugin or field) is a migration too: run pnpm db:auth:migrate the same way first. Migrations are additive, idempotent and ledgered, so the live build keeps serving happily against the migrated schema; merging first promotes a build that expects a schema the database does not have.

Run the gate from the PR's branch as it is pushed, with nothing uncommitted or untracked in the checkout. The runner ledgers each file under a checksum, so whatever it applies is what production keeps. If review later changes a file that was applied from a local edit, every later migrate against production aborts on the checksum mismatch (§8). drk-deploy migrate (§1.3) runs this gate with those checks enforced. It refuses a dirty or unpushed checkout, checks the URL against production's own settings, runs both migrators, and prints the commit it migrated from.

Getting that order wrong is verifiable without credentials: GET https://<domain>/api/health/ready returns 503 schema_behind while a live build is ahead of one of its migrations, and 200 ready once the ledger is complete (§4). Since F-26 that covers the Better Auth half as well: readiness asks Better Auth's own schema check, which compares every table and column the running configuration writes with what the database holds. Before anyone looks at the probe the symptom is 500s confined to the routes that touch the new column or table, or, for a Better Auth table or column, a 500 on every sign-in and every authenticated page. Migration 0004-oauth-client-secret-rotated-at.sql documents the worked case (review #43); the runbook entry is in Troubleshooting.

1.2 The Actions pipeline: optional, and not configured (DEPLOY-1)

deploy.yml encodes the migrate-first contract as CI. It runs when the CI workflow completes successfully for main (workflow_run trigger — a failed CI run never deploys) or on manual workflow_dispatch, and does, in order:

  1. pnpm db:auth:migrate, then pnpm db:app:migrate, against the direct (non-pooled) Neon endpoint (PRODUCTION_DIRECT_DATABASE_URL). The Better Auth step was missing until F-26. It runs with CI placeholder values for the variables @/lib/auth validates at load, because the Better Auth schema does not depend on them, so no production secret besides the database URL reaches it.
  2. vercel pull → vercel build --prod → vercel deploy --prebuilt --prod — builds in CI and promotes the prebuilt output.

The new build goes live only after migrations succeed; if step 1 fails, nothing is promoted and production keeps running the previous deployment.

None of that is in force. This repository holds no value for any of the workflow's four credentials, so its preflight job reports "not configured", the deploy job is skipped, and the run finishes green with the reason in its job summary (DEPLOY-1). Before that guard existed the job ran anyway and failed at the migrate step on every push from 2026-09-07 onward — a permanently red run on main that everyone had learned to ignore.

preflight reports three states, not two, and the middle one is the one to remember while you are adopting this path:

Credentials present preflight deploy The run
0 of 4 reports "not configured" skipped green — nobody has adopted this path, so there is nothing to warn about
1–3 of 4 fails, naming the missing ones skipped red — a partial set is evidence somebody expects deploys here, and a green run would hide that production is no longer being deployed
4 of 4 reports "configured" runs green if migrate-then-promote succeeded

A typo'd secret name, a rotated-to-empty value, or a secret added to an environment this job cannot see all land in the middle row. That is deliberate: once half two below is done, a silently-skipping pipeline means production stops receiving deploys with every check still green, and main auto-merges on green.

Turning the pipeline on is an operator decision with two halves, and doing one without the other is worse than doing neither:

  1. Add VERCEL_TOKEN, VERCEL_ORG_ID, VERCEL_PROJECT_ID and PRODUCTION_DIRECT_DATABASE_URL (§3).
  2. Turn Vercel's automatic production Git deploys OFF — Vercel → Project → Settings → Git: disable production-branch auto-deploy, or set an "Ignored Build Step". With both paths live, the Vercel half still promotes a build ahead of the migration, which is the exact race the pipeline exists to prevent. Nothing in vercel.json enforces this; it is a dashboard setting.

The output: "standalone" setting in next.config.mjs is for the Docker image only — Vercel ignores it; no action needed.

1.3 drk-deploy: the safe order, on demand

vercel-cli/README.md documents drk-deploy, a local client that encodes the same order — migrate → build → promote → verify — and needs no repository secrets. drk-deploy deploy applies migrations first and promotes only if they succeed; drk-deploy up syncs the environment first. It migrates only the database it is told to (--database-url, or PRODUCTION_DIRECT_DATABASE_URL in the shell or the --from-env file, never a shell's DATABASE_URL), and before migrating it pulls production's settings read-only and refuses that URL unless its host, port and database are production's DATABASE_URL (or DATABASE_URL_UNPOOLED), migrating production's DB_SCHEMA (F-47). It refuses a pooled connection string for migrations, never prints a secret value, and probes /api/health, /api/health/ready and a deliberately-wrong sign-in after promoting. A build that fails that probe is promoted back to the deployment production served before the run under --rollback-on-fail, the default under --yes (exit 4 when that restores a healthy production, 5 when it does not). Otherwise it stays live, exits 3, and the error prints the exact vercel promote command (F-51). If you want a deploy that is ordered end to end today, this is the path that gives it.

It releases only a clean, pushed commit, and prints which (F-49). Before anything writes, deploy, up and migrate read each checkout they release from (the kit, and a satellite's app folder when that is what is built) with read-only git, and refuse unless all of these hold:

  • git status --porcelain is empty, untracked files included. No flag overrides this. The one exception is a next-env.d.ts modified in the working tree. Every next build rewrites it, the one deploy runs included, so it is named and set aside rather than refused, and deploy puts it back after its build.
  • HEAD is pushed: a remote-tracking branch points at it.
  • For deploy and up, which promote: HEAD is origin's default branch (origin/HEAD, else origin/main), or the pushed ref named with --allow-ref <ref>.
  • For a satellite that owns its database: the kit checkout its migrations come from is at the kit's default branch, for deploy, up and migrate alike. No flag moves it.

The kit's own migrate may run from any pushed branch. That is how the §1.1 gate works: migrate from the open PR's branch, then merge. The gate is the kit's own, so it does not reach a satellite's database. Nothing is fetched. "Pushed" and "origin/main" are the remote-tracking refs as last fetched, so fetch first if the branch moved elsewhere. The commit, branch and tree state are printed before anything writes and again next to the database being migrated. That printed line is the only record of which commit a migration ran from: the ledger has no column for it. The promoted build's commit is recorded by Vercel.

One caveat that keeps §1.1 binding: running drk-deploy does not stop Vercel's Git integration from deploying the same push. While auto-deploy is on, a merge can still promote a build ahead of its migration no matter what you do afterwards, so the hand-migration gate stands until half two of §1.2 is done. drk-deploy deploy and up therefore read the project's git connection first, and refuse to migrate while it auto-deploys production (F-49), unless you pass --allow-git-integration-race. Auto-deploy counts as off only when no repository is connected, when vercel.json sets git.deploymentEnabled false for the production branch, or when the Ignored Build Step is exactly exit 0. This project's production is deployed by the git integration today, so on the live path use drk-deploy migrate from the PR's branch and let the merge promote.

1.4 Two environment stores

They are separate and serve different phases:

Store Holds Read by
Vercel env (Project → Settings → Environment Variables → Production) runtime + build vars (DATABASE_URL, auth secrets, feature flags, NEXT_PUBLIC_*, …) vercel build and the deployed functions at runtime — this is the live path
GitHub Actions secrets (the production GitHub Environment, or repository secrets) VERCEL_TOKEN, VERCEL_ORG_ID, VERCEL_PROJECT_ID, PRODUCTION_DIRECT_DATABASE_URL (optional variable DB_SCHEMA) the Actions pipeline only — unset today (§1.2)

2. One-time database bootstrap

No deploy path seeds. The Actions pipeline (§1.2) and drk-deploy (§1.3) apply both migrators on every deploy and the live Vercel Git path applies neither (§1.1), but none of them creates the baseline org, roles and first admin — you run the bootstrap once against a fresh database. Skipping it is the most common first-deploy mistake.

Run it from your own machine (laptop or a one-off runner) with the repo checked out, Node 24 / pnpm 10+, and the Neon direct (unpooled) connection string — it connects over the network; it is not SQL you paste into Neon's console.

pnpm install --frozen-lockfile

# Windows PowerShell: use `$env:NAME = "value"` instead of `export`.
export DATABASE_URL="postgresql://USER:PASSWORD@ep-xxxx.us-east-1.aws.neon.tech/neondb?sslmode=require"  # DIRECT / unpooled
export DB_SCHEMA="auth"
export SEED_ADMIN_EMAIL="you@example.com"
export SEED_ADMIN_PASSWORD="<a strong password>"

pnpm db:provision

pnpm db:provision (src/db/provision.ts) runs the full setup in order, fail-fast:

  1. db:auth:migrate — Better Auth tables (user/session/account/verification/rateLimit). On the live Vercel Git path nothing runs this for you: re-run it by hand, before merging, whenever src/db/migrations/better-auth-schema.sql changes in a release you are deploying (§1.1); the Actions pipeline and drk-deploy run it on every deploy. The rateLimit table (Better Auth's shared sign-in limiter store, review #199) was added this way. Better Auth does not tolerate a missing table or column: an instance that starts without it refuses every auth request, and keeps refusing until it restarts, even once the migration has run. /api/health/ready reports it as schema_behind (§4).
  2. db:app:migrate — extensions (pgcrypto, pg_trgm, created in public) + the core app schema, then the localized data under src/db/migrations/locales/. The Actions pipeline and drk-deploy re-run this on every deploy; on the Vercel Git path you run it by hand (§1.1).
  3. db:seed — the default org (the org flagged is_default is reused whatever its slug; the initial default org is created only when none is flagged, so a re-run never adds a second default, F-40), the admin.* permission catalog, baseline roles, and your first admin (from SEED_ADMIN_EMAIL / SEED_ADMIN_PASSWORD). The roles and the admin go into the platform org, the one holding the seeded superuser role (the original default), not into whichever org is flagged default now, so a re-run after a superadmin moved the default to a tenant writes nothing into that tenant. The demo satellite apps it registers for the local dev rig are skipped under NODE_ENV=production (opt in with SEED_DEMO_APPS=1); register real enterprise apps via the admin console instead.

Every step is idempotent and ledgered — applied migrations are recorded in app_schema_migrations, so re-running db:provision (or letting a deploy path re-run the migrators) only applies new work. All tables land in the schema named by DB_SCHEMA (default auth); extensions stay in public so every schema resolves them.

Notes:

  • English-only install: the email templates live under src/db/migrations/locales/, one file per locale, applied by default. Set DB_MIGRATE_LOCALES=0 (or false/no/off) to skip the localized files — the non-English email-template rows are then absent and those recipients fall back to English. The core schema and the English base (locales/0000-email-templates-en.sql, always applied) install regardless, so an English-only database still has every template.
  • db:seed is safe to re-run: every insert is on conflict do nothing, and its one update — relaxing the platform sign-up default from the migration's fail-closed admin_approval to auto_active — is first-run only. It is gated on the row never having been edited by an administrator (updated_by IS NULL; the admin API stamps it on every edit), so a re-run after you tightened the policy under Administrator → Platform sign-up defaults leaves it exactly as configured and prints [seed] platform sign-up policy left as configured (admin-managed). See Sign-up policy §5. In production, change the seeded admin password immediately or supply non-default SEED_ADMIN_* values.
  • The seed admin is provenance-gated (src/db/seeds/default-admin.ts). The seed fully escalates (verifies, activates, grants admin + admin.platform + superuser) only an account it creates itself in that run. On a re-run it recognises its own admin — the account is already email-verified and already holds superuser — and re-inserts any missing grants without touching anything else: a seed admin you have since blocked, suspended or deactivated stays that way (the run prints [seed] admin … left as configured (status=blocked)). Any other pre-existing account matching SEED_ADMIN_EMAIL — e.g. someone who self-registered that address before you bootstrapped, or an admin whose verification was revoked — makes the seed refuse ([seed] REFUSED to escalate pre-existing account …, exit code 1, nothing written). If that account really is yours, re-run with SEED_ADMIN_ADOPT_EXISTING=1 to confer the admin grants on it; even then its password, emailVerified flag and status are left as found, and a later plain re-run keeps refusing until the account is verified. Otherwise point SEED_ADMIN_EMAIL at an unregistered address.
  • Never run db:seed:dev against production — it creates 24 accounts (21 org-scoped + 3 cross-org members) sharing one weak password, three of them cross-tenant superusers. Two independent guards make this hard to do by accident (src/db/guards.ts): the seed refuses under NODE_ENV=production (override DEV_SEED_ALLOW_PROD=1), and — whatever NODE_ENV says, since it is routinely unset in a shell whose .env holds a production URL — it refuses any DATABASE_URL whose host is not local (localhost / 127.0.0.1 / ::1 / 0.0.0.0 / none; an unparseable URL counts as remote). Both checks run before a connection is opened, so a refusal writes nothing. The host guard is lifted only by --force or DEV_SEED_ALLOW_REMOTE=1; db:reset shares the same host check.

3. Vercel project + environment

Step 1 is the live path. Steps 2 and 3 are needed only if you are adopting the optional Actions pipeline (§1.2).

  1. From the repo root: vercel link (or import the repo in the dashboard). Framework preset: Next.js. Importing the repo is also what enables the Git integration that deploys production today (§1.1).
  2. (Actions pipeline only.) Capture the identifiers from .vercel/project.json → GitHub secrets: orgId → VERCEL_ORG_ID, projectId → VERCEL_PROJECT_ID. Create a deploy token (Vercel Account Settings → Tokens) → VERCEL_TOKEN.
  3. (Actions pipeline only.) In the repo, create the production GitHub Environment (Settings → Environments) and add the secrets above plus PRODUCTION_DIRECT_DATABASE_URL (Neon direct/unpooled URL). Repository secrets work too; what the environment buys is a narrower audience — an environment secret is readable only by a job that declares environment: production, but by any such job, in any workflow in this repository. It is scoped to the environment, not to this workflow: it fences the jobs that exist today (only deploy.yml names the environment) and not a workflow somebody adds tomorrow. Optionally add required reviewers so each deploy needs approval — with two consequences worth knowing before you turn it on. Both preflight and deploy name the environment (the DEPLOY-1 comment in deploy.yml explains why preflight must), so a real deploy is approved twice; and preflight asks on every successful CI run on main, including while the pipeline is unconfigured and the deploy would skip anyway. Doing this step before you have the credentials therefore buys an approval prompt per merge for a job that deploys nothing. And do not stop here: half two of §1.2 — turning Vercel's production auto-deploy off — belongs in the same sitting, or you now have two paths to production.

Set runtime env in Vercel (Production). Configuration is the authoritative list of every variable (≈60); set it there. Validation is at runtime, not build time — a missing required var will not fail next build. It fails the deployment when it starts: the Node boot hook (register() in src/instrumentation.ts) parses the whole schema first (F-26), so every server-rendered page and API route of that deployment answers 500, /api/health and /api/health/ready included, and the function log names each invalid variable and its rule. Set everything before sending real traffic. The deployment-critical must-set production secrets:

  • BETTER_AUTH_SECRET — strong random string (≥ 32 chars).
  • BETTER_AUTH_URL — https://<your-domain> (also set NEXT_PUBLIC_APP_URL and NEXT_PUBLIC_PRODUCTION_HOST to the same origin/host).
  • DATABASE_URL — see the endpoint decision below.
  • The three SSO_HANDOFF_* vars — SSO_HANDOFF_ISSUER (the primary's origin URL), SSO_HANDOFF_AUDIENCE_PREFIX, SSO_HANDOFF_APPLICATION_ID. Required even if you never use cross-app SSO — they are validated at boot. The prefix and application id can be placeholders, but SSO_HANDOFF_ISSUER must be an exact https:// origin (use BETTER_AUTH_URL's value when SSO is unused). If this deployment issues handoffs (it is the primary of a satellite fleet) also set SSO_HANDOFF_PRIVATE_KEY — an Ed25519 private JWK, generated per Configuration → SSO handoff, distinct from API_JWT_PRIVATE_KEY. Satellites never get it; they verify against https://<primary>/api/sso/jwks.json.

NODE_ENV=production is set by Vercel automatically. The NEXT_PUBLIC_* values are inlined at build time, so changing a domain requires a redeploy.

Operator gate before MCP_ENABLED goes on (review #57). While the gateway is enabled the env schema refuses an API_JWT_ISSUER that is not the same identifier as BETTER_AUTH_URL (and, since F-22, one with a trailing slash in any case, because the value is stamped into tokens as iss) — the OAuth discovery documents are served from BETTER_AUTH_URL and RFC 8414 requires the advertised issuer to be that location. The env schema is parsed when the deployment starts (F-26), so a deployment that already has both set to different values answers 500 on every server-rendered page and API route after the deploy rather than merely serving undiscoverable metadata. CI cannot catch this: next build parses placeholder env (buildPhasePlaceholders in src/lib/env.ts), so every check stays green and the failure appears only on the deployed instance. It is therefore a manual pre-deploy check — in Vercel → Settings → Environment Variables, for Production and Preview, confirm API_JWT_ISSUER is unset or equal to BETTER_AUTH_URL wherever MCP_ENABLED is set, and fix the env before the deploy. (The gateway is dark by default, so a deployment that never set MCP_ENABLED is unaffected.)

Operator gate for the origin and key rules (F-22). The env schema refuses, when the deployment starts (F-26), any origin-valued variable that is not an http(s) origin, an http:// one in production off loopback, a trailing slash on SSO_HANDOFF_ISSUER / API_JWT_ISSUER, a COOKIE_DOMAIN that does not cover BETTER_AUTH_URL or is written with a trailing dot, and an Ed25519 key of the wrong shape; the same boot hook refuses a key whose x is not d's public half (Configuration §1). The same blind spot applies as for the MCP gate: next build parses placeholders, so CI stays green and only the deployed instance fails. Vercel runs Preview with NODE_ENV=production too. Before deploying a build that carries these rules, check each of these for Production and Preview, reading sensitive values back from where they are observable when the dashboard cannot show them: BETTER_AUTH_URL, SSO_HANDOFF_ISSUER (a minted handoff's iss), ADMIN_TRUSTED_ORIGINS, COOKIE_DOMAIN (the Domain= of the sign-in Set-Cookie), API_JWT_ISSUER, MCP_DISPATCH_BASE_URL, MAILGUN_BASE_URL, and the four *_PRIVATE_KEY values. drk-deploy env:check covers only part of this gate (F-46). Of the values listed above it checks only BETTER_AUTH_URL and SSO_HANDOFF_ISSUER (and, on an Option C satellite, COOKIE_DOMAIN). For Production it reads each one back, applies the same rule the boot does, and compares it with the value the recorded config derives. It also reports any public value stored sensitive as a problem, because nothing can read it back. drk-deploy env:sync applies the same checks to every variable it leaves in place, so up stops on them too. Neither reads a secret, so the *_PRIVATE_KEY values are checked only when written (env:sync applies the key import to SSO_HANDOFF_PRIVATE_KEY). Everything else in this gate stays a manual check: ADMIN_TRUSTED_ORIGINS (outside the kit's contract; a satellite's is read back, but no origin rule is applied), the kit's own COOKIE_DOMAIN (outside its contract), API_JWT_ISSUER, MCP_DISPATCH_BASE_URL and MAILGUN_BASE_URL. And env:check covers Production only: drk-deploy env:sync --target preview --dry-run checks Preview.

DATABASE_URL: direct vs. pooled. Neon gives two connection strings for the same database. By default point DATABASE_URL at the direct/unpooled endpoint (no -pooler in the host). To use the pooled endpoint (better serverless concurrency), make the app pooler-compatible first — see §5. Keep ?sslmode=require on both. Migrations always use the direct endpoint.


4. Deploy + post-deploy verification

Apply any new migration to production first (§1.1), then push to main and let Vercel build and promote it — or, on the tooling paths, run drk-deploy deploy (§1.3) or the workflow from the Actions tab (§1.2). Once the deployment is live, verify:

  • GET https://<domain>/ → the landing page returns 200.
  • GET https://<domain>/api/health/ready → 200 {"status":"ready"}. This proves the environment passes its schema, the database is reachable, the ledger holds every core migration the live build depends on (REQUIRED_CORE_MIGRATIONS in src/db/migrations/migration-plan.ts) and Better Auth's own schema check finds every table and column its configuration writes (F-26). A 503 {"status":"unavailable","reason":"schema_behind"} means the build went live ahead of one of its migrations: the server log says which half. kind: "schema-behind" lists missing core ids — run pnpm db:app:migrate against production now. kind: "auth-schema-behind" lists missing Better Auth tables or columns — run pnpm db:auth:migrate, then redeploy, because an instance that already saw the gap keeps refusing auth until it restarts. A 503 config_invalid names the invalid variables in the log (kind: "config-invalid"). None of these details is ever in the response.
  • Sign in with the seed admin from §2; the session persists.
  • GET https://<domain>/api/internal/outbox-drain and /api/internal/mcp-registration-reap without the bearer header → 401 (confirms both cron endpoints are fail-closed).
  • In Neon's SQL editor, select id from auth.app_schema_migrations order by id lists the applied ids — the core 000N-*.sql files (0001-initial-schema.sql, 0002-admin-groups-permissions.sql, 0003-outbox-delivery-payload.sql, 0004-oauth-client-secret-rotated-at.sql, 0005-integrity-constraints.sql, 0006-rate-limit-buckets.sql, …), the always-applied locales/0000-email-templates-en.sql, and (unless DB_MIGRATE_LOCALES=0) the localized locales/0001-… files.
  • If Sentry is configured, trigger a test error and confirm it lands (Observability).
  • If METRICS_TOKEN is set, GET /api/metrics with Authorization: Bearer <token> returns Prometheus text.

The cron jobs. vercel.json declares two scheduled jobs: GET /api/internal/outbox-drain daily at 08:00 UTC retries pending rows in app_outbox (the serverless substitute for a long-running drain worker), and GET /api/internal/mcp-registration-reap daily at 08:30 UTC expires MCP self-registrations still pending after MCP_REGISTRATION_PENDING_TTL_DAYS (the substitute for pnpm mcp:reap). Vercel Cron calls both with Authorization: Bearer <CRON_SECRET>, so set CRON_SECRET in Vercel env or the routes return 401 and neither mail is retried nor junk registrations expired. Daily works on all Vercel plans (Hobby included, which allows two daily crons); higher frequencies need Pro. Note what "daily" means for token-bearing mail: a password-reset / verification token lives one hour, so a row whose inline attempt failed is already expired when the next tick runs. The drain fails such a row as token_expired instead of delivering a dead link (review #90) — the user simply requests a new one. If you need those retries to actually land, run the drain more often (Pro cron, an external scheduler, or pnpm outbox:drain). pnpm db:prune (token-revocation + audit/outbox retention) is not scheduled in this repo — add a cron or external caller if you want it.


5. Operations & gotchas

Rate limiting and instance count. Two limiters run in the application tier, with different topologies. The pre-auth floors — /api/v1/auth/token (per-IP, then global), /api/mcp/register (per-IP, then global), the CSP report sink (per-IP, then global), /api/sso/consume and a signed-out /api/sso/launch (per-IP only, no global floor; F-19) and /api/invitations/accept (per user) — and Better Auth's built-in sign-in / password-reset limiter keep their buckets in Postgres (app_rate_limits, migration 0006, and Better Auth's rateLimit table), so they enforce one budget across every instance or serverless invocation; if the database is unreachable the app floors fall back to a per-instance bucket and log a warning (counted in devresponsekit_rate_limit_shared_fallbacks_total). The per-actor abuse guard on admin mutations, bulk operations, CSV export and a signed-in SSO launch (src/lib/admin/rate-limit.server.ts) is in-process — its budget lives in one Node process's memory, resets on restart, and with more than one instance is enforced per instance (it effectively multiplies by the instance count). That guard layers on top of the real authorization checks and its fan-out is bounded by the credentials an actor holds, so multi-instance is a supported topology; only the per-actor UX limit is best-effort there. (This applies only to the application tier — Postgres is external and unaffected.)

Direct endpoint by default; pooled needs two changes. The app sends three per-connection startup parameters: search_path (-c search_path=…) plus the statement_timeout / idle_in_transaction_session_timeout ceilings from src/db/database.ts (pg puts those in the startup packet too). A transaction pooler rejects startup parameters — every connection fails with 08P01 unsupported startup parameter in options: search_path — along with the DDL + advisory locks the migrator needs. So both migrations and runtime use the direct/unpooled endpoint by default. To run the runtime on the pooled endpoint:

  1. Set all three as role defaults the pooler honors, once against the database (<app_role> is the user in your connection string, e.g. Neon's neondb_owner). The 30s values match what the code sends by default (PG_STATEMENT_TIMEOUT_MS / PG_IDLE_IN_TX_TIMEOUT_MS = 30000); mirror any override you set:

    ALTER ROLE <app_role> SET search_path = "auth", public;
    ALTER ROLE <app_role> SET statement_timeout = '30s';
    ALTER ROLE <app_role> SET idle_in_transaction_session_timeout = '30s';
  2. Set DB_SEARCH_PATH_VIA_OPTIONS=0 in Vercel so the app stops sending all three rejected parameters (review #20 — the flag used to strip only search_path, and the pooler rejected the timeouts just the same). Verify with show statement_timeout; on a pooled connection: it must read 30s, not 0.

Migrations still use the direct endpoint. Keep PGPOOL_MAX small on serverless (each function instance opens its own pool).

Shutdown on Vercel is a no-op; on a long-running server it is a two-step drain. Vercel never delivers SIGTERM to a warm function in the normal freeze/teardown path, and the app's shutdown watchdog (src/lib/shutdown.server.ts) registers nothing when the platform's VERCEL variable is set — pool connections are simply dropped when the function instance is recycled. On next start / the container (docs/docker.md §7), Next's own cleanup drains HTTP and exits 143/130; the watchdog only ends the pool and exits with the same code if that drain overruns SHUTDOWN_TIMEOUT_MS (review #24).

Schema changes ship as new numbered files in src/db/migrations/ — never edit an applied migration:

  • Core — 0001-initial-schema.sql is the frozen baseline; further schema changes are added as new numbered NNNN-*.sql files, applied in lexical order after it and recorded once each in the ledger.
  • Email templates — one file per locale — go in src/db/migrations/locales/. The English base locales/0000-email-templates-en.sql is ALWAYS applied (the fallback every locale resolves to); the localized files (locales/0001-…+) apply unless DB_MIGRATE_LOCALES=0. Ledger ids are path-prefixed (locales/<file>) so they can never collide with a core filename.

On the live path you apply them, against production, before the PR merges (§1.1). The tooling paths apply them migrate-first on the next deploy (§1.2, §1.3).

Rollback. Roll the app back by promoting a previous deployment in Vercel (dashboard → previous deployment → "Promote to Production", or vercel promote <deployment>). Prefer promoting to vercel rollback: after an Instant Rollback Vercel stops assigning production domains to new deployments until one is promoted, so on the live path (§1.1) the next merge is built and never goes live. Migrations are forward-only — additive, with no down-migrations — so the older build runs safely against the newer schema (forward-compatible by design). Migrations always land before the build that needs them — by hand on the live path (§1.1), by the pipeline on the tooling paths — so a rollback needs no DB change. A migration that genuinely must be reverted is authored as a new forward migration. To recover lost data (not a bad deploy), use your provider's PITR/snapshot, not a schema revert.

See Troubleshooting for operational issues.


6. Self-host / container

A production-ready multi-stage Dockerfile (built from the Next.js standalone output, non-root) is provided. Build/configure/run, running migrations as a separate init step, required env, a docker compose example, and hardening are all in Docker.


7. CI

CI is .github/workflows/ (source of truth). ci.yml runs on push + pull_request and validates quality and behavior — typecheck, lint, format, build, tests + coverage gate, DB-backed integration tests, Playwright e2e + accessibility, SDK/schema/doc-link drift checks, and the deploy CLI's own typecheck, build, tests and format check (vercel-cli/, F-45) — but does not itself deploy: deploy.yml fires only after this workflow succeeds on main, and then skips itself because its credentials are unset (§1.2). Nothing in CI deploys production today — Vercel's Git integration does, independently of CI's result (§1.1). Separate workflows run the pnpm audit hard gate over both lockfiles, the app's and vercel-cli/'s (dependency-audit.yml, which also reads the Dependabot alerts weekly), plus Trivy, CodeQL, gitleaks, and an advisory Stryker mutation-testing pass on the security core (mutation.yml). See Testing.


8. Least-privilege runtime role (optional, recommended)

By default the application connects as the same role that runs migrations and owns every table (Neon's neondb_owner, the local devresponse). That role can do anything to the schema — including deleting audit rows — so the audit log's append-only trigger is a guard against accidents, not a privilege boundary. Migration 0005-integrity-constraints.sql (review #83) splits the two:

  • Owner / migration role — whatever DATABASE_URL you run pnpm db:app:migrate / db:provision with. Owns the tables, the trigger and the SECURITY DEFINER retention function app_audit_events_prune(days, batch).
  • Runtime role <DB_SCHEMA>_runtime (auth_runtime by default) — created by 0005 as NOLOGIN with no password, holding USAGE on the schema, SELECT/INSERT/UPDATE/DELETE on every table except UPDATE/DELETE/TRUNCATE on app_audit_events (INSERT/SELECT only), and EXECUTE on the retention function. Default privileges are set so tables a later migration creates are covered automatically. The append-only trigger permits a DELETE only when the effective role is the table owner and the transaction-local app.audit_retention marker is on — both hold inside that function whoever calls it; the owner half never holds for the runtime role — so a stolen runtime credential cannot purge or rewrite audit history even by setting the marker itself. Until you switch, the app connects as the owner and a session that sets the marker can delete: the trigger guards against accidents, the role switch is the privilege boundary. The function itself clamps the requested window to a 30-day floor and each batch to 10 000 rows, so even through its one sanctioned path the runtime credential cannot purge recent history.

Nothing changes until you switch the app's connection string. To adopt it, once, against the direct endpoint as the owner role:

alter role auth_runtime login password '<a strong secret>';
-- Only if the runtime uses the POOLED endpoint (§5): the pooler drops the
-- startup search_path, so pin it on the role as well.
alter role auth_runtime set search_path = "auth", public;

Then set the runtime DATABASE_URL (Vercel → Environment Variables, or the container's env) to the same host/database with auth_runtime as the user, redeploy, and verify: sign in, open an admin page, and confirm select count(*) from auth.app_audit_events keeps growing. Keep the owner role's connection string for pnpm db:app:migrate / db:provision / db:seed / db:reset (the deploy pipeline and CI keep using it). pnpm db:prune works under either role — the retention job goes through the definer function.

If the migrating role lacks CREATEROLE (some managed providers), 0005 prints a NOTICE with the manual steps instead of failing: create the role yourself (create role auth_runtime nologin;) and re-run pnpm db:app:migrate — the grant block is idempotent and runs on any later pass. (Neon's neondb_owner can create roles.)

Ledger checksums (review #86). app_schema_migrations now records a sha256 checksum per applied file; the runner refuses to proceed if an applied file's hash differs from the ledger, printing the id and both hashes. Rows ledgered before this column existed are backfilled on the next run (logged as [migrate] backfilled checksum for …). The hash is of the file's normalised content (comments stripped, whitespace collapsed — normalizeMigrationSql), so the comment-only edits this repo deliberately makes to frozen files never move it; only a change to what an applied file does trips the check, and that is a bug to revert (restore the file from main, put the change in a new numbered file), not bookkeeping. Only a deliberate DDL change that was already applied by hand to that database needs the pin in tests/unit/migration-checksums.test.ts updated and the ledger row corrected with the exact update … set checksum = … statement the error prints. The runner also holds pg_advisory_lock(hashtext('app_schema_migrations')) for the whole run (review #85), so a redeploy racing a manual migrate serialises instead of colliding.


Next: Configuration · Docker · Troubleshooting