| title | Deployment |
|---|---|
| description | How the repo ships to production on Vercel and Neon, the DB bootstrap, and release checks. |
| group | General |
| order | 60 |
Audience: DevOps and release engineers. How this repo ships to production, the one-time database bootstrap, and how to verify a release.
This repo deploys to Vercel with a Neon serverless Postgres database. For the full environment-variable catalog see Configuration; for the self-host/container path see Docker.
Production ships through Vercel's own Git integration: every push to main is built and promoted by Vercel, and database migrations are applied by hand beforehand. That is the whole of the live path today — read §1.1 before merging anything that adds a migration.
The repo also carries two pieces of tooling that automate the safe ordering: a GitHub Actions pipeline (deploy.yml, §1.2) and a local command-line client (drk-deploy, §1.3). The Actions pipeline is written but not configured, and skips itself; the CLI works today.
| What happens today | |
|---|---|
| Trigger | every push to main — a merge is a production deploy |
| Deployer | Vercel's Git integration (deployments appear under the repo's Production environment, created by vercel[bot]) |
| Migrations | none — Vercel builds and promotes; it knows nothing about the database |
| Env read at build + runtime | Vercel → Project → Settings → Environment Variables → Production |
Because that path cannot migrate, the ordering contract is enforced by a person. It is a standing operator gate, not a suggestion:
A pull request that adds or changes a migration is applied to production FIRST and merged SECOND.
While the PR is still open, run
pnpm db:app:migrateagainst the productionDATABASE_URL(the direct/unpooled endpoint — §5), confirm it succeeded, and only then merge. A PR that changessrc/db/migrations/better-auth-schema.sql(a better-auth upgrade, a new plugin or field) is a migration too: runpnpm db:auth:migratethe same way first. Migrations are additive, idempotent and ledgered, so the live build keeps serving happily against the migrated schema; merging first promotes a build that expects a schema the database does not have.
Run the gate from the PR's branch as it is pushed, with nothing uncommitted or untracked in the checkout. The runner ledgers each file under a checksum, so whatever it applies is what production keeps. If review later changes a file that was applied from a local edit, every later migrate against production aborts on the checksum mismatch (§8). drk-deploy migrate (§1.3) runs this gate with those checks enforced. It refuses a dirty or unpushed checkout, checks the URL against production's own settings, runs both migrators, and prints the commit it migrated from.
Getting that order wrong is verifiable without credentials: GET https://<domain>/api/health/ready returns 503 schema_behind while a live build is ahead of one of its migrations, and 200 ready once the ledger is complete (§4). Since F-26 that covers the Better Auth half as well: readiness asks Better Auth's own schema check, which compares every table and column the running configuration writes with what the database holds. Before anyone looks at the probe the symptom is 500s confined to the routes that touch the new column or table, or, for a Better Auth table or column, a 500 on every sign-in and every authenticated page. Migration 0004-oauth-client-secret-rotated-at.sql documents the worked case (review #43); the runbook entry is in Troubleshooting.
deploy.yml encodes the migrate-first contract as CI. It runs when the CI workflow completes successfully for main (workflow_run trigger — a failed CI run never deploys) or on manual workflow_dispatch, and does, in order:
pnpm db:auth:migrate, thenpnpm db:app:migrate, against the direct (non-pooled) Neon endpoint (PRODUCTION_DIRECT_DATABASE_URL). The Better Auth step was missing until F-26. It runs with CI placeholder values for the variables@/lib/authvalidates at load, because the Better Auth schema does not depend on them, so no production secret besides the database URL reaches it.vercel pull→vercel build --prod→vercel deploy --prebuilt --prod— builds in CI and promotes the prebuilt output.
The new build goes live only after migrations succeed; if step 1 fails, nothing is promoted and production keeps running the previous deployment.
None of that is in force. This repository holds no value for any of the workflow's four credentials, so its preflight job reports "not configured", the deploy job is skipped, and the run finishes green with the reason in its job summary (DEPLOY-1). Before that guard existed the job ran anyway and failed at the migrate step on every push from 2026-09-07 onward — a permanently red run on main that everyone had learned to ignore.
preflight reports three states, not two, and the middle one is the one to remember while you are adopting this path:
| Credentials present | preflight |
deploy |
The run |
|---|---|---|---|
| 0 of 4 | reports "not configured" | skipped | green — nobody has adopted this path, so there is nothing to warn about |
| 1–3 of 4 | fails, naming the missing ones | skipped | red — a partial set is evidence somebody expects deploys here, and a green run would hide that production is no longer being deployed |
| 4 of 4 | reports "configured" | runs | green if migrate-then-promote succeeded |
A typo'd secret name, a rotated-to-empty value, or a secret added to an environment this job cannot see all land in the middle row. That is deliberate: once half two below is done, a silently-skipping pipeline means production stops receiving deploys with every check still green, and main auto-merges on green.
Turning the pipeline on is an operator decision with two halves, and doing one without the other is worse than doing neither:
- Add
VERCEL_TOKEN,VERCEL_ORG_ID,VERCEL_PROJECT_IDandPRODUCTION_DIRECT_DATABASE_URL(§3). - Turn Vercel's automatic production Git deploys OFF — Vercel → Project → Settings → Git: disable production-branch auto-deploy, or set an "Ignored Build Step". With both paths live, the Vercel half still promotes a build ahead of the migration, which is the exact race the pipeline exists to prevent. Nothing in
vercel.jsonenforces this; it is a dashboard setting.
The output: "standalone" setting in next.config.mjs is for the Docker image only — Vercel ignores it; no action needed.
vercel-cli/README.md documents drk-deploy, a local client that encodes the same order — migrate → build → promote → verify — and needs no repository secrets. drk-deploy deploy applies migrations first and promotes only if they succeed; drk-deploy up syncs the environment first. It migrates only the database it is told to (--database-url, or PRODUCTION_DIRECT_DATABASE_URL in the shell or the --from-env file, never a shell's DATABASE_URL), and before migrating it pulls production's settings read-only and refuses that URL unless its host, port and database are production's DATABASE_URL (or DATABASE_URL_UNPOOLED), migrating production's DB_SCHEMA (F-47). It refuses a pooled connection string for migrations, never prints a secret value, and probes /api/health, /api/health/ready and a deliberately-wrong sign-in after promoting. A build that fails that probe is promoted back to the deployment production served before the run under --rollback-on-fail, the default under --yes (exit 4 when that restores a healthy production, 5 when it does not). Otherwise it stays live, exits 3, and the error prints the exact vercel promote command (F-51). If you want a deploy that is ordered end to end today, this is the path that gives it.
It releases only a clean, pushed commit, and prints which (F-49). Before anything writes, deploy, up and migrate read each checkout they release from (the kit, and a satellite's app folder when that is what is built) with read-only git, and refuse unless all of these hold:
git status --porcelainis empty, untracked files included. No flag overrides this. The one exception is anext-env.d.tsmodified in the working tree. Everynext buildrewrites it, the onedeployruns included, so it is named and set aside rather than refused, anddeployputs it back after its build.- HEAD is pushed: a remote-tracking branch points at it.
- For
deployandup, which promote: HEAD is origin's default branch (origin/HEAD, elseorigin/main), or the pushed ref named with--allow-ref <ref>. - For a satellite that owns its database: the kit checkout its migrations come from is at the kit's default branch, for
deploy,upandmigratealike. No flag moves it.
The kit's own migrate may run from any pushed branch. That is how the §1.1 gate works: migrate from the open PR's branch, then merge. The gate is the kit's own, so it does not reach a satellite's database. Nothing is fetched. "Pushed" and "origin/main" are the remote-tracking refs as last fetched, so fetch first if the branch moved elsewhere. The commit, branch and tree state are printed before anything writes and again next to the database being migrated. That printed line is the only record of which commit a migration ran from: the ledger has no column for it. The promoted build's commit is recorded by Vercel.
One caveat that keeps §1.1 binding: running drk-deploy does not stop Vercel's Git integration from deploying the same push. While auto-deploy is on, a merge can still promote a build ahead of its migration no matter what you do afterwards, so the hand-migration gate stands until half two of §1.2 is done. drk-deploy deploy and up therefore read the project's git connection first, and refuse to migrate while it auto-deploys production (F-49), unless you pass --allow-git-integration-race. Auto-deploy counts as off only when no repository is connected, when vercel.json sets git.deploymentEnabled false for the production branch, or when the Ignored Build Step is exactly exit 0. This project's production is deployed by the git integration today, so on the live path use drk-deploy migrate from the PR's branch and let the merge promote.
They are separate and serve different phases:
| Store | Holds | Read by |
|---|---|---|
| Vercel env (Project → Settings → Environment Variables → Production) | runtime + build vars (DATABASE_URL, auth secrets, feature flags, NEXT_PUBLIC_*, …) |
vercel build and the deployed functions at runtime — this is the live path |
GitHub Actions secrets (the production GitHub Environment, or repository secrets) |
VERCEL_TOKEN, VERCEL_ORG_ID, VERCEL_PROJECT_ID, PRODUCTION_DIRECT_DATABASE_URL (optional variable DB_SCHEMA) |
the Actions pipeline only — unset today (§1.2) |
No deploy path seeds. The Actions pipeline (§1.2) and drk-deploy (§1.3) apply both migrators on every deploy and the live Vercel Git path applies neither (§1.1), but none of them creates the baseline org, roles and first admin — you run the bootstrap once against a fresh database. Skipping it is the most common first-deploy mistake.
Run it from your own machine (laptop or a one-off runner) with the repo checked out, Node 24 / pnpm 10+, and the Neon direct (unpooled) connection string — it connects over the network; it is not SQL you paste into Neon's console.
pnpm install --frozen-lockfile
# Windows PowerShell: use `$env:NAME = "value"` instead of `export`.
export DATABASE_URL="postgresql://USER:PASSWORD@ep-xxxx.us-east-1.aws.neon.tech/neondb?sslmode=require" # DIRECT / unpooled
export DB_SCHEMA="auth"
export SEED_ADMIN_EMAIL="you@example.com"
export SEED_ADMIN_PASSWORD="<a strong password>"
pnpm db:provisionpnpm db:provision (src/db/provision.ts) runs the full setup in order, fail-fast:
db:auth:migrate— Better Auth tables (user/session/account/verification/rateLimit). On the live Vercel Git path nothing runs this for you: re-run it by hand, before merging, wheneversrc/db/migrations/better-auth-schema.sqlchanges in a release you are deploying (§1.1); the Actions pipeline anddrk-deployrun it on every deploy. TherateLimittable (Better Auth's shared sign-in limiter store, review #199) was added this way. Better Auth does not tolerate a missing table or column: an instance that starts without it refuses every auth request, and keeps refusing until it restarts, even once the migration has run./api/health/readyreports it asschema_behind(§4).db:app:migrate— extensions (pgcrypto,pg_trgm, created inpublic) + the core app schema, then the localized data undersrc/db/migrations/locales/. The Actions pipeline anddrk-deployre-run this on every deploy; on the Vercel Git path you run it by hand (§1.1).db:seed— the default org (the org flaggedis_defaultis reused whatever its slug; the initialdefaultorg is created only when none is flagged, so a re-run never adds a second default, F-40), theadmin.*permission catalog, baseline roles, and your first admin (fromSEED_ADMIN_EMAIL/SEED_ADMIN_PASSWORD). The roles and the admin go into the platform org, the one holding the seededsuperuserrole (the original default), not into whichever org is flagged default now, so a re-run after a superadmin moved the default to a tenant writes nothing into that tenant. The demo satellite apps it registers for the local dev rig are skipped underNODE_ENV=production(opt in withSEED_DEMO_APPS=1); register real enterprise apps via the admin console instead.
Every step is idempotent and ledgered — applied migrations are recorded in app_schema_migrations, so re-running db:provision (or letting a deploy path re-run the migrators) only applies new work. All tables land in the schema named by DB_SCHEMA (default auth); extensions stay in public so every schema resolves them.
Notes:
- English-only install: the email templates live under
src/db/migrations/locales/, one file per locale, applied by default. SetDB_MIGRATE_LOCALES=0(orfalse/no/off) to skip the localized files — the non-English email-template rows are then absent and those recipients fall back to English. The core schema and the English base (locales/0000-email-templates-en.sql, always applied) install regardless, so an English-only database still has every template. db:seedis safe to re-run: every insert ison conflict do nothing, and its one update — relaxing the platform sign-up default from the migration's fail-closedadmin_approvaltoauto_active— is first-run only. It is gated on the row never having been edited by an administrator (updated_by IS NULL; the admin API stamps it on every edit), so a re-run after you tightened the policy under Administrator → Platform sign-up defaults leaves it exactly as configured and prints[seed] platform sign-up policy left as configured (admin-managed). See Sign-up policy §5. In production, change the seeded admin password immediately or supply non-defaultSEED_ADMIN_*values.- The seed admin is provenance-gated (
src/db/seeds/default-admin.ts). The seed fully escalates (verifies, activates, grantsadmin+admin.platform+superuser) only an account it creates itself in that run. On a re-run it recognises its own admin — the account is already email-verified and already holdssuperuser— and re-inserts any missing grants without touching anything else: a seed admin you have since blocked, suspended or deactivated stays that way (the run prints[seed] admin … left as configured (status=blocked)). Any other pre-existing account matchingSEED_ADMIN_EMAIL— e.g. someone who self-registered that address before you bootstrapped, or an admin whose verification was revoked — makes the seed refuse ([seed] REFUSED to escalate pre-existing account …, exit code 1, nothing written). If that account really is yours, re-run withSEED_ADMIN_ADOPT_EXISTING=1to confer the admin grants on it; even then its password,emailVerifiedflag and status are left as found, and a later plain re-run keeps refusing until the account is verified. Otherwise pointSEED_ADMIN_EMAILat an unregistered address. - Never run
db:seed:devagainst production — it creates 24 accounts (21 org-scoped + 3 cross-org members) sharing one weak password, three of them cross-tenant superusers. Two independent guards make this hard to do by accident (src/db/guards.ts): the seed refuses underNODE_ENV=production(overrideDEV_SEED_ALLOW_PROD=1), and — whateverNODE_ENVsays, since it is routinely unset in a shell whose.envholds a production URL — it refuses anyDATABASE_URLwhose host is not local (localhost/127.0.0.1/::1/0.0.0.0/ none; an unparseable URL counts as remote). Both checks run before a connection is opened, so a refusal writes nothing. The host guard is lifted only by--forceorDEV_SEED_ALLOW_REMOTE=1;db:resetshares the same host check.
Step 1 is the live path. Steps 2 and 3 are needed only if you are adopting the optional Actions pipeline (§1.2).
- From the repo root:
vercel link(or import the repo in the dashboard). Framework preset: Next.js. Importing the repo is also what enables the Git integration that deploys production today (§1.1). - (Actions pipeline only.) Capture the identifiers from
.vercel/project.json→ GitHub secrets:orgId→VERCEL_ORG_ID,projectId→VERCEL_PROJECT_ID. Create a deploy token (Vercel Account Settings → Tokens) →VERCEL_TOKEN. - (Actions pipeline only.) In the repo, create the
productionGitHub Environment (Settings → Environments) and add the secrets above plusPRODUCTION_DIRECT_DATABASE_URL(Neon direct/unpooled URL). Repository secrets work too; what the environment buys is a narrower audience — an environment secret is readable only by a job that declaresenvironment: production, but by any such job, in any workflow in this repository. It is scoped to the environment, not to this workflow: it fences the jobs that exist today (onlydeploy.ymlnames the environment) and not a workflow somebody adds tomorrow. Optionally add required reviewers so each deploy needs approval — with two consequences worth knowing before you turn it on. Bothpreflightanddeployname the environment (the DEPLOY-1 comment indeploy.ymlexplains whypreflightmust), so a real deploy is approved twice; andpreflightasks on every successful CI run onmain, including while the pipeline is unconfigured and the deploy would skip anyway. Doing this step before you have the credentials therefore buys an approval prompt per merge for a job that deploys nothing. And do not stop here: half two of §1.2 — turning Vercel's production auto-deploy off — belongs in the same sitting, or you now have two paths to production.
Set runtime env in Vercel (Production). Configuration is the authoritative list of every variable (≈60); set it there. Validation is at runtime, not build time — a missing required var will not fail next build. It fails the deployment when it starts: the Node boot hook (register() in src/instrumentation.ts) parses the whole schema first (F-26), so every server-rendered page and API route of that deployment answers 500, /api/health and /api/health/ready included, and the function log names each invalid variable and its rule. Set everything before sending real traffic. The deployment-critical must-set production secrets:
BETTER_AUTH_SECRET— strong random string (≥ 32 chars).BETTER_AUTH_URL—https://<your-domain>(also setNEXT_PUBLIC_APP_URLandNEXT_PUBLIC_PRODUCTION_HOSTto the same origin/host).DATABASE_URL— see the endpoint decision below.- The three
SSO_HANDOFF_*vars —SSO_HANDOFF_ISSUER(the primary's origin URL),SSO_HANDOFF_AUDIENCE_PREFIX,SSO_HANDOFF_APPLICATION_ID. Required even if you never use cross-app SSO — they are validated at boot. The prefix and application id can be placeholders, butSSO_HANDOFF_ISSUERmust be an exacthttps://origin (useBETTER_AUTH_URL's value when SSO is unused). If this deployment issues handoffs (it is the primary of a satellite fleet) also setSSO_HANDOFF_PRIVATE_KEY— an Ed25519 private JWK, generated per Configuration → SSO handoff, distinct fromAPI_JWT_PRIVATE_KEY. Satellites never get it; they verify againsthttps://<primary>/api/sso/jwks.json.
NODE_ENV=production is set by Vercel automatically. The NEXT_PUBLIC_* values are inlined at build time, so changing a domain requires a redeploy.
Operator gate before
MCP_ENABLEDgoes on (review #57). While the gateway is enabled the env schema refuses anAPI_JWT_ISSUERthat is not the same identifier asBETTER_AUTH_URL(and, since F-22, one with a trailing slash in any case, because the value is stamped into tokens asiss) — the OAuth discovery documents are served fromBETTER_AUTH_URLand RFC 8414 requires the advertised issuer to be that location. The env schema is parsed when the deployment starts (F-26), so a deployment that already has both set to different values answers 500 on every server-rendered page and API route after the deploy rather than merely serving undiscoverable metadata. CI cannot catch this:next buildparses placeholder env (buildPhasePlaceholdersinsrc/lib/env.ts), so every check stays green and the failure appears only on the deployed instance. It is therefore a manual pre-deploy check — in Vercel → Settings → Environment Variables, for Production and Preview, confirmAPI_JWT_ISSUERis unset or equal toBETTER_AUTH_URLwhereverMCP_ENABLEDis set, and fix the env before the deploy. (The gateway is dark by default, so a deployment that never setMCP_ENABLEDis unaffected.)
Operator gate for the origin and key rules (F-22). The env schema refuses, when the deployment starts (F-26), any origin-valued variable that is not an http(s) origin, an
http://one in production off loopback, a trailing slash onSSO_HANDOFF_ISSUER/API_JWT_ISSUER, aCOOKIE_DOMAINthat does not coverBETTER_AUTH_URLor is written with a trailing dot, and an Ed25519 key of the wrong shape; the same boot hook refuses a key whosexis notd's public half (Configuration §1). The same blind spot applies as for the MCP gate:next buildparses placeholders, so CI stays green and only the deployed instance fails. Vercel runs Preview withNODE_ENV=productiontoo. Before deploying a build that carries these rules, check each of these for Production and Preview, reading sensitive values back from where they are observable when the dashboard cannot show them:BETTER_AUTH_URL,SSO_HANDOFF_ISSUER(a minted handoff'siss),ADMIN_TRUSTED_ORIGINS,COOKIE_DOMAIN(theDomain=of the sign-inSet-Cookie),API_JWT_ISSUER,MCP_DISPATCH_BASE_URL,MAILGUN_BASE_URL, and the four*_PRIVATE_KEYvalues.drk-deploy env:checkcovers only part of this gate (F-46). Of the values listed above it checks onlyBETTER_AUTH_URLandSSO_HANDOFF_ISSUER(and, on an Option C satellite,COOKIE_DOMAIN). For Production it reads each one back, applies the same rule the boot does, and compares it with the value the recorded config derives. It also reports any public value storedsensitiveas a problem, because nothing can read it back.drk-deploy env:syncapplies the same checks to every variable it leaves in place, soupstops on them too. Neither reads a secret, so the*_PRIVATE_KEYvalues are checked only when written (env:syncapplies the key import toSSO_HANDOFF_PRIVATE_KEY). Everything else in this gate stays a manual check:ADMIN_TRUSTED_ORIGINS(outside the kit's contract; a satellite's is read back, but no origin rule is applied), the kit's ownCOOKIE_DOMAIN(outside its contract),API_JWT_ISSUER,MCP_DISPATCH_BASE_URLandMAILGUN_BASE_URL. Andenv:checkcovers Production only:drk-deploy env:sync --target preview --dry-runchecks Preview.
DATABASE_URL: direct vs. pooled. Neon gives two connection strings for the same database. By default point DATABASE_URL at the direct/unpooled endpoint (no -pooler in the host). To use the pooled endpoint (better serverless concurrency), make the app pooler-compatible first — see §5. Keep ?sslmode=require on both. Migrations always use the direct endpoint.
Apply any new migration to production first (§1.1), then push to main and let Vercel build and promote it — or, on the tooling paths, run drk-deploy deploy (§1.3) or the workflow from the Actions tab (§1.2). Once the deployment is live, verify:
-
GET https://<domain>/→ the landing page returns 200. -
GET https://<domain>/api/health/ready→ 200{"status":"ready"}. This proves the environment passes its schema, the database is reachable, the ledger holds every core migration the live build depends on (REQUIRED_CORE_MIGRATIONSinsrc/db/migrations/migration-plan.ts) and Better Auth's own schema check finds every table and column its configuration writes (F-26). A 503{"status":"unavailable","reason":"schema_behind"}means the build went live ahead of one of its migrations: the server log says which half.kind: "schema-behind"lists missing core ids — runpnpm db:app:migrateagainst production now.kind: "auth-schema-behind"lists missing Better Auth tables or columns — runpnpm db:auth:migrate, then redeploy, because an instance that already saw the gap keeps refusing auth until it restarts. A 503config_invalidnames the invalid variables in the log (kind: "config-invalid"). None of these details is ever in the response. - Sign in with the seed admin from §2; the session persists.
-
GET https://<domain>/api/internal/outbox-drainand/api/internal/mcp-registration-reapwithout the bearer header → 401 (confirms both cron endpoints are fail-closed). - In Neon's SQL editor,
select id from auth.app_schema_migrations order by idlists the applied ids — the core000N-*.sqlfiles (0001-initial-schema.sql,0002-admin-groups-permissions.sql,0003-outbox-delivery-payload.sql,0004-oauth-client-secret-rotated-at.sql,0005-integrity-constraints.sql,0006-rate-limit-buckets.sql, …), the always-appliedlocales/0000-email-templates-en.sql, and (unlessDB_MIGRATE_LOCALES=0) the localizedlocales/0001-…files. - If Sentry is configured, trigger a test error and confirm it lands (Observability).
- If
METRICS_TOKENis set,GET /api/metricswithAuthorization: Bearer <token>returns Prometheus text.
The cron jobs.
vercel.jsondeclares two scheduled jobs:GET /api/internal/outbox-draindaily at 08:00 UTC retriespendingrows inapp_outbox(the serverless substitute for a long-running drain worker), andGET /api/internal/mcp-registration-reapdaily at 08:30 UTC expires MCP self-registrations still pending afterMCP_REGISTRATION_PENDING_TTL_DAYS(the substitute forpnpm mcp:reap). Vercel Cron calls both withAuthorization: Bearer <CRON_SECRET>, so setCRON_SECRETin Vercel env or the routes return 401 and neither mail is retried nor junk registrations expired. Daily works on all Vercel plans (Hobby included, which allows two daily crons); higher frequencies need Pro. Note what "daily" means for token-bearing mail: a password-reset / verification token lives one hour, so a row whose inline attempt failed is already expired when the next tick runs. The drain fails such a row astoken_expiredinstead of delivering a dead link (review #90) — the user simply requests a new one. If you need those retries to actually land, run the drain more often (Pro cron, an external scheduler, orpnpm outbox:drain).pnpm db:prune(token-revocation + audit/outbox retention) is not scheduled in this repo — add a cron or external caller if you want it.
Rate limiting and instance count. Two limiters run in the application tier, with different topologies. The pre-auth floors — /api/v1/auth/token (per-IP, then global), /api/mcp/register (per-IP, then global), the CSP report sink (per-IP, then global), /api/sso/consume and a signed-out /api/sso/launch (per-IP only, no global floor; F-19) and /api/invitations/accept (per user) — and Better Auth's built-in sign-in / password-reset limiter keep their buckets in Postgres (app_rate_limits, migration 0006, and Better Auth's rateLimit table), so they enforce one budget across every instance or serverless invocation; if the database is unreachable the app floors fall back to a per-instance bucket and log a warning (counted in devresponsekit_rate_limit_shared_fallbacks_total). The per-actor abuse guard on admin mutations, bulk operations, CSV export and a signed-in SSO launch (src/lib/admin/rate-limit.server.ts) is in-process — its budget lives in one Node process's memory, resets on restart, and with more than one instance is enforced per instance (it effectively multiplies by the instance count). That guard layers on top of the real authorization checks and its fan-out is bounded by the credentials an actor holds, so multi-instance is a supported topology; only the per-actor UX limit is best-effort there. (This applies only to the application tier — Postgres is external and unaffected.)
Direct endpoint by default; pooled needs two changes. The app sends three per-connection startup parameters: search_path (-c search_path=…) plus the statement_timeout / idle_in_transaction_session_timeout ceilings from src/db/database.ts (pg puts those in the startup packet too). A transaction pooler rejects startup parameters — every connection fails with 08P01 unsupported startup parameter in options: search_path — along with the DDL + advisory locks the migrator needs. So both migrations and runtime use the direct/unpooled endpoint by default. To run the runtime on the pooled endpoint:
-
Set all three as role defaults the pooler honors, once against the database (
<app_role>is the user in your connection string, e.g. Neon'sneondb_owner). The30svalues match what the code sends by default (PG_STATEMENT_TIMEOUT_MS/PG_IDLE_IN_TX_TIMEOUT_MS= 30000); mirror any override you set:ALTER ROLE <app_role> SET search_path = "auth", public; ALTER ROLE <app_role> SET statement_timeout = '30s'; ALTER ROLE <app_role> SET idle_in_transaction_session_timeout = '30s';
-
Set
DB_SEARCH_PATH_VIA_OPTIONS=0in Vercel so the app stops sending all three rejected parameters (review #20 — the flag used to strip onlysearch_path, and the pooler rejected the timeouts just the same). Verify withshow statement_timeout;on a pooled connection: it must read30s, not0.
Migrations still use the direct endpoint. Keep PGPOOL_MAX small on serverless (each function instance opens its own pool).
Shutdown on Vercel is a no-op; on a long-running server it is a two-step drain. Vercel never delivers SIGTERM to a warm function in the normal freeze/teardown path, and the app's shutdown watchdog (src/lib/shutdown.server.ts) registers nothing when the platform's VERCEL variable is set — pool connections are simply dropped when the function instance is recycled. On next start / the container (docs/docker.md §7), Next's own cleanup drains HTTP and exits 143/130; the watchdog only ends the pool and exits with the same code if that drain overruns SHUTDOWN_TIMEOUT_MS (review #24).
Schema changes ship as new numbered files in src/db/migrations/ — never edit an applied migration:
- Core —
0001-initial-schema.sqlis the frozen baseline; further schema changes are added as new numberedNNNN-*.sqlfiles, applied in lexical order after it and recorded once each in the ledger. - Email templates — one file per locale — go in
src/db/migrations/locales/. The English baselocales/0000-email-templates-en.sqlis ALWAYS applied (the fallback every locale resolves to); the localized files (locales/0001-…+) apply unlessDB_MIGRATE_LOCALES=0. Ledger ids are path-prefixed (locales/<file>) so they can never collide with a core filename.
On the live path you apply them, against production, before the PR merges (§1.1). The tooling paths apply them migrate-first on the next deploy (§1.2, §1.3).
Rollback. Roll the app back by promoting a previous deployment in Vercel (dashboard → previous deployment → "Promote to Production", or vercel promote <deployment>). Prefer promoting to vercel rollback: after an Instant Rollback Vercel stops assigning production domains to new deployments until one is promoted, so on the live path (§1.1) the next merge is built and never goes live. Migrations are forward-only — additive, with no down-migrations — so the older build runs safely against the newer schema (forward-compatible by design). Migrations always land before the build that needs them — by hand on the live path (§1.1), by the pipeline on the tooling paths — so a rollback needs no DB change. A migration that genuinely must be reverted is authored as a new forward migration. To recover lost data (not a bad deploy), use your provider's PITR/snapshot, not a schema revert.
See Troubleshooting for operational issues.
A production-ready multi-stage Dockerfile (built from the Next.js standalone output, non-root) is provided. Build/configure/run, running migrations as a separate init step, required env, a docker compose example, and hardening are all in Docker.
CI is .github/workflows/ (source of truth). ci.yml runs on push + pull_request and validates quality and behavior — typecheck, lint, format, build, tests + coverage gate, DB-backed integration tests, Playwright e2e + accessibility, SDK/schema/doc-link drift checks, and the deploy CLI's own typecheck, build, tests and format check (vercel-cli/, F-45) — but does not itself deploy: deploy.yml fires only after this workflow succeeds on main, and then skips itself because its credentials are unset (§1.2). Nothing in CI deploys production today — Vercel's Git integration does, independently of CI's result (§1.1). Separate workflows run the pnpm audit hard gate over both lockfiles, the app's and vercel-cli/'s (dependency-audit.yml, which also reads the Dependabot alerts weekly), plus Trivy, CodeQL, gitleaks, and an advisory Stryker mutation-testing pass on the security core (mutation.yml). See Testing.
By default the application connects as the same role that runs migrations and owns every table (Neon's neondb_owner, the local devresponse). That role can do anything to the schema — including deleting audit rows — so the audit log's append-only trigger is a guard against accidents, not a privilege boundary. Migration 0005-integrity-constraints.sql (review #83) splits the two:
- Owner / migration role — whatever
DATABASE_URLyou runpnpm db:app:migrate/db:provisionwith. Owns the tables, the trigger and theSECURITY DEFINERretention functionapp_audit_events_prune(days, batch). - Runtime role
<DB_SCHEMA>_runtime(auth_runtimeby default) — created by 0005 asNOLOGINwith no password, holdingUSAGEon the schema,SELECT/INSERT/UPDATE/DELETEon every table exceptUPDATE/DELETE/TRUNCATEonapp_audit_events(INSERT/SELECTonly), andEXECUTEon the retention function. Default privileges are set so tables a later migration creates are covered automatically. The append-only trigger permits aDELETEonly when the effective role is the table owner and the transaction-localapp.audit_retentionmarker is on — both hold inside that function whoever calls it; the owner half never holds for the runtime role — so a stolen runtime credential cannot purge or rewrite audit history even by setting the marker itself. Until you switch, the app connects as the owner and a session that sets the marker can delete: the trigger guards against accidents, the role switch is the privilege boundary. The function itself clamps the requested window to a 30-day floor and each batch to 10 000 rows, so even through its one sanctioned path the runtime credential cannot purge recent history.
Nothing changes until you switch the app's connection string. To adopt it, once, against the direct endpoint as the owner role:
alter role auth_runtime login password '<a strong secret>';
-- Only if the runtime uses the POOLED endpoint (§5): the pooler drops the
-- startup search_path, so pin it on the role as well.
alter role auth_runtime set search_path = "auth", public;Then set the runtime DATABASE_URL (Vercel → Environment Variables, or the container's env) to the same host/database with auth_runtime as the user, redeploy, and verify: sign in, open an admin page, and confirm select count(*) from auth.app_audit_events keeps growing. Keep the owner role's connection string for pnpm db:app:migrate / db:provision / db:seed / db:reset (the deploy pipeline and CI keep using it). pnpm db:prune works under either role — the retention job goes through the definer function.
If the migrating role lacks CREATEROLE (some managed providers), 0005 prints a NOTICE with the manual steps instead of failing: create the role yourself (create role auth_runtime nologin;) and re-run pnpm db:app:migrate — the grant block is idempotent and runs on any later pass. (Neon's neondb_owner can create roles.)
Ledger checksums (review #86). app_schema_migrations now records a sha256 checksum per applied file; the runner refuses to proceed if an applied file's hash differs from the ledger, printing the id and both hashes. Rows ledgered before this column existed are backfilled on the next run (logged as [migrate] backfilled checksum for …). The hash is of the file's normalised content (comments stripped, whitespace collapsed — normalizeMigrationSql), so the comment-only edits this repo deliberately makes to frozen files never move it; only a change to what an applied file does trips the check, and that is a bug to revert (restore the file from main, put the change in a new numbered file), not bookkeeping. Only a deliberate DDL change that was already applied by hand to that database needs the pin in tests/unit/migration-checksums.test.ts updated and the ledger row corrected with the exact update … set checksum = … statement the error prints. The runner also holds pg_advisory_lock(hashtext('app_schema_migrations')) for the whole run (review #85), so a redeploy racing a manual migrate serialises instead of colliding.
Next: Configuration · Docker · Troubleshooting