Skip to content

15 · Platform reliability — environments, releases, monitoring, recovery

How changes reach production and how we notice and recover when something breaks. These rules protect every other rule file: a correct migration applied by hand to the wrong project, or a backup nobody has ever restored, loses money and data just as surely as a bug.

Findings that triggered this file (2026-09-17)

  • npx tsc --noEmit checked nothing: the root tsconfig.json has "files": [], and tsconfig.app.json had an ignoreDeprecations value TypeScript 5.9 rejects, so type checking aborted before the first file. 41 real type errors had accumulated, including a filter control that silently never rendered and realised-FX gain/loss journal lines that read a non-existent amount field (so they always post zero).
  • CI ran lint, unit tests, build and e2e — no type check, no database tests, no edge-function type check, no migration checks. Two migrations had already shared a version number once (commit 70c8db1).
  • There is no staging environment: migrations and edge functions go straight from a laptop to production.
  • Production is on the Supabase Free plan: no point-in-time recovery. The only recovery path is the daily logical dump (supabase-backup.yml), which has never been test-restored.
  • Edge functions read ~40 secrets; nothing lists which ones a function needs, and several fail open (e.g. whatsapp-webhook skips signature checks when WHATSAPP_APP_SECRET is missing).
  • No uptime check or alerting; errors are noticed when staff complain.

Environments

PLT-001 · Staging before production

Status: PROPOSED · Owner: IT Every migration and edge-function change is applied to a staging Supabase environment and smoke-tested (PLT-032) before production. Staging is either a Supabase branch (Pro plan) or a separate free project with the same migrations, functions, secrets (test values) and auth settings. Staging never holds real customer data.

PLT-002 · Production on a plan with point-in-time recovery

Status: PROPOSED · Owner: Management Production runs on Supabase Pro (or higher) with PITR enabled, giving a recovery point measured in minutes instead of up to 24 hours. Until then, PLT-040 daily dumps are the only recovery and the risk is accepted in decisions-log.md.

PLT-003 · Local database matches production major version

Status: PROPOSED · Owner: IT supabase/config.toml [db] major_version matches the production project's Postgres major version (Dashboard → Settings → Infrastructure), so local and CI database tests exercise the same engine. (CI currently runs on the configured version 15; the migrations and DB tests also pass on 17.)

Releases

PLT-010 · Migrations are applied by the pipeline, never by hand

Status: PROPOSED · Owner: IT Production migrations are applied only by a CI workflow (supabase db push) that runs after merge to main, requires a manual approval (GitHub Environment production with required reviewers), takes a fresh backup first (PLT-041) and records the run. Nobody runs supabase db push, psql DDL or dashboard SQL against production by hand except during a declared incident, which is written up afterwards.

PLT-011 · Migrations are forward-only and idempotent

Status: DECIDED (existing, CLAUDE.md) · Owner: IT A migration merged to main is never edited, renamed or deleted; fixes are new migrations. Every statement is safe to run twice (IF NOT EXISTS, CREATE OR REPLACE, DROP … IF EXISTS before CREATE POLICY). Filenames are YYYYMMDDHHMMSS_description.sql and each version is unique. Enforced by: scripts/check-migrations.mjs (CI step + pretest) — errors on bad names/duplicate versions, warns on unguarded CREATEs.

PLT-012 · Edge function deploy checklist

Status: PROPOSED · Owner: IT Before deploying an edge function: (1) deno check passes; (2) every secret it reads is set in the target project (see the table in deploy runbook); (3) verify_jwt in supabase/config.toml matches the function's own auth (functions with verify_jwt = false must call guardRequest or verify a signature); (4) any migration it depends on is already applied; (5) it is smoke-tested on staging; (6) the deploy is recorded (function, commit, who, when).

PLT-013 · Security-relevant secrets fail closed

Status: PROPOSED · Owner: IT A function whose security depends on a secret (webhook signing secret, CAPTCHA secret, dispatch key) refuses the request with 503 when the secret is missing, instead of skipping the check. Currently fails open: whatsapp-webhook (WHATSAPP_APP_SECRET).

PLT-014 · Frontend and backend ship in a safe order

Status: PROPOSED · Owner: IT Database changes ship first and are backward-compatible with the frontend currently live; the frontend that depends on them ships after. Destructive changes (drop column, tighten RLS) ship only once no live frontend uses the old shape.

PLT-015 · No third-party secret in the browser bundle

Status: PROPOSED · Owner: IT A key that lets someone spend our quota or act as us at an outside provider is an edge-function secret (supabase secrets set), never a VITE_* variable: everything prefixed VITE_ is inlined into the public JS bundle. The browser calls an edge function that holds the key, checks the caller is signed-in staff with the permission of the screen that uses it, and returns only the answer. Only public identifiers stay in VITE_* (the Supabase URL and anon key, the Turnstile site key, the Sentry DSN, the VAPID public key, the Google OAuth client id). A key found in a bundle is treated as leaked and rotated at the provider. First case (2026-09-26): the AirLabs flight-schedule key was VITE_AIRLABS_API_KEY and shipped in the bundle. It is now the secret AIRLABS_API_KEY, read only by the flight-lookup function (inventory.view). Browser code reads each Vite variable by name (import.meta.env.VITE_X). Handing import.meta.env around as an object makes Vite inline every VITE_* variable of the build into the bundle, read or not; until 2026-09-26 observability.ts, ErrorBoundary.tsx, gstnApiClient.ts and api.ts did this, so any variable set in Cloudflare shipped. A variable that is no longer read is also deleted from Cloudflare Pages and GitHub Actions. Enforced by: supabase/functions/flight-lookup/, src/lib/flightLookup.ts, src/lib/flightLookup.test.ts (no key, no AirLabs call in the browser), src/test/viteEnvByName.test.ts (no whole-object use of import.meta.env). Not built: a CI check that scans dist/ for provider keys. VITE_OXR_API_KEY and VITE_EXCHANGE_RATE_API_KEY (exchange rates) still ship in the bundle and have not been moved yet.

CI gates

PLT-020 · Required checks before merge

Status: PROPOSED · Owner: IT A PR can merge to main only when these pass:

Gate Command Runs
Migration guard node scripts/check-migrations.mjs every run
Lint npm run lint every run
Type check (ratchet) npm run typecheck:ratchet every run
Unit tests npm run test:run every run
Build + bundle budgets npm run build + scripts/check-bundle-budgets.mjs every run
DB behaviour tests scripts/db-test.sh after supabase start migrations / supabase/tests changed
Edge functions scripts/deno-check-functions.sh supabase/functions changed
End-to-end npm run e2e every non-Dependabot run

Enforced by: .github/workflows/ci.yml; branch protection on main must list these jobs as required.

PLT-021 · The type-error baseline only shrinks

Status: PROPOSED · Owner: IT scripts/typecheck-baseline.txt lists known type errors. New errors fail CI. Nobody adds lines to the baseline to get a PR green; a PR that fixes errors regenerates it (npm run typecheck:ratchet -- --update). Target: empty.

PLT-022 · Every database rule has a database test

Status: PROPOSED · Owner: IT A rule enforced in the database (RLS, trigger, constraint, RPC) has a test in supabase/tests/*.sql that proves it, named with the rule ID, wrapped in BEGIN … ROLLBACK, ending with \echo ALL … PASSED.

Monitoring

PLT-030 · Errors are captured and alert someone

Status: PROPOSED · Owner: IT Frontend errors go to Sentry (VITE_SENTRY_DSN, already wired) with release tags; edge-function errors are logged with a correlation ID. A new error type or an error-rate spike notifies the on-call person (email/WhatsApp) within 15 minutes.

PLT-031 · Uptime and database health checks

Status: PROPOSED · Owner: IT An external uptime check (e.g. Better Stack, UptimeRobot, Cloudflare health check) hits the app URL and a lightweight health endpoint every 5 minutes and alerts after 2 failures. Supabase usage/log alerts are configured for database CPU, disk, connection count and auth/5xx error spikes; Free-plan quota warnings (database size, egress, function invocations) go to the same inbox.

PLT-032 · Smoke test after every production change

Status: PROPOSED · Owner: Operations After any migration or function deploy, someone runs the smoke checklist in the deploy runbook (staff login, customer sign-up, partner sign-up, finance PIN, create booking, record payment) and records the result.

PLT-033 · A fallback that is taken is reported

Status: DECIDED · Owner: IT · Enforced by: src/lib/degraded.ts, src/test/noSilentCatch.test.ts Where the API layer survives a failure — a display name it cannot look up, a default prefix, a best-effort mirror — it reports that it fell back, through degraded(): to Sentry (PLT-030) and to the browser console, naming the place and why the fallback is safe. It reports at most three times per place per page load. A failure that would make a number wrong (a price, a tax base, a balance) does not fall back: it fails the request, and the person sees the error. catch {} and .catch(() => null) in src/lib/api.ts fail the build.

Found in the V1 audit (2026-09-24): 78 such sites, including the two head-count lookups behind the GST base on the ground margin (PAX-034), which fell back to the plain head count — the exact over-charge PAX-034 exists to prevent. Those now fail loudly.

Backups and recovery

PLT-040 · Daily offsite backup

Status: DECIDED (existing) · Owner: IT supabase-backup.yml dumps roles, schema and data daily at 02:00 UTC, with checksums, kept 30 days as artifacts (optionally S3).

PLT-041 · Backup before every risky change

Status: PROPOSED · Owner: IT A fresh backup (workflow_dispatch of supabase-backup.yml) completes successfully — and its artifact is downloaded or confirmed — before any production migration, bulk data fix or restore.

PLT-042 · Monthly restore test

Status: PROPOSED · Owner: IT Once a month the latest backup is restored into a scratch project or local Supabase, row counts of key tables (Booking, Payment, JournalEntry, JournalLine, Customer, User) are compared with production, and scripts/db-test.sh runs against it. Duration (actual RTO) and any problems are recorded. A backup that has not been restored is not counted as a backup.

PLT-043 · Rollback plan written before release

Status: PROPOSED · Owner: IT Every production release notes how to undo it: for code, the previous frontend deploy / function commit; for migrations, the forward "undo" migration or the restore point. Irreversible data migrations are called out explicitly and need management sign-off.

List completeness

A list screen is a promise: "this is what exists". The rules below make that promise testable. They exist because an audit found fetchBookings() asking for one 300-row page and six screens — the ops approval queue, the corrections queue, finance's payment workflow, airline seat reconciliation, the group booking picker and the sidebar badge — treating that page as "all bookings". Nothing failed; the queues were simply short.

PLT-050 · A list never presents one page as the whole set

Status: PROPOSED · Owner: IT Every list endpoint returns the shared paged shape (src/lib/paging.ts: { data, nextCursor, total?, totalIsEstimate? }) and every caller either pages to exhaustion or states, in the UI, that it is showing a page of a larger set. Page size is clamped to MAX_PAGE_SIZE (200); a caller asking for more gets 200, not a silent truncation at PostgREST's max_rows. Paging is keyset (cursor on the sort column plus id), so inserts during a walk cannot make a row appear twice or be skipped. A helper that walks every page (exportAll) throws when it hits its ceiling — a short export is never returned as if it were complete. Enforced by: 20260920130000_bookings_paging_search.sql (booking_page); tests: src/lib/exportAll.test.ts, src/lib/api.bookingsPaging.test.ts, supabase/tests/bookings_paging.sql.

PLT-051 · Filtering, searching and counting happen in the database

Status: PROPOSED · Owner: IT Filters and free-text search run server-side against the whole table, not over the rows the browser happens to hold. Counters and KPI tiles come from a count query with the same filters as the list — never from rows.length of a loaded page. Exact counts are opt-in (withTotal=1); the default is a planner estimate, because an exact count is a full scan. The bookings tiles (DECIDED 2026-10-02, owner, through Hamid). Cancelled counts fully cancelled bookings only; partly cancelled and transferred bookings keep their own counts, because their travellers still travel or have moved. Each tile counts exactly what its Status filter word lists (bkg_status_set, 20261008140000). Enforced by: booking_page / booking_stats in 20260920130000_bookings_paging_search.sql; test: supabase/tests/bookings_paging.sql. A tile and the filter word with the same name count the same bookings: each status word (pending, awaiting_approval, cancelled, …) is defined once, in bkg_status_set(), and both the Status filter (bkg_status_db) and the tiles (booking_stats) read it (20261008140000_booking_screens.sql; test: supabase/tests/booking_screens.sql). Every sidebar badge comes from badge_counts() (20260927194500), one call, each count taken the way its list decides what to show and under the caller's own row policies; test: supabase/tests/every_badge_counted.sql.

PLT-052 · Bulk actions resolve their target set on the server

Status: PROPOSED · Owner: IT A bulk action carries either an explicit id list or a filter plus exclusions — never "the rows currently on screen". When it carries a filter, the server resolves the set, and the confirmation (UX-001, level typed for cancel and delete) shows the count the server returned, not the number of visible rows. The action executes in batched server calls with the permission checked server-side per batch; it is not fanned out one HTTP request per row from the browser. Enforced by: POST /sales/bookings/bulk-resolve and booking_filter_ids; tests: src/lib/api.bookingsPaging.test.ts, supabase/tests/bookings_paging.sql.

PLT-053 · Id batches sent to the database are bounded

Status: PROPOSED · Owner: IT Any .in('<column>', ids) built from a caller-supplied or query-derived list is chunked at ID_BATCH_SIZE (100). An unbounded list grows the request URI with the data set and fails wholesale once it is too long, which reads as a broken page rather than a too-large query. Enforced by: src/lib/chunk.ts; test: src/lib/chunk.test.ts.

The same audit found the problem is not only about bookings. PostgREST runs with max_rows = 1000 (supabase/config.toml), so every list read without an explicit window was silently truncated there — no error, no signal — and the customers, visa, ticketing, requests, leads, quotations, group-roster and portal screens all filtered, sorted and counted that capped page in the browser. The audit log was worse: a hard .limit(500) with no offset or cursor, so compliance evidence older than the newest 500 events was simply unreachable. The three rules below add what PLT-050 … PLT-053 do not already cover.

PLT-056 · Keyset cursors, whitelisted sorts, clamped limits

Status: PROPOSED · Owner: IT The sort column a cursor walks is NOT NULL and comes from the endpoint's whitelist; any other sort value is a 400, never silently ignored, so a sort key can never name a column the caller was not meant to order (or probe) by. limit is clamped to MAX_PAGE_SIZE (src/lib/paging.ts, shared byte-for-byte across workstreams) and a nonsense limit falls back to DEFAULT_PAGE_SIZE. A malformed cursor is rejected, never treated as "start from the top". A read that still uses the legacy unpaged shape logs a breadcrumb when it comes back at the max_rows cap (warnIfTruncated()), and no .limit(N) above 1000 exists anywhere. Enforced by: src/lib/api.paging.test.ts, supabase/tests/list_paging.sql.

PLT-057 · A capped section says what it is hiding

Status: PROPOSED · Owner: IT Where a section keeps a fixed cap for good reason — the Customer 360 requests (50) and messages (20) tabs, a group's "needs attention" rollup, an incident list — it returns the total and a way to page the rest, so nothing disappears without the reader knowing. A rollup that aggregates an unbounded set (the journey board aggregated every group and booking in a 60-day window) takes its window on the outer keyset before the expensive aggregation runs. Enforced by: 20260920120300_list_paging_rpcs.sql (get_customer_360_tab_meta, get_customer_360_tab, get_journey_board_page, list_incidents_page); test: supabase/tests/list_paging.sql.

PLT-058 · A cursor is a position, never a capability

Status: PROPOSED · Owner: IT A cursor carries only a sort value and a row id the caller was already shown. It grants nothing: every page re-applies the same tenant scope (agentId, customerId) and the same RLS, so a cursor minted by one partner selects nothing for another (ACC-010). Cursors are never signed, trusted, or used in place of a permission check, and a filter value is quoted into the PostgREST predicate so a comma or bracket in a customer's name cannot be read as filter syntax. Enforced by: src/lib/api.paging.test.ts (partner cursor replay), supabase/tests/list_paging.sql.

Apps

PLT-060 · The Android app is a shell around the live site

Status: DECIDED 2026-09-25 (owner: "the app should have an APK version also") · Owner: IT The Android app (in.alhudatravels.app) opens https://alhudatravels.in/m — the phone layout of the ERP (docs/features/mobile-app.md), the same routes and functions as the desktop, gated the same way — in a full-screen WebView and bundles no copy of the app. A phone-sized browser lands on /m as well, with a way back to the desktop layout. Notifications inside the app go through Firebase Cloud Messaging, since a WebView has no Web Push; browsers keep Web Push. A release of the web app is therefore a release of the Android app: nothing is rebuilt or reinstalled, and a phone never runs an old version. The APK is built by a hand-run workflow, never by CI; a release APK is signed only with the company keystore held as GitHub secrets, never with a throwaway key. Enforced by: capacitor.config.ts (server.url), android/, .github/workflows/android-apk.yml, docs/operations/android-app.md; the /m routes in src/App.tsx and src/pages/mobile/; src/lib/nativePush.ts and the FCM path in push-send. Not built: a Play Store listing, iOS, app-links.

PLT-061 · The native app takes JavaScript fixes over the air

Status: PROPOSED 2026-09-26 · Owner: IT The native app (apps/mobile) receives JavaScript-only changes through EAS Update, on the production channel, without a reinstall. An update is published only from main, after the change is merged, from the machine that holds apps/mobile/.env; the publish script refuses when any EXPO_PUBLIC_* value is missing. An update is offered only to APKs built with the same version in app.json (runtime version policy appVersion); any native change — a native package, an SDK upgrade, a permission, anything in app.json or app.config.js — raises version and ships as a new APK. A phone checks on each cold start without waiting and runs a downloaded update on the next cold start, so a phone without signal opens as before. Enforced by: apps/mobile/app.json (runtimeVersion, updates), apps/mobile/app.config.js (channel), apps/mobile/eas.json, apps/mobile/scripts/check-update-env.js, docs/operations/android-app.md → Updates without a new APK. Not built: signed updates, iOS updates, an in-app restart prompt.