15 · Platform reliability — environments, releases, monitoring, recovery
How changes reach production and how we notice and recover when something breaks. These rules protect every other rule file: a correct migration applied by hand to the wrong project, or a backup nobody has ever restored, loses money and data just as surely as a bug.
Findings that triggered this file (2026-09-17)
npx tsc --noEmitchecked nothing: the roottsconfig.jsonhas"files": [], andtsconfig.app.jsonhad anignoreDeprecationsvalue TypeScript 5.9 rejects, so type checking aborted before the first file. 41 real type errors had accumulated, including a filter control that silently never rendered and realised-FX gain/loss journal lines that read a non-existentamountfield (so they always post zero).- CI ran lint, unit tests, build and e2e — no type check, no database tests, no edge-function type check, no migration checks. Two migrations had already shared a version number once (commit
70c8db1). - There is no staging environment: migrations and edge functions go straight from a laptop to production.
- Production is on the Supabase Free plan: no point-in-time recovery. The only recovery path is the daily logical dump (
supabase-backup.yml), which has never been test-restored. - Edge functions read ~40 secrets; nothing lists which ones a function needs, and several fail open (e.g.
whatsapp-webhookskips signature checks whenWHATSAPP_APP_SECRETis missing). - No uptime check or alerting; errors are noticed when staff complain.
Environments
PLT-001 · Staging before production
Status: PROPOSED · Owner: IT Every migration and edge-function change is applied to a staging Supabase environment and smoke-tested (PLT-032) before production. Staging is either a Supabase branch (Pro plan) or a separate free project with the same migrations, functions, secrets (test values) and auth settings. Staging never holds real customer data.
PLT-002 · Production on a plan with point-in-time recovery
Status: PROPOSED · Owner: Management Production runs on Supabase Pro (or higher) with PITR enabled, giving a recovery point measured in minutes instead of up to 24 hours. Until then, PLT-040 daily dumps are the only recovery and the risk is accepted in decisions-log.md.
PLT-003 · Local database matches production major version
Status: PROPOSED · Owner: IT
supabase/config.toml [db] major_version matches the production project's Postgres major version (Dashboard → Settings → Infrastructure), so local and CI database tests exercise the same engine. (CI currently runs on the configured version 15; the migrations and DB tests also pass on 17.)
Releases
PLT-010 · Migrations are applied by the pipeline, never by hand
Status: PROPOSED · Owner: IT
Production migrations are applied only by a CI workflow (supabase db push) that runs after merge to main, requires a manual approval (GitHub Environment production with required reviewers), takes a fresh backup first (PLT-041) and records the run. Nobody runs supabase db push, psql DDL or dashboard SQL against production by hand except during a declared incident, which is written up afterwards.
PLT-011 · Migrations are forward-only and idempotent
Status: DECIDED (existing, CLAUDE.md) · Owner: IT
A migration merged to main is never edited, renamed or deleted; fixes are new migrations. Every statement is safe to run twice (IF NOT EXISTS, CREATE OR REPLACE, DROP … IF EXISTS before CREATE POLICY). Filenames are YYYYMMDDHHMMSS_description.sql and each version is unique.
Enforced by: scripts/check-migrations.mjs (CI step + pretest) — errors on bad names/duplicate versions, warns on unguarded CREATEs.
PLT-012 · Edge function deploy checklist
Status: PROPOSED · Owner: IT
Before deploying an edge function: (1) deno check passes; (2) every secret it reads is set in the target project (see the table in deploy runbook); (3) verify_jwt in supabase/config.toml matches the function's own auth (functions with verify_jwt = false must call guardRequest or verify a signature); (4) any migration it depends on is already applied; (5) it is smoke-tested on staging; (6) the deploy is recorded (function, commit, who, when).
PLT-013 · Security-relevant secrets fail closed
Status: PROPOSED · Owner: IT
A function whose security depends on a secret (webhook signing secret, CAPTCHA secret, dispatch key) refuses the request with 503 when the secret is missing, instead of skipping the check. Currently fails open: whatsapp-webhook (WHATSAPP_APP_SECRET).
PLT-014 · Frontend and backend ship in a safe order
Status: PROPOSED · Owner: IT Database changes ship first and are backward-compatible with the frontend currently live; the frontend that depends on them ships after. Destructive changes (drop column, tighten RLS) ship only once no live frontend uses the old shape.
PLT-015 · No third-party secret in the browser bundle
Status: PROPOSED · Owner: IT
A key that lets someone spend our quota or act as us at an outside provider is an edge-function secret (supabase secrets set), never a VITE_* variable: everything prefixed VITE_ is inlined into the public JS bundle. The browser calls an edge function that holds the key, checks the caller is signed-in staff with the permission of the screen that uses it, and returns only the answer. Only public identifiers stay in VITE_* (the Supabase URL and anon key, the Turnstile site key, the Sentry DSN, the VAPID public key, the Google OAuth client id). A key found in a bundle is treated as leaked and rotated at the provider.
First case (2026-09-26): the AirLabs flight-schedule key was VITE_AIRLABS_API_KEY and shipped in the bundle. It is now the secret AIRLABS_API_KEY, read only by the flight-lookup function (inventory.view).
Browser code reads each Vite variable by name (import.meta.env.VITE_X). Handing import.meta.env around as an object makes Vite inline every VITE_* variable of the build into the bundle, read or not; until 2026-09-26 observability.ts, ErrorBoundary.tsx, gstnApiClient.ts and api.ts did this, so any variable set in Cloudflare shipped. A variable that is no longer read is also deleted from Cloudflare Pages and GitHub Actions.
Enforced by: supabase/functions/flight-lookup/, src/lib/flightLookup.ts, src/lib/flightLookup.test.ts (no key, no AirLabs call in the browser), src/test/viteEnvByName.test.ts (no whole-object use of import.meta.env).
Not built: a CI check that scans dist/ for provider keys. VITE_OXR_API_KEY and VITE_EXCHANGE_RATE_API_KEY (exchange rates) still ship in the bundle and have not been moved yet.
CI gates
PLT-020 · Required checks before merge
Status: PROPOSED · Owner: IT
A PR can merge to main only when these pass:
| Gate | Command | Runs |
|---|---|---|
| Migration guard | node scripts/check-migrations.mjs |
every run |
| Lint | npm run lint |
every run |
| Type check (ratchet) | npm run typecheck:ratchet |
every run |
| Unit tests | npm run test:run |
every run |
| Build + bundle budgets | npm run build + scripts/check-bundle-budgets.mjs |
every run |
| DB behaviour tests | scripts/db-test.sh after supabase start |
migrations / supabase/tests changed |
| Edge functions | scripts/deno-check-functions.sh |
supabase/functions changed |
| End-to-end | npm run e2e |
every non-Dependabot run |
Enforced by: .github/workflows/ci.yml; branch protection on main must list these jobs as required.
PLT-021 · The type-error baseline only shrinks
Status: PROPOSED · Owner: IT
scripts/typecheck-baseline.txt lists known type errors. New errors fail CI. Nobody adds lines to the baseline to get a PR green; a PR that fixes errors regenerates it (npm run typecheck:ratchet -- --update). Target: empty.
PLT-022 · Every database rule has a database test
Status: PROPOSED · Owner: IT
A rule enforced in the database (RLS, trigger, constraint, RPC) has a test in supabase/tests/*.sql that proves it, named with the rule ID, wrapped in BEGIN … ROLLBACK, ending with \echo ALL … PASSED.
Monitoring
PLT-030 · Errors are captured and alert someone
Status: PROPOSED · Owner: IT
Frontend errors go to Sentry (VITE_SENTRY_DSN, already wired) with release tags; edge-function errors are logged with a correlation ID. A new error type or an error-rate spike notifies the on-call person (email/WhatsApp) within 15 minutes.
PLT-031 · Uptime and database health checks
Status: PROPOSED · Owner: IT An external uptime check (e.g. Better Stack, UptimeRobot, Cloudflare health check) hits the app URL and a lightweight health endpoint every 5 minutes and alerts after 2 failures. Supabase usage/log alerts are configured for database CPU, disk, connection count and auth/5xx error spikes; Free-plan quota warnings (database size, egress, function invocations) go to the same inbox.
PLT-032 · Smoke test after every production change
Status: PROPOSED · Owner: Operations After any migration or function deploy, someone runs the smoke checklist in the deploy runbook (staff login, customer sign-up, partner sign-up, finance PIN, create booking, record payment) and records the result.
PLT-033 · A fallback that is taken is reported
Status: DECIDED · Owner: IT · Enforced by: src/lib/degraded.ts, src/test/noSilentCatch.test.ts
Where the API layer survives a failure — a display name it cannot look up, a default prefix, a best-effort mirror — it reports that it fell back, through degraded(): to Sentry (PLT-030) and to the browser console, naming the place and why the fallback is safe. It reports at most three times per place per page load. A failure that would make a number wrong (a price, a tax base, a balance) does not fall back: it fails the request, and the person sees the error. catch {} and .catch(() => null) in src/lib/api.ts fail the build.
Found in the V1 audit (2026-09-24): 78 such sites, including the two head-count lookups behind the GST base on the ground margin (PAX-034), which fell back to the plain head count — the exact over-charge PAX-034 exists to prevent. Those now fail loudly.
Backups and recovery
PLT-040 · Daily offsite backup
Status: DECIDED (existing) · Owner: IT
supabase-backup.yml dumps roles, schema and data daily at 02:00 UTC, with checksums, kept 30 days as artifacts (optionally S3).
PLT-041 · Backup before every risky change
Status: PROPOSED · Owner: IT
A fresh backup (workflow_dispatch of supabase-backup.yml) completes successfully — and its artifact is downloaded or confirmed — before any production migration, bulk data fix or restore.
PLT-042 · Monthly restore test
Status: PROPOSED · Owner: IT
Once a month the latest backup is restored into a scratch project or local Supabase, row counts of key tables (Booking, Payment, JournalEntry, JournalLine, Customer, User) are compared with production, and scripts/db-test.sh runs against it. Duration (actual RTO) and any problems are recorded. A backup that has not been restored is not counted as a backup.
PLT-043 · Rollback plan written before release
Status: PROPOSED · Owner: IT Every production release notes how to undo it: for code, the previous frontend deploy / function commit; for migrations, the forward "undo" migration or the restore point. Irreversible data migrations are called out explicitly and need management sign-off.
List completeness
A list screen is a promise: "this is what exists". The rules below make that
promise testable. They exist because an audit found fetchBookings() asking for
one 300-row page and six screens — the ops approval queue, the corrections
queue, finance's payment workflow, airline seat reconciliation, the group
booking picker and the sidebar badge — treating that page as "all bookings".
Nothing failed; the queues were simply short.
PLT-050 · A list never presents one page as the whole set
Status: PROPOSED · Owner: IT
Every list endpoint returns the shared paged shape (src/lib/paging.ts:
{ data, nextCursor, total?, totalIsEstimate? }) and every caller either pages
to exhaustion or states, in the UI, that it is showing a page of a larger set.
Page size is clamped to MAX_PAGE_SIZE (200); a caller asking for more gets
200, not a silent truncation at PostgREST's max_rows. Paging is keyset
(cursor on the sort column plus id), so inserts during a walk cannot make a
row appear twice or be skipped. A helper that walks every page (exportAll)
throws when it hits its ceiling — a short export is never returned as if it
were complete.
Enforced by: 20260920130000_bookings_paging_search.sql (booking_page);
tests: src/lib/exportAll.test.ts, src/lib/api.bookingsPaging.test.ts,
supabase/tests/bookings_paging.sql.
PLT-051 · Filtering, searching and counting happen in the database
Status: PROPOSED · Owner: IT
Filters and free-text search run server-side against the whole table, not over
the rows the browser happens to hold. Counters and KPI tiles come from a count
query with the same filters as the list — never from rows.length of a loaded
page. Exact counts are opt-in (withTotal=1); the default is a planner
estimate, because an exact count is a full scan.
The bookings tiles (DECIDED 2026-10-02, owner, through Hamid). Cancelled counts fully cancelled bookings only; partly cancelled and transferred bookings keep their own counts, because their travellers still travel or have moved. Each tile counts exactly what its Status filter word lists (bkg_status_set, 20261008140000).
Enforced by: booking_page / booking_stats in
20260920130000_bookings_paging_search.sql; test: supabase/tests/bookings_paging.sql.
A tile and the filter word with the same name count the same bookings: each status word
(pending, awaiting_approval, cancelled, …) is defined once, in bkg_status_set(), and
both the Status filter (bkg_status_db) and the tiles (booking_stats) read it
(20261008140000_booking_screens.sql; test: supabase/tests/booking_screens.sql).
Every sidebar badge comes from badge_counts() (20260927194500), one call, each count
taken the way its list decides what to show and under the caller's own row policies;
test: supabase/tests/every_badge_counted.sql.
PLT-052 · Bulk actions resolve their target set on the server
Status: PROPOSED · Owner: IT
A bulk action carries either an explicit id list or a filter plus exclusions —
never "the rows currently on screen". When it carries a filter, the server
resolves the set, and the confirmation (UX-001, level typed for cancel and
delete) shows the count the server returned, not the number of visible
rows. The action executes in batched server calls with the permission checked
server-side per batch; it is not fanned out one HTTP request per row from the
browser.
Enforced by: POST /sales/bookings/bulk-resolve and booking_filter_ids;
tests: src/lib/api.bookingsPaging.test.ts, supabase/tests/bookings_paging.sql.
PLT-053 · Id batches sent to the database are bounded
Status: PROPOSED · Owner: IT
Any .in('<column>', ids) built from a caller-supplied or query-derived list is
chunked at ID_BATCH_SIZE (100). An unbounded list grows the request URI with
the data set and fails wholesale once it is too long, which reads as a broken
page rather than a too-large query.
Enforced by: src/lib/chunk.ts; test: src/lib/chunk.test.ts.
The same audit found the problem is not only about bookings. PostgREST runs with
max_rows = 1000 (supabase/config.toml), so every list read without an
explicit window was silently truncated there — no error, no signal — and the
customers, visa, ticketing, requests, leads, quotations, group-roster and portal
screens all filtered, sorted and counted that capped page in the browser. The
audit log was worse: a hard .limit(500) with no offset or cursor, so
compliance evidence older than the newest 500 events was simply unreachable.
The three rules below add what PLT-050 … PLT-053 do not already cover.
PLT-056 · Keyset cursors, whitelisted sorts, clamped limits
Status: PROPOSED · Owner: IT
The sort column a cursor walks is NOT NULL and comes from the endpoint's
whitelist; any other sort value is a 400, never silently ignored, so a sort
key can never name a column the caller was not meant to order (or probe) by.
limit is clamped to MAX_PAGE_SIZE (src/lib/paging.ts, shared byte-for-byte
across workstreams) and a nonsense limit falls back to DEFAULT_PAGE_SIZE. A
malformed cursor is rejected, never treated as "start from the top". A read that
still uses the legacy unpaged shape logs a breadcrumb when it comes back at the
max_rows cap (warnIfTruncated()), and no .limit(N) above 1000 exists
anywhere.
Enforced by: src/lib/api.paging.test.ts, supabase/tests/list_paging.sql.
PLT-057 · A capped section says what it is hiding
Status: PROPOSED · Owner: IT
Where a section keeps a fixed cap for good reason — the Customer 360 requests
(50) and messages (20) tabs, a group's "needs attention" rollup, an incident
list — it returns the total and a way to page the rest, so nothing disappears
without the reader knowing. A rollup that aggregates an unbounded set (the
journey board aggregated every group and booking in a 60-day window) takes its
window on the outer keyset before the expensive aggregation runs.
Enforced by: 20260920120300_list_paging_rpcs.sql (get_customer_360_tab_meta,
get_customer_360_tab, get_journey_board_page, list_incidents_page);
test: supabase/tests/list_paging.sql.
PLT-058 · A cursor is a position, never a capability
Status: PROPOSED · Owner: IT
A cursor carries only a sort value and a row id the caller was already shown. It
grants nothing: every page re-applies the same tenant scope (agentId,
customerId) and the same RLS, so a cursor minted by one partner selects
nothing for another (ACC-010). Cursors are never signed, trusted, or used in
place of a permission check, and a filter value is quoted into the PostgREST
predicate so a comma or bracket in a customer's name cannot be read as filter
syntax.
Enforced by: src/lib/api.paging.test.ts (partner cursor replay),
supabase/tests/list_paging.sql.
Apps
PLT-060 · The Android app is a shell around the live site
Status: DECIDED 2026-09-25 (owner: "the app should have an APK version also") · Owner: IT
The Android app (in.alhudatravels.app) opens https://alhudatravels.in/m — the phone layout of the ERP (docs/features/mobile-app.md), the same routes and functions as the desktop, gated the same way — in a full-screen WebView and bundles no copy of the app. A phone-sized browser lands on /m as well, with a way back to the desktop layout. Notifications inside the app go through Firebase Cloud Messaging, since a WebView has no Web Push; browsers keep Web Push. A release of the web app is therefore a release of the Android app: nothing is rebuilt or reinstalled, and a phone never runs an old version. The APK is built by a hand-run workflow, never by CI; a release APK is signed only with the company keystore held as GitHub secrets, never with a throwaway key.
Enforced by: capacitor.config.ts (server.url), android/, .github/workflows/android-apk.yml, docs/operations/android-app.md; the /m routes in src/App.tsx and src/pages/mobile/; src/lib/nativePush.ts and the FCM path in push-send.
Not built: a Play Store listing, iOS, app-links.
PLT-061 · The native app takes JavaScript fixes over the air
Status: PROPOSED 2026-09-26 · Owner: IT
The native app (apps/mobile) receives JavaScript-only changes through EAS Update, on the production channel, without a reinstall. An update is published only from main, after the change is merged, from the machine that holds apps/mobile/.env; the publish script refuses when any EXPO_PUBLIC_* value is missing. An update is offered only to APKs built with the same version in app.json (runtime version policy appVersion); any native change — a native package, an SDK upgrade, a permission, anything in app.json or app.config.js — raises version and ships as a new APK. A phone checks on each cold start without waiting and runs a downloaded update on the next cold start, so a phone without signal opens as before.
Enforced by: apps/mobile/app.json (runtimeVersion, updates), apps/mobile/app.config.js (channel), apps/mobile/eas.json, apps/mobile/scripts/check-update-env.js, docs/operations/android-app.md → Updates without a new APK.
Not built: signed updates, iOS updates, an in-app restart prompt.