Skip to content

Release rehearsal — integration/ops-upgrade against production data (2026-09-19)

Verdict: GO, with the fix committed alongside this page (20260922130000_journal_group_dimension.sql, back-fill under break-glass). Without that fix the call is NO-GO: supabase db push would stop after nine of the sixteen pending migrations and leave production half-released.

What this was: the 2026-09-19 production backup restored into a throw-away local Postgres, the pending release applied to it the way a deploy would, and every database suite run against the result. No connection was made to the production project; the only thing taken from it is the backup artifact the nightly workflow publishes. The procedure is the one in the 18 Sep rehearsal, with two corrections recorded below.

1. What was restored

Source GitHub Actions run 35427530457, artifact supabase-backup-20260919T064709Z
Checksums all three verified (roles, schema, data)
Target local stack alhuda-rehearsal, ports 593xx, Postgres 17.6 with production's pinned auth / REST / storage versions
Result 124 public tables, zero restore errors

Production is at main's head: every object the newest main migrations create is present, and none of the sixteen pending migrations' objects are.

Dropped to fit the local auth version (as on 18 Sep): empty rows of seven auth tables, auth.one_time_tokens.expires_at (3 rows) and storage.buckets.versioning_status (2 rows). None affects this release.

Correction 1 — use Postgres 17

A scratch workdir without supabase/.temp/postgres-version starts Postgres 15 (the major_version in config.toml). Production is 17. Copy the *-version files into the scratch workdir — never project-ref, pooler-url or start-secrets, which would tie it to the production project.

Correction 2 — restored grants must match production

The local image's default privileges grant anon, authenticated and service_role everything on objects postgres creates in public. pg_dump records production's grants relative to the built-in default only, so a plain restore leaves every table and function with grants production does not have. The first pass of this rehearsal failed twelve suites that way — anonymous callers able to run get_customer_360, the paging functions and the numbering functions, a partner able to rename a permission — and every one of them was the restore, not production (the raw dump grants none of it). Revoking those default privileges before loading schema.sql restores the grants exactly. Backup & restore now does this: the same trap applies to a real disaster-recovery restore into a new project.

2. What ran

Sixteen pending migrations, filename order, one transaction per file plus its migration-history row — the shape of supabase db push. The migration history itself is not in the backup (known since 18 Sep), so it was rebuilt from main first.

Migrations applied 16
Failures 0 (after the fix)
psql time 0.8 s

Substantive notices: operational write grants: 8 new RolePermission rows; journal group dimension: 3 booking voucher(s) tagged, 2 hotel-consumption voucher(s) tagged. The three starter cancellation policies were added; the existing default, Standard Umrah Cancellation Policy, stays the default.

3. What broke, and what was fixed

NO-GO (fixed) — the departure back-fill hit the voucher guard · FIN-031

20260922130000_journal_group_dimension.sql:160: ERROR:
  Approved voucher PV-… cannot be edited (FIN-031): groupId. Reverse it and post a new one.

The migration back-fills JournalEntry.groupId on existing vouchers. Production holds approved vouchers; fin_journal_entry_guard refuses any change to an approved voucher. A freshly built database has none, so every suite passed. The departure tag is a reporting dimension, not an amount or account, so the back-fill now runs under the sanctioned break-glass (fin_break_glass()), transaction-local and switched off at the end; every row it touches is still audited. The guard itself is unchanged. supabase/tests/release_backfills.sql re-runs the migration over approved, untagged vouchers — it fails without the fix with the production error.

4. The database suites on the production copy

20 of 25 pass, including every security suite. The five that fail assert over whole tables or insert fixtures that production already has — the known limitation T8 in the accounting test, not release defects:

Suite Why it fails on real data
bookings_paging counts every booking; production has six more
role_dashboards counts every pending item
posting_failures fixture supplier "Marjan Group" already exists
tickets_visa_comms fixture setting whatsapp_phone_id already exists
inventory_integrity finds a genuinely over-allocated FIT row (below)

On a freshly built database all suites pass.

5. For the owner

  • One inventory row needs an operations decision: FIT 6E-26MAY26-A-FIT-OW records 3,003 cancelled seats against 3 bought. inventory_drift_report() lists it; the release leaves it untouched (AUD-010).
  • Cancellation policies: no active traveller's cancellation would be refused after the release. None of the nine live groups has a visa rate on its rate sheet, so a group moved to Umrah — standard refuses cancellations of travellers whose visa is applied for until its visa rate is set.
  • The restrictive-only RLS guard returns nothing.

6. Deployed — 19 Sep 2026

Backup before the change run 35439772185 (supabase-backup-20260919T111912Z), checksums verified
Migration check production at main (307 / 307, none remote-only); dry run listed exactly the 16 rehearsed files
supabase db push 16 migrations, exit 0, ~10 s; afterwards 323 / 323 applied, 0 pending
Edge functions none changed in this release
Secrets none added
Frontend PR #379 squash-merged as 1099653 — not published: the squash message carried the previous release's [skip ci], so Actions and Cloudflare skipped it. Published by this follow-up merge (runbook §6 now warns about it)
Operator / Checker Claude Code (operator) / owner (checker; production holds test data only)

Smoke tests are run by the owner once the new frontend is live.