JahUs: deploys that refuse to go wrong
A two-player web game used as the proving ground for environment isolation, guarded production releases and a client that cannot write to its own database.
- Role
- Platform, security and release engineering
- Period
- Sep – Oct 2026
- Status
- Private · staging; production provisioned, not yet released
- Source
- private · architecture and practices described
- Problem
- Several Firebase projects live on one account: an emulator project, staging, production, an earlier beta and unrelated ones. A mistyped alias or a default project is all it takes to ship to the wrong place.
- Key decision
- Make the safe path the only path: allowlisted targets, tagged releases, a typed confirmation, and checks after every deploy.
- Result
- Separate emulator, staging and production projects; a client with no write access; App Check enforced on staging after a monitored rollout; 468 automated test cases including 34 security-rule tests.
Architecture
Components and flows as text
| Component | Kind | Technology | Flows out |
|---|---|---|---|
| Guarded deploy | pipeline | operator terminal · release tags only | ⇢ Firebase Hosting: deploy at release tag, verified after |
| Web app | client | React SPA · App Check | → Firebase Hosting: static assets; → Firebase Auth: sign-in, ID token; → Callable API: callable: ID token + App Check token; → Firestore: read-only listeners, rules-scoped |
| GitHub Actions | pipeline | emulator suite · no cloud credentials | — |
| Firebase Hosting | edge | per-environment targets | — |
| Firebase Auth | external service | email + Google sign-in | — |
| Callable API | service | Cloud Functions v2 · 16 callables | → Firestore: all writes, transactional; → Cloud Monitoring: privacy-safe structured logs |
| Daily jobs | scheduled job | purge · scheduling · 0 retries | → Firestore: daily purge |
| Cloud Monitoring | external service | log metrics · uptime · alerts | — |
| Firestore | datastore | deny writes · TTL · PITR | — |
Trust boundaries: Firebase project, one per environment (Firebase Hosting, Firebase Auth, Callable API, Daily jobs, Cloud Monitoring, Firestore).
Context and constraints
JahUs is a relationship game for two people. The product is small; the platform around it was built as if it were not. The interesting problems were operational: keeping environments apart, making production changes deliberate, and making sure personal answers can never leak between players.
- Three isolated projects. An emulator-only demo project, staging and production. The default alias points at the emulator project and never at production.
- Private by construction. A private answer is readable only by its author and cannot be listed. Operational logs have no free-form payload field, resource ids are one-way hashed, and errors are logged by category only.
- A data lifecycle. Users can export their data and delete their account; closed pairs and expired content are purged daily; Firestore has delete protection and 7-day point-in-time recovery.
- Cost-bounded. Functions are pinned to one region with 256 MiB, a 60-second timeout and a maximum of 10 instances, scaling to zero.
Verification
- 468 automated test cases: unit, component and integration tests, 34 security-rule tests on the emulator, and 31 Playwright specs including real two-browser journeys and 146 visual baselines.
- CI on every push and pull request in three jobs, against the Firebase Emulator Suite only; CI holds no cloud credentials.
- The deploy guards are tested themselves, including tests that staging and production configuration cannot cross.
- Monitoring as configuration: log-based metrics for function errors, failed scheduled runs, callable rejections and App Check rejections, an uptime check, and alert policies.
Decisions
Clients never write to Firestore
Every mutation goes through one of 16 callable functions, all built on one wrapper: authentication required, the user id taken only from the verified token, strict input schemas that reject unknown fields, and App Check enforced by a parameter that defaults to on.
Options considered
- Validate writes in security rules: no function cold starts, but validation split across rules and code.
- Server-only writes: the rules become read-only and trivially reviewable.
Trade-off accepted
Higher latency and function cost on every write. In exchange, multi-document writes are transactional and the attack surface is one wrapper.
Production only from a release tag, never from CI
The production deploy refuses to run in CI, on a dirty tree, or on a commit without an annotated vX.Y.Z tag already on origin/main, and the operator must type the project id. It re-runs typecheck, lint, unit and emulator tests first.
Options considered
- Continuous deployment from CI: faster, and puts production credentials in CI.
- Operator-run, tag-gated deploys: every production change is deliberate.
Trade-off accepted
No continuous deployment, and releases depend on one operator workstation.
Allowlist the targets, blocklist the neighbours
A deploy target must match exactly one approved project id; the earlier beta, staging, demo and unrelated projects are blocklisted explicitly. The env file and the built bundle are checked to reference only the target project.
Options considered
- Trust Firebase aliases and the gcloud default project.
- Allowlist-first, with tests that keep staging and production configs isolated.
Trade-off accepted
Adding an environment is a code change with tests.
Verify the deploy, do not assume it
After deploying, the script checks the Firestore location and delete protection, compares the deployed function inventory and region with an expected list, and checks that the latest Hosting release carries the deployed commit SHA.
Options considered
- Treat a zero exit code from the deploy CLI as success.
- Assert the intended end state.
Trade-off accepted
More script to maintain; any drift fails loudly.
App Check: monitor first, then enforce
App Check was rolled out in stages on staging, from monitor mode to enforced on Functions, Firestore and Auth, with a probe test that proves enforcement. The enforcement parameter defaults to on, so a new environment is protected unless it opts out.
Options considered
- Enforce immediately: risks locking out real users before verified traffic is seen.
- Staged rollout with a fail-closed default.
Trade-off accepted
Production remains in monitor mode until its own monitoring window completes.
Failure modes
| Failure | Detection | Handling | Evidence |
|---|---|---|---|
| Deploy aimed at the wrong project | Target not on the allowlist, or blocklisted | Refused before anything runs | targets.mjs |
| Production deploy started from CI | CI environment detected | Refused | production-guard.mjs |
| Untagged or dirty commit | No annotated tag on origin/main | Refused | production-guard.test.ts |
| Deploy drifts from intent | Post-deploy assertions | Fails with the specific mismatch | production.mjs |
| Scheduled job partly fails | Partial failure reported as failed execution | Alert policy fires; next run resumes | functions/src/index.ts |
| Abuse of invites and actions | Per-user fixed-window counters | Rate-limited; invite codes stored only as hashes | services/rateLimit.ts |
What I would change next
- Re-baseline the visual regression suite so CI on main is green again.
- Move project, billing, budget and Auth configuration into Terraform; today only rules, indexes, functions, hosting and monitoring are code.
- Add a Content-Security-Policy and frame-ancestors to the Hosting headers.
- Enforce App Check on production before the first release.
- Add a missed-run watchdog for the daily jobs.
Evidence index
The source is private; references are to files and pull requests in it.