Wathba Stack Map

wathba-platform · wathba-platform-frontend · wathba-infra

Wathba Stack Map

How the backend, the builder portal and the infrastructure fit together, written for an engineer who knows Laravel, Next.js and deployments and is new to this codebase.

Read 2026-10-08 platform main 839510be frontend main 6cc22fd facts from code; docs used where marked

01Wathba in one minute

Wathba (وثبة) is a platform for Saudi builders, many of them working through AI coding agents. A builder creates a project, turns on services such as payments, OTP, shipping or domains, and calls one stable Wathba API. Behind that API, adapters talk to local providers: Moyasar, Authentica, Torod, Souq T2 and NHC/Ejar.

Agents are first-class clients. Builders reach the platform four ways: the web portal, a Go CLI, a TypeScript SDK, and a hosted MCP server that coding agents connect to over OAuth.

  • wathba-platform is the product. A NestJS backend that serves the API, the MCP server, the hosted checkout pages and the webhook intake, and also runs the Temporal workers.
  • wathba-platform-frontend is the builder portal. A Next.js app with no database and no server-side API layer of its own; the browser calls the backend directly.
Marketing vs. code

ZATCA, KYC and StreamPay appear in product copy and docs, but have no adapters in code. Payments run through Moyasar only. StreamPay material in docs/streampay/ is a product reference.

02System map

Every client ends at the same backend. The backend is one container image deployed as two Cloud Run services (API and worker) plus one-shot jobs.

flowchart LR
  subgraph Clients
    B["Builder browser"]
    P["Payer browser"]
    A["AI agent"]
    O["Operator"]
  end
  subgraph GCP["GCP me-central2"]
    LB["HTTPS load balancer + Cloud Armor"]
    FE["Portal (Next.js)"]
    BO["Backoffice (Next.js SPA)"]
    API["wathba-platform: api role"]
    WK["wathba-platform: worker role"]
    PG[("Cloud SQL Postgres 16")]
    RD[("Memorystore Redis")]
    GCS[("GCS protected payloads")]
    NAT["Cloud NAT static IPs"]
  end
  subgraph Outside["Outside GCP"]
    TMP["Temporal (Cloud in prod, Dokploy in dev)"]
    PRV["Moyasar, Authentica, Torod, Souq T2, NHC/Ejar"]
    SES["AWS SES"]
    OBS["ClickStack / HyperDX (Dokploy)"]
  end
  B --> LB --> FE
  B -. "API calls with cookie" .-> LB
  P --> LB
  A -- "CLI / SDK / MCP" --> LB
  O --> LB --> BO
  LB --> API
  API --> PG
  API --> RD
  API --> GCS
  WK --> PG
  API --> NAT
  WK --> NAT
  NAT --> TMP
  NAT --> PRV
  NAT --> SES
  NAT --> OBS
  PRV -. "webhooks" .-> LB

Two details shape everything else. The frontends have no VPC access and no backend of their own, so all data flows through the API. The API and worker send all outbound traffic through Cloud NAT with fixed IPs, because providers allowlist those IPs.

03Stack, translated

ConcernHereClosest thing you know
RuntimeNode 24, pnpm 11, TypeScriptPHP-FPM + Composer
Backend frameworkNestJS 11 on Fastify 5Laravel: modules ≈ service providers, guards ≈ middleware, DI by abstract-class tokens
ORMMikroORM 6 with EntitySchema (no decorators), plus a lot of hand-written SQL repositoriesDoctrine-style unit of work, not Eloquent's active record
MigrationsHand-written immutable .sql files and a custom runnerLaravel migrations, but raw SQL and checksummed
Validationzod 4 for config and contracts; class-validator in the global pipeForm Requests
Background workTemporal workflows, activities and schedulesQueues + Horizon + scheduler, but durable and replayed
Domain eventsPostgres outbox and inbox, polled every second in-processEvents + listeners with a transactional outbox
CacheRedis, used only by auth (sessions, deny lists, rate limits, API key cache)Redis session and rate-limit driver
Authargon2id; hand-rolled Ed25519 JWTs via node:crypto; signed cookiesSanctum / Passport, written in house
MailAWS SES v2; a local SES-compatible inbox in DockerMailpit
Agent surfaceMCP server (@modelcontextprotocol/* v2) at /mcpNo equivalent
TelemetryOpenTelemetry → ClickStack (HyperDX UI on ClickHouse)Telescope + Sentry + a log stack
PortalNext.js 16 App Router, React 19, TanStack Query 5, react-hook-form + zod, Tailwind 4, shadcn/RadixWhat you already use
TestsJest 30 + Testcontainers (backend); Vitest 4 + MSW + Playwright (portal)PHPUnit / Pest with a real DB
InfraTerraform on GCP: Cloud Run, Cloud SQL, Memorystore, LB, Secret Manager, KMSForge / Vapor, but explicit

04How much rides on external infrastructure

Most infrastructure is managed GCP, provisioned in wathba-infra. A separate group of services runs on a self-hosted Dokploy server at ploy.jsa.sa, which no repo in this workspace provisions.

DependencyUsed forWho runs it
Temporal (prod)Workflows and schedulesTemporal Cloud Terraform only accepts a *.tmprl.cloud:7233 address; API key from Secret Manager
Temporal (dev)Same, shared by the deployed DEV workerDokploy temporal-grpc-dev.ploy.jsa.sa:443, TLS, no auth, namespace wathba-dev
ClickStack / HyperDXTraces, logs, metrics (ClickHouse underneath)Dokploy OTLP at otel[dev].wathba.ploy.jsa.sa; config lives in a wathba-observability repo that isn't in this workspace
SonarQubeQuality gate in the pre-commit hookDokploy sonar.ploy.jsa.sa
Docs app + MeilisearchCustomer docs, docs MCP, searchDokploy DEV only; prod not provisioned yet
Postgres 16All stateCloud SQL private IP; regional HA in prod
Redis 7Auth onlyMemorystore TLS + AUTH; HA in prod
GCS, Secret Manager, KMSProtected payloads, env secrets, CMEK and release signingGCP
AWS SES (ap-south-1)OTP codes, invites, notificationsAWS static IAM user keys injected from Secret Manager
MoyasarPayments, merchant onboarding, Apple Pay, webhooksProvider API the only provider live in prod
Authentica, Torod, Souq T2, NHC/EjarOTP/SMS, shipping, domains, rental contractsProvider APIs dev or off in prod
GitHub App + Slack webhookFeedback reports become GitHub issues and a Slack noticeSaaS
GitHub Actions on BlacksmithCI and deploys (keyless via Workload Identity Federation)SaaS
What this means in practice
  • Temporal is the dependency to understand deeply. Every provider call that can end ambiguously, every poll and about twenty recurring sweeps go through it.
  • ClickHouse is never touched by app code. It is only the storage behind the telemetry UI. Prod export is currently off (WATHBA_TELEMETRY_ENABLED=false).
  • The Dokploy server is a single point for DEV Temporal, telemetry, Sonar and docs. Who operates it, and its backups, aren't written down in any repo here.

05One image, three roles

There is one Dockerfile and one entry point, node dist/src/main.js. The environment variable WATHBA_RUNTIME_ROLE picks what the process does.

api

Serves HTTP: the REST API, /mcp, the hosted checkout pages, webhooks. Default when NODE_ENV=production.

worker

Also serves HTTP, but additionally starts one Temporal worker per task queue and upserts the Temporal schedules. Default everywhere else, including your laptop.

job

One-shot scripts in src/cli/: SQL migrations, smoke tests, API-contract census, provider-secret bootstrap. Run as Cloud Run jobs.

Every role polls the Postgres outbox once a second and runs event subscribers in-process, so subscribers must be idempotent. They are, through an inbox table keyed by event id.

Boot order in src/main.ts: OpenTelemetry registers first, then config is parsed with zod (bad config stops startup), then Fastify with a custom body parser that keeps raw bodies for webhook verification, then cookie, CORS and helmet plugins, then a global validation pipe.

06Module map

Each module in src/modules/<name>/ has the same folders: api/ use-cases/ domain/ ports/ adapters/ repositories/ migrations/ contracts/. Other modules may import only from contracts/. A shared src/platform/ layer holds infrastructure that no module may bypass.

The boundary rules are enforced by a test, not by ESLint: test/wathba-boundary-rules.e2e-spec.ts checks cross-module imports, that domain/ has no Nest or ORM imports, and that platform/ imports no module.

Identity and members

  • auth: logins, CLI device flow, MCP OAuth, API keys, operators, guards
  • member, registration, member-onboarding
  • verification, business-entities-and-profiles

Projects and catalog

  • projects: projects, environments, service bindings, pins
  • service-catalog: signed versioned releases and manifests
  • agent-access: MCP server, CLI workspace
  • member-dashboard: read models

Calling providers

  • operation-execution: the single doorway and all runtime adapters
  • provider-connections: credentials, onboarding
  • provider-routing, provider-health
  • eligibility, traffic-control: gates

Money

  • billing-wallet: quotes, wallet, ledger
  • budget-controls, usage-metering
  • payment-links: intents, checkout, refunds
  • domains: Souq T2 domain purchase

In and out

  • inbound-provider-updates: webhooks from providers
  • member-notifications: notifications and signed webhooks to builders' apps

Operations

  • administration: backoffice cases, approvals, feedback
  • audit-and-evidence: immutable audit facts
  • lookups: recorded Moyasar exchanges

integration-sessions holds only historical SQL; the module itself was removed.

07Data and migrations

  • One Postgres schema per module (projects, auth, payment_links, billing_wallet, …). Modules reference each other by opaque prefixed IDs (prj_, env_, mem_), never by foreign keys.
  • Tenancy: a member owns projects; a project owns environments and service bindings. Access goes through ProjectAccess.memberOwnsProject(...). Project API keys are bound to one project and environment.
  • Aggregates carry an optimistic version. Idempotency keys are backed by unique constraints, and every mutating HTTP call takes an Idempotency-Key header.

Migrations

MikroORM's migrator is configured but unused. The real runner is src/platform/migrations/sql-migration-runner.ts:

  • Files are YYYYMMDDnnnn_name.sql in each module's migrations/ folder, about 290 of them.
  • They run in the order of the sqlMigrationDirs array, then by filename. A new module's folder must be added to that array.
  • Each file is recorded with a SHA-256 checksum. Editing an applied file fails CI.
  • In deploys, migrations run before new code gets traffic, and the old image is smoke-tested against the new schema. Every schema change has to work with the previous release.

08Temporal

If you've used Laravel queues: a Temporal workflow is a job whose progress is persisted at every step, so it survives restarts and can wait for hours. The cost is that workflow code is replayed from history, so it must be deterministic.

How the code uses it

  • A port in src/platform/workflows/ wraps the SDK. Each module registers a workflow contribution: name, task queue, activities, schedules.
  • About 20 task queues, one per area: wathba-operation-execution, wathba-payment-links, wathba-provider-health, wathba-billing-wallet and so on.
  • Main uses:
    • reconciling a provider call whose outcome is unknown;
    • polling Moyasar for a pending payment;
    • multi-step provider setup;
    • health probes every 15–30 s;
    • expiry and stale-state sweeps.
  • Workflows are started after the DB commit, from outbox subscribers. Workflow IDs are business keys, so a duplicate start is harmless.

Rules the team follows

  • *.workflow.ts files import only @temporalio/workflow and pure types. No DB, HTTP, Date.now() or randomness.
  • Payloads carry IDs and handles only; a guard rejects secret-looking keys.
  • Versioning: a breaking change ships as a new workflow name (for example serviceSetupWorkflowV2 on a new queue), or with patched(). CI fails any PR that touches workflow code without a note in docs/plans/temporal-compatibility/.
WhereAddressNotes
Locallocalhost:7233temporalio/temporal dev server in the Compose local profile
DEVtemporal-grpc-dev.ploy.jsa.sa:443Self-hosted, no auth, UI at temporal-dev.ploy.jsa.sa
PROD*.tmprl.cloud:7233Temporal Cloud with an API key
Laptop gotcha

With nothing configured, the code's defaults point at shared DEV, and the local role defaults to worker. Because queue names are fixed, a laptop started without settings polls the same queues as the deployed DEV worker and can take its jobs. Use the local Temporal from .env.local.example, or your own namespace on the shared cluster.

09Providers and the service catalog

Catalog, manifests, pins

A service (payments.moyasar, shipping.torod, messaging.otp.authentica, messaging.otp.wathba, …) is defined in a manifest. A manifest declares the provider, the adapter class, the operations, the setup flow, scopes and pricing. Manifests ship in signed, immutable catalog releases, known by number (Catalog 016, 019, 023, …). The number names a release, not a code version.

When a project enables a service, its binding records a pin to the exact catalog release and API-contract version. The stated rule is that a new release never changes behaviour for a builder who already integrated; they move only through an explicit upgrade.

The execution doorway

Every provider call goes through ExecuteOperationUseCase in operation-execution, in a fixed order:

  1. Claim the idempotency key.
  2. Check the pin and catalog availability.
  3. Approvals, eligibility, traffic admission.
  4. Connection readiness, and resolve the credential handle.
  5. Reserve wallet funds where the service is billed by Wathba (Authentica, Ejar).
  6. Choose a provider route.
  7. Write a "dispatch armed" marker, then call the adapter. A timeout becomes UNKNOWN and a Temporal reconciliation workflow takes over.

Adapters that exist

Moyasar (hosted checkout, payment read, full refund), Torod (shipping), Authentica (OTP/SMS, with the legacy Authenta adapter beside it), Wathba OTP (first-party, emails the code), NHC/Ejar (rental contracts), Souq T2 (domains). Provider fallback is limited: if the chosen route fails, the call fails closed with no_executable_provider.

Credentials

Provider credentials are AES-256-GCM sealed into a Postgres table by LocalSecretStoreAdapter. Despite the name, it is the only implementation and is used in deployed environments too, with the key held in Secret Manager. Code receives opaque handles and sees the plain value only inside a callback.

10API surfaces and auth

PrefixWho calls it
/v1/platform/*Portal, and builders' own servers through the SDK
/v1/cli/*CLI
/v1/backoffice/*Operators
/v1/public/*, /pay/:slug, /checkout/*Payers; the backend renders the checkout pages itself
/v1/provider-callbacks/*Provider webhooks
/mcp, /.well-known/oauth-*Coding agents

Three global guards run on every route, and every endpoint must declare @Public() or @RequireScopes(...). The guard turns the request into a principal:

CredentialPrincipalNotes
wtb_sid cookiemember, webSigned cookie pointing at a Redis session (12 h). SameSite=Lax.
Bearer JWT (Ed25519)member, cli or mcp_hostCLI via OAuth device flow; MCP via OAuth code flow with PKCE
x-api-key wtb_proj_live_…api_keyBound to project and environment; optional HMAC request signing
wtb_ops_sid cookieoperatorBackoffice, email OTP, sessions in Postgres
Google ID tokenworkloadCI-only bootstrap and catalog publishing endpoints

URLs stay at /v1. Compatibility is chosen by the pins on a binding, or a wathba-version header, not by URL. Errors are application/problem+json with a correlationId.

The MCP server exposes tools such as list_projects, create_project, get_service_integration_docs and recommend_services_for_repository. A CI workflow fails a backend change that renames a tool the public plugin depends on.

The contracts/ folder

Generated and committed: the public OpenAPI, version manifests, route inventory, AI-integration schemas, signed catalog artifacts. Regenerate with pnpm contracts:generate. The SDK and docs import the OpenAPI from here. The portal does not generate its API client from it; its types are hand-written.

11Trace: a payment link, end to end

This flow shows most of the architecture in one pass.

sequenceDiagram
  autonumber
  participant B as Builder (portal)
  participant API as wathba-platform API
  participant DB as Postgres
  participant Pay as Payer browser
  participant OE as operation-execution
  participant M as Moyasar
  participant T as Temporal worker
  B->>API: POST /v1/platform/projects/:id/payment-links
  API->>DB: save link + outbox event
  Pay->>API: GET /pay/:slug (server-rendered page)
  Pay->>API: POST /v1/public/payment-links/:slug/checkout
  API->>DB: payment intent + checkout session
  Pay->>M: card tokenised directly with Moyasar
  Pay->>API: POST /v1/public/checkout-sessions/:id/submit
  API->>OE: execute payments.moyasar startHostedCheckout
  OE->>M: create payment
  OE-->>T: if pending: start paymentLinkPollWorkflow
  M-->>API: webhook /v1/provider-callbacks/moyasar/:reg
  T->>M: poll payment status
  API->>DB: ProviderObservationApplied → link paid
  B->>API: GET .../payments (sees it paid)
  1. Create. Portal page payments/links/new → postCreatePaymentLink() → CreatePaymentLinkUseCase. Saved with an outbox event in the payment_links schema.
  2. Open. The payer hits /pay/:slug, a page the backend renders. Starting checkout creates a payment intent and a checkout session, and hands back a one-time token.
  3. Pay. Card data goes from the browser straight to Moyasar. The submit call runs through the execution doorway with idempotency key payment-intent:<sessionId>.
  4. Wait. A pending or 3-D Secure payment starts a polling workflow on wathba-payment-links. A timeout starts a reconciliation workflow instead.
  5. Learn the outcome. Two independent paths: Moyasar's webhook (verified by a shared secret in the body) and the poll. Both feed ApplyProviderObservation; competing observations are merged.
  6. Update. A subscriber moves the payment, intent and link to paid or failed, through an inbox so redelivery can't apply twice. Notifications and signed webhooks to the builder's app can follow.
Availability

PaymentLinksModule loads only when a Moyasar checkout flag is set. In prod it is enabled for one configured project and environment. The DEV environment file sets none of those flags, so payment links likely don't run on DEV. A separate sandbox-candidate deployment exists for them.

12The portal (wathba-platform-frontend)

Shape

App Router with output: "standalone". Every page.tsx is a thin Server Component that renders a client component from src/features/<area>/. API calls live in src/lib/wathba-api/, one file per backend area.

No backend-for-frontend

No middleware and no proxy routes. The only route handler is /healthz. The browser calls the API origin directly with credentials: "include", and apiFetch refuses paths outside an allowlist.

Auth

Client-side only: RequireAuth calls /auth/status and redirects to /login. Sign-in is password-first, with email OTP as a second option. Onboarding is invite-only.

Agents

/device approves CLI logins (device flow). /oauth/authorize is the consent screen for MCP clients. /app/agents/connect shows install commands.

Things that work differently from a typical Next app

  • Same-site, different origin. The session cookie is SameSite=Lax, so the portal and API must be sibling subdomains (platformdev. and apidev.wathba.info). A pair of default *.run.app URLs won't work.
  • All NEXT_PUBLIC_* values are baked in at build. Changing Cloud Run env vars does nothing; each environment has its own image. A validator rejects bad combinations during the Docker build.
  • Hand-written i18n. src/lib/i18n/en.ts and ar.ts (about 5,000 lines each) are typed catalogs. Locale is client-side in localStorage; an inline script sets dir before hydration; layouts use logical Tailwind classes (ms-*, me-*).
  • Fixture mode. NEXT_PUBLIC_WATHBA_FIXTURE_MODE=msw runs the whole portal against MSW handlers, with named scenarios (payments-active, unauthenticated, …). All unit and Playwright tests run this way; none hit a real backend. It is compiled out of staging and prod builds.
  • Feature flags default off and are forced off in staging and prod. Locally they can be turned on with localStorage["wathba.dev.flag.<name>"]="true".
  • Agentation is a feedback toolbar on local and DEV. Annotations go to the backend, which opens a GitHub issue and posts to Slack.

Root files ralph/, prompt.md and prmopt.md are leftovers from the autonomous-agent loop that first built the portal. Treat them as history.

13Observability

  • The OpenTelemetry SDK registers before anything else. It exports traces, metrics and logs over OTLP, and instruments HTTP, undici, pg, ioredis and pino. Resource attributes include the release SHA and runtime role.
  • The destination is a ClickStack gateway on Dokploy. In DEV you search in HyperDX at clickstackdev.wathba.ploy.jsa.sa. Cloud Logging still receives stdout.
  • Bad telemetry config disables export but never stops the app.
  • Browser telemetry from the portal is proxied through the API at /v1/telemetry/browser/* with a short-lived token. It is off by default, sampled, limited to an allowlist of routes, and masks inputs.
  • On Cloud Run with idle CPU, export batches get dropped. DEV runs with CPU always allocated; prod would need the same before traces can be trusted.

GCP monitoring adds uptime checks on every /healthz, Cloud SQL connection and Redis memory alerts, and two auth log-metric alerts. Prod alert policies have no notification destination set in Terraform.

14GCP topology and deploys

Two projects

  • wathba-platform: the DEV runtime plus shared pieces: Terraform state, Artifact Registry, the wathba.info DNS zone, the GitHub identity pool.
  • wathba-platform-prod: production. Hosts are api., platform., backoffice.wathba.info; DEV prefixes each with dev.

Traffic reaches Cloud Run through one global HTTPS load balancer per environment, with Cloud Armor rate rules on auth and OAuth paths. API and frontends accept traffic only from the load balancer; the worker is internal only. The backoffice host also routes /v1/* to the API, so the backoffice calls it same-origin. In prod the backoffice is behind Google IAP as well.

DEVPROD
Cloud SQL1 vCPU, zonal2 vCPU, regional HA, deletion protection
RedisBasicStandard HA
API instances0–51–20, CPU always on
Worker instances11–2
FrontendsRun with next devProduction builds
NAT IPs12

Backend release pipeline

flowchart LR
  M["push to main"] --> V["full verify gate"] --> I["docker build, push to Artifact Registry"] --> S["Trivy scan"] --> D["deploy DEV"]
  D -. "manual dispatch, confirm_production" .-> P["deploy PROD"]
  subgraph R["Each deploy (deploy-cloud-run.sh)"]
    direction TB
    r1["run migrate job on new image"] --> r2["smoke OLD image on NEW schema"] --> r3["new worker + API revisions, no traffic"] --> r4["smoke candidate"] --> r5["shift traffic: DEV 100%, PROD 1→10→50→100"] --> r6["public smoke; auto-rollback on failure"]
  end
  • Every pipeline authenticates to GCP with GitHub OIDC (Workload Identity Federation). There are no service-account keys anywhere.
  • Terraform ignores the container image, so the running version isn't in Terraform state; release workflows own it.
  • Secrets come from Secret Manager as env vars pinned to latest, so a rotated secret needs a new revision.
  • The portal and backoffice follow the same shape. Push to main deploys DEV; prod is a manual dispatch.
  • Terraform changes go through gates.yml: static checks and plans on every PR, DEV apply on main, and a staged, manually confirmed prod apply limited to an allowlist of people.

15Local dev and quality gates

  • Compose: Postgres 16 and Redis 7 by default. The local profile adds Temporal and an SES-compatible inbox. Walkthrough: docs/runbooks/local-development.md.
  • Backend scripts:
    • pnpm start:dev, pnpm db:migrate:local, pnpm dev:seed, pnpm dev:smoke;
    • pnpm test (about 980 unit files) and pnpm test:e2e (about 150);
    • pnpm contracts:generate.
  • Tests need Docker. Testcontainers starts one Postgres and Redis pair; each suite gets its own database cloned from a migrated template.
  • verify:precommit runs lint, build, contract checks, secret scan, migration replay, coverage and e2e. CI runs the same gate. The optional pre-commit hook adds a SonarQube quality gate. It is heavy and needs Docker; budget close to an hour on a small laptop.
  • Portal: pnpm dev serves on 3000, though .env.example assumes 3200. Its pre-commit hook runs format, lint, typecheck, unit tests, build and Playwright on three browsers.
  • CI runners are Blacksmith. Repo rules require Blacksmith for any new job.

16Sibling repos

wathba-backoffice

Internal admin, English only. Next.js 16 used as a client-only SPA, bun, ts-rest client from a local @wathba/contracts package. Operator session cookie wtb_ops_sid. Behind IAP in prod.

wathba-cli

Go and Cobra; every command prints one JSON object, for agents. Login is the OAuth device flow approved at /device on the portal, with tokens kept only in the OS keychain. Released with goreleaser and cosign to install.wathba.info and npm. Owns the agent skill source, mirrored to the public wathba-skill repo.

wathba-sdk-typescript

Server-only client (@wathba-cli/sdk), zero dependencies, authenticated with project API keys. Its OpenAPI and release.json are imported from the platform's contracts/. Published with npm provenance.

wathba-infra

All the Terraform above: two env roots, about 17 modules, custom Checkov policies (for example, no service-account keys, region locked to me-central2), daily drift detection.

wathba-docs

Fumadocs on Next.js, English and Arabic. Serves agents through /llms.txt, markdown routes and a read-only docs MCP at /api/mcp, plus an "Ask the docs" chat backed by DeepSeek. Deployed to Dokploy, DEV only so far.

wathba-plugin

A generator repo that packages the hosted MCP server and a short skill for Claude Code, Codex and other clients. Everything under plugins/, .claude-plugin/ and similar is generated by scripts/sync.mjs; never edit it by hand.

17What to read, and what to skip

Read

  • wathba-platform/AGENTS.md (not CLAUDE.md, an older copy)
  • docs/runbooks/local-development.md and the other runbooks
  • docs/modules/README.md and the per-module pages
  • docs/plans/wathba-mvp-impl-plans/00-conventions-and-temporal.md
  • docs/plans/ci-cd/, the Temporal compatibility policy
  • docs/design/public-api-backward-compatibility-v1.md, provider-gateway-mvp-decision.md, wathba-checkout-universal.md
  • docs/operations/clickstack-telemetry.md
  • Portal: docs/backend-integration.md (a dated field log), environment.md, deployment.md
  • wathba-infra/README.md

Skip unless you need it

  • Most of docs/plans/*.md: per-feature fix plans
  • docs/superpowers/, research/, roadmaps/, evidence/, launch/, learnings/
  • docs/operations/nestjs-observe*: a retired pilot
  • docs/streampay/
  • About 100 scripts/generate-catalog-* and sign-* files: catalog publishing tools
  • docs/knowledge/, docs/skills/: content shipped to agents, not engineering docs

18Week-one traps

  1. A plain pnpm start:dev runs as worker and, without config, connects to shared DEV Temporal (section 08).
  2. A new module's migration folder must be added to sqlMigrationDirs. The migration linter keeps its own list, which currently lacks domains, lookups and two others.
  3. Schema changes must be backward compatible with the previous image; the deploy tests that.
  4. Changing a route, schema or manifest usually means pnpm contracts:generate, or CI fails.
  5. jose is in package.json but unused; JWTs are hand-rolled on node:crypto.
  6. The Moyasar webhook secret travels in the request body. Never log raw webhook bodies.
  7. The onboarding progress SSE feed is in-process. With several API instances, a client may miss pushes; the polling endpoint is the reliable path.
  8. Authenta (legacy) and Authentica (current) are different services; payments.wathba is a retired alias of payments.moyasar.
  9. Portal auth is client-side; a 401 doesn't always mean the session is gone. Check problem.code first.
  10. Several docs reference a docs/Wathba Build Guide/ folder and a wathba-nestjs-arch-skill that aren't in the repo.

19Open questions found while reading

Places where the code and the docs disagree, or where the answer isn't in any repo here. Worth asking the team.

  • Prod Temporal. Platform AGENTS.md says prod Temporal is on Dokploy with no public gRPC. The prod Terraform only accepts a Temporal Cloud address. The infra repo looks more recent.
  • The Dokploy server. Who operates ploy.jsa.sa, what is backed up, and where is the wathba-observability repo?
  • Portal pay page vs. backend. The portal's /pay/[slug] calls /pay/:slug/status, /checkout and /attempts/:id/status. None of those exist on backend main, which renders /pay/:slug itself and serves checkout under /v1/public/* and /checkout/*. The portal page looks superseded; confirm before relying on it.
  • Contract drift. The portal's vendored contracts/ai-integration/v1 (47 files) doesn't match the backend's v1 (32 files), and the digests differ. It is unclear which side is current.
  • Prod alerting. Alert policies exist, but no notification destination is set in Terraform.
  • Payment links scope. Is the single-project gate in prod temporary, and when does it open up?
  • Who can release to prod. The PROD_APPLY_ACTORS list and GitHub environment settings aren't visible from the repos.