Decisions
The open questions this spec currently depends on. This page is a log, not a one-time list — update a row’s status in place as it gets resolved, and add new rows as new forks appear. Don’t let an answer live only in a chat thread; if it’s decided, it belongs here.
- 1. Identity provider
- 2. Organizations
- 3. Downstream app architecture
- 4. Billing system of record
- 5. Settings schema ownership
- 6. Session model for revocation
- 7. AWS account and region
- 8. Storage primitive scope
- 9. App developer/publisher model
- 10. Compliance scope
- 11. Tenant vs. Organization
- 12. Tenancy tiers & dedicated infrastructure
- 13. Organization-vs-User entitlement precedence
- 14. MCP server authorization scope
- 15. Platform subscription billing processor
| # | Decision | Status |
|---|---|---|
| 1 | Identity provider | 🟢 Resolved |
| 2 | Organizations | 🟢 Resolved |
| 3 | Downstream app architecture | 🟢 Resolved |
| 4 | Billing system of record | 🟢 Resolved |
| 5 | Settings schema ownership | 🟢 Resolved |
| 6 | Session model for revocation | 🟢 Resolved |
| 7 | AWS account and region | 🟢 Resolved |
| 8 | Storage primitive scope | 🟢 Resolved |
| 9 | App developer/publisher model | 🟢 Resolved |
| 10 | Compliance scope | 🟢 Resolved |
| 11 | Tenant vs. Organization | 🟢 Resolved |
| 12 | Tenancy tiers & dedicated infrastructure | 🟢 Resolved |
| 13 | Organization-vs-User entitlement precedence | 🟢 Resolved |
| 14 | MCP server authorization scope | 🟢 Resolved |
| 15 | Platform subscription billing processor | 🟢 Resolved |
1. Identity provider
Resolved: both, not either/or — native in-house auth, architected from the start for pluggable, per-Organization SSO.
- Native email/password is owned here, not delegated — a real
password_hash, real sessions, built and live from Phase 1, requiring zero external integration to use: no Auth0/WorkOS/Cognito account, no API key for a third-party provider, nothing to configure. A single developer standing up their first Application has a fully working login system the moment Phase 1 ships. Not “in-house vs. delegate”; in-house is the default every User has available, delegation is additive on top of it. - SSO is a per-Organization bridge, not a platform-wide or per-Application choice: an Organization admin connects their own company’s identity provider (Okta, Azure AD, Google Workspace, …) through a federation broker (leaning WorkOS, for exactly this “bring your own enterprise IdP” use case), and it becomes available to every member of that Organization. See Domain Model → SSOConnection.
- A User can hold both at once — a
passwordidentity and anssoidentity simultaneously — choosing either at login, not locked to one method per account. See Domain Model → UserIdentity. - MFA (TOTP) is optional and user-enabled for native accounts, via self-service enrollment (
POST /v1/users/{id}/mfa/totp) — not required, not yet built for SSO identities, whose MFA policy belongs to the member’s own IdP. - Sequencing: the
UserIdentity/SSOConnectionschema and the auth endpoints are built for this shape now, so adding a real SSO broker integration later is additive, not a breaking migration — but the broker integration itself (the actual Auth0/WorkOS/Cognito wiring) is explicitly not being built yet.POST /v1/auth/sso/{provider}/callbackexists as a route with its contract still open — see API Reference → Auth.
This replaces the earlier in-house-vs-delegate framing entirely — the real question was never which one, it was how both coexist on the same User without a later rework, which is what UserIdentity now answers.
2. Organizations
Both individual and team/business end users are expected, so Organizations isn’t a someday-maybe feature. It’s pulled forward to Phase 2 rather than Phase 3 — not the MVP itself (the first dozens of users are expected to be mostly individual early adopters), but needed well before the 10x/100x growth horizon in Deployment Architecture → Growth trajectory hits, where team accounts are assumed to matter.
organization_id being load-bearing in the schema from day one (per Non-Functional Requirements) was the right call regardless of timing — this just confirms it wasn’t a hedge against a hypothetical.
Note this organization_id (end-user team/seat grouping, scoped within one Application) is a different axis from the infrastructure-placement tenant_id introduced in Decision #11 — don’t conflate the two when reading Database Schema.
3. Downstream app architecture
Resolved: Applications are separately hosted, and not just web apps. An Application built on this platform can be an iOS, macOS, Windows, Linux, or web app, or any other platform entirely — the one thing every option has in common is that it’s a separate piece of software Substratal doesn’t host, which this API serves purely as a backend: API methods for authenticating a user and for reading/writing that user’s settings and related collateral, nothing about how or where the app itself runs. This confirms the “separately hosted services” assumption Trust Model already builds on — the “apps as modules/iframes inside one deployment” alternative is ruled out, since a native iOS or desktop app can’t be an iframe inside anything.
One real implication, not yet reflected in Trust Model: that page’s launch flow is written as a web redirect carrying a JWT. A native mobile/desktop app can’t receive a browser redirect the same way — it needs the equivalent of an OAuth 2.0 Authorization Code flow with PKCE (industry-standard for exactly this: a public client with no safe place to hold a secret, handing control back via a custom URL scheme or app link rather than a server-side redirect). The underlying claims and verification model Trust Model already specifies (short-lived JWT, effective_permissions, introspection, webhooks) don’t change — only the mechanics of handing the token to the app need a second, platform-appropriate flow alongside the existing web redirect. Flagged here as a known gap to close in Trust Model, not a reason to revisit anything else already decided.
4. Billing system of record
Resolved, fully: Substratal Apps is not a payment processor or marketplace billing engine, for any Application, ever — not even once the marketplace/eventual-UI phase ships. Each Application’s developer owns the billing relationship with their own end users (their own Stripe or equivalent) entirely outside this API, and simply calls this API’s Entitlements endpoints to reflect the outcome. Order/order_id is reference metadata the developer supplies for their own reconciliation, never a record this API’s own billing system pushes webhooks about. When the marketplace ships, Applications are made available through it based on each developer’s own pricing and billing mechanism — discovery, not a checkout Substratal runs.
The one and only payment flow that runs through Substratal Apps itself is its own Pricing subscription — what an Application owner pays Substratal for the platform (Starter/Team/Enterprise). See Decision #15 for how that one is processed.
5. Settings schema ownership
Resolved: schema-on-file with the hub. Each Application registers its own JSON Schema (settings_schema), and the hub validates every write against it — confirming what AppSettings already assumed, rather than falling back to a fully opaque blob. Centralized validation is why GET /v1/users/{id}/apps/{appId}/settings can promise a resolved, valid object instead of “whatever blob was last written,” and it’s the same schema the typed generated-column mechanism from Decision #8 reads to know what to type and index.
6. Session model for revocation
Resolved: immediately. “Turn off this user’s access” must take effect inside an already-open app session right away — not “by next login or token refresh.” This is a single global guarantee, not a per-Application opt-in.
What this requires of every Application, not just permits: a short JWT TTL alone can no longer be called sufficient, since a live session must not survive a revocation for the length of that TTL. In practice this makes the Trust Model’s three mechanisms a required layering, not a menu an app picks from — an Application must subscribe to the entitlement.revoked/entitlement.disabled/role.removed webhooks and force-expire the affected session the moment one arrives, with the JWT TTL as a backstop rather than the primary mechanism. Trust Model → How fast does revocation need to land? needs updating to reflect this as settled rather than an open tradeoff — done alongside this decision, see that page.
7. AWS account and region
Resolved: a dedicated AWS account, under AWS Organizations, in us-east-1. Elaborated in full — the reasoning for both the dedicated account and the specific region, not just the choice — in root DEPLOYMENT.md → AWS account & region, alongside the rest of the infrastructure build-out this decision feeds into.
8. Storage primitive scope
The product pitch names “storage” as one of the five things a developer shouldn’t have to build (alongside user, settings, organization, tenancy) — but the spec as written only has two narrow storage primitives: AppProfile’s custom field and AppSettings’ overrides, both flat JSON blobs with no server-side structure beyond the Application’s own settings_schema. Is that actually what “storage” means, or does a developer need something more general — arbitrary collections, file/blob storage — before this product does what it says on the label?
Resolved: Postgres-backed, and typed where it’s declared. The engine is Postgres (see Deployment Architecture → Database engine), so the storage primitive follows that directly rather than needing a separate decision:
- The source of truth stays
jsonb—AppProfile.customandAppSettings.overridesremain flexible JSON blobs, so a developer never has to pre-declare a migration to store a new field. - Every field a developer has declared in their
settings_schemagets a real Postgres type, not just JSON. The implementation backs each declared schema property with a Postgres generated column (GENERATED ALWAYS AS (overrides->>'week_start') STORED, cast to the schema’s declared type —text,boolean,integer,timestamptz, whatever it specifies) and an index on it. This is what “map to a respective PostgreSQL data type” means concretely: the JSON is where a value lives, the generated column is how it’s queried and type-checked once a developer has told the platform what shape to expect. - A general-purpose storage API (arbitrary collections, file/blob storage) is still not in scope today — the above is enough for the flag/preference/small-record use case every Application needs, and it’s the same mechanism regardless of how large that use case grows, since adding a new declared field just adds another generated column rather than requiring a new kind of storage object. Revisit only if a real Application — including the three-plus the platform’s own developer is building — hits something a typed JSON field genuinely can’t represent (e.g. actual file bytes).
9. App developer/publisher model
Third-party developers registering their own Applications is an explicit later phase, not day one (see Home) — but “later” still needs a real shape: self-service submission into review_status: pending_review, who reviews it and against what criteria, what SLA a developer should expect, and what happens to an already-launched app that gets suspended mid-flight (do its existing users’ Entitlements stay active, or does suspension cascade to them?).
Resolved — elaborated so the review-queue mechanics are no longer open:
POST /v1/applicationsstays platform-admin-only through MVP and Phase 2, confirmed. Theowner_user_id/owner_organization_id/review_statusfields on Application exist now specifically so this doesn’t require a breaking schema change when self-service registration ships in Phase 3.- Reviewer assignment: no dedicated assignment system — any platform User holding
applications.managecan review. At this team’s actual scale, a queue (GET /v1/applications?review_status=pending_review, admin-only) is enough; auto-assignment is exactly the kind of process worth skipping until submission volume makes a FIFO queue genuinely insufficient, which isn’t knowable in advance. - SLA: a 5-business-day review target, stated as policy, not a system-enforced guarantee — nothing automatically escalates or refunds over a miss. Revisit if real volume makes that insufficient.
- A fourth
review_statusvalue,rejected, closes a real gap the original three-value enum left open: there was no way to represent “reviewed, declined” distinct from “never submitted.” A rejected Application is not deleted — it’s retained with a requiredreview_notesexplaining why, visible to its owner, who may edit and resubmit (review_status: rejected → pending_review, same record, not a new Application). - Suspension does not cascade to existing Entitlements, by default. Setting
review_status: suspendedon an already-launched Application immediately removes it from the public catalog, blocks new Entitlement grants, and blocks new launch-JWT issuance — but a User who already holds an active Entitlement keeps it, and an already-open session isn’t force-killed by this action alone. The alternative (suspension silently cutting off every existing paying user) is the same “an org’s decision reaching into something an individual holds” problem Decision #13 already ruled against, applied to a different relationship. A platform admin can still separately, explicitly revoke the affected Entitlements via the ordinary Entitlements endpoint for an egregious case — that’s a distinct, deliberate action, not an automatic consequence of suspension.review_notesis required on asuspendedtransition too, same asrejected.
See API Reference → Applications for the endpoint-level detail this resolution implies.
10. Compliance scope
Resolved: the recommendation stands as decided. GDPR/CCPA/CPRA mechanisms are built now — Compliance & Data Protection specifies the actual hard-delete cascade, the GET /v1/users/{id}/export endpoint, and the 72-hour breach-notification commitment, not just a table of gaps. SOC 2 posture is built now (access control, audit logging, encryption, change management, vendor risk management — all already load-bearing parts of this spec anyway); the formal Type I → Type II audit is deferred until a customer’s security review actually requires it. HIPAA is explicitly out of scope — not deferred-as-a-maybe, decided: nothing here is built for PHI, and that’s revisited only if a real healthcare-vertical developer wants to build on the platform, as its own scoped project at that point.
This page’s Data residency row, previously open, is now resolved by Decision #12 — a customer needing EU residency gets a dedicated_region Tenant, not a platform-wide region change.
11. Tenant vs. Organization
Does “tenancy” — named in the product pitch alongside user/settings/organization/storage (see Home) — mean the same thing as Organization, or is it a separate concept?
Resolved: separate, and tied to different things. Organization is a domain/grouping object — a company, or a group within a company — used to organize which Users share admin standing and which Applications a group is granted access to as a whole. It carries no infrastructure meaning. Tenant is new: the infrastructure-placement and data-isolation boundary, tied to the subscription — concretely, to whoever owns an Application’s catalog entry (owner_user_id or owner_organization_id, see Applications), since that’s the only subscription relationship Substratal has directly (see Decision #4: a developer’s own end-user billing is their own concern, not this API’s).
Practically: a User or Organization can hold membership in many Organizations (even, now, across multiple Tenants — a person building one app and also using a seat on someone else’s app), but there is exactly one Tenant governing where a given Application’s data physically lives. See Domain Model → Tenancy for the full entity and Decision #13 for how Organization-level access decisions interact with individual Users — a related but orthogonal question.
12. Tenancy tiers & dedicated infrastructure
Some customers (Application owners) want — or need, for compliance — their own dedicated infrastructure rather than the shared Tier 0 database, and some need a specific geographic region for their data. Four sub-questions, all resolved together:
- How many tiers? Three:
shared(default — Tier 0, logical isolation only),isolated(a dedicated Aurora Serverless v2 cluster, same region),dedicated_region(a dedicated cluster in a customer-chosen AWS region — real data residency). See Domain Model → Tenancy and Deployment Architecture → Tenancy tiers. - Self-serve or gatekept? Gatekept by support today — no public API lets a customer trigger their own migration.
tenants.manage(support/superadmin only) is required even to request a tier change. Automated, customer-initiated tier changes are a later, explicitly revenue-gated step, not a roadmap-phase trigger: it means accepting a real amount of migration risk (a failed cutover, a narrower support safety net) that a human currently absorbs step by step, and that trade only makes sense once the volume of tier-change requests justifies building it. - Downtime during a tier change? A brief, scheduled maintenance window is acceptable — this is support-run and infrequent at current scale, so a snapshot/restore cutover is enough. Zero-downtime (logical replication) migration is a legitimate future upgrade once this is self-serve and frequent enough to need it, not a day-one requirement.
- Does this reach downstream Applications? No — tenancy governs where this API’s own data lives (Entitlements, AppProfile, AppSettings, the relevant Audit Events), never an Application’s own separately-hosted infrastructure. An Application’s
tenant_id/tier/regionare exposed as static metadata on the Tenant and Application resources, for the owning developer’s own benefit — not as a live JWT claim on every request, since placement is set once per Application, not computed per end-user per request the wayeffective_permissionsis.
13. Organization-vs-User entitlement precedence
When an end-user Organization holds an org-wide (org_seat) Entitlement to an Application, and a member of that Organization also holds (or could hold) their own individual standing for the same app, which wins?
Resolved, with one deliberate scoping refinement flagged below:
- Within one Organization’s own grant, the Organization’s decision is authoritative. A member cannot opt themselves in or out of their org’s seat grant — see Entitlements → Org-wide entitlements for the
member_scopemechanism (all_members/allowlist/denylist) that lets an org admin include or exclude specific members. - A User’s own personal Entitlement to the same Application (purchased or granted independently of any Organization) is a separate, untouched access path. An Organization’s exclusion of a member from its own org-wide grant does not reach into and revoke a personal Entitlement that member holds some other way. This is the one place this resolution departs from a literal “Organization always overrides User” rule — the alternative (an org silently revoking something a member individually holds) creates a real billing/legal defensibility problem (“the company turned off access to something I personally paid for”), and the scoped version below still satisfies the actual goal — an org’s decision about its own grant is final — without that side effect. Revisit this if it doesn’t match intent.
- A User’s effective access to an Application is the union of every active path: their own personal Entitlement (if any) OR any Organization they belong to whose grant includes them. Because each Organization’s grant is independently evaluated, a User in multiple Organizations (even across different Tenants) never hits a real “Org A says yes, Org B says no” conflict — Org B’s answer only ever governs Org B’s own grant.
- Attribution is always surfaced. Reading a User’s entitlement to an app that came from (or was blocked by) an Organization’s grant shows
source: org_seat, theorganization_id, and whethermember_scopeincluded or excluded this specific member — so anyone pulling a User’s access record can see which Organization is responsible, rather than seeing a bare allow/deny. See Entitlements for the exact shape.
14. MCP server authorization scope
MCP Server deliberately introduces no new authorization model — every tool call carries the same Bearer credential (user token or API Key) as the equivalent REST call, and is checked against the exact same permissions. The open question is narrower: should Substratal recommend — or eventually require — a purpose-scoped class of API Key specifically for agent/MCP callers, rather than relying on whatever key a human happens to hand their agent?
The case for a narrower default: an LLM deciding which tool to call based on a prompt (possibly influenced by untrusted data it has read, e.g. a disabled_reason or an AppProfile.custom field written by someone else) is a different risk shape than deterministic service code making the same call — not because the platform’s own enforcement is any weaker (it isn’t; Access Control doesn’t know or care whether its caller is an agent), but because the decision to call a destructive tool at all is now made by something a prompt can influence, where a service integration’s call sites are fixed at write time.
Resolved — elaborated for implementation, not just operational guidance. No new API Key type (that would be a parallel, redundant scoping system next to the one that already exists). Two concrete mechanisms instead, both shipping with API Keys:
intended_use: "service" | "agent"on API Key, set at creation. Pure metadata with one real effect: it changes the default of the field below, steering an agent-facing key toward the safer posture without forcing it.restrict_destructive: boolean, defaulttruewhenintended_use: "agent"(andfalsefor"service"). Whentrue, any request this key authorizes that classifies as destructive is rejected with403/code: "destructive_operation_restricted"— regardless of what permissions the key otherwise carries. This is the “harden into something enforced server-side” option from the original open question, resolved yes rather than left pending an incident.
“Destructive” is not a second classification scheme — it’s the exact same rule MCP Server → Tool annotations & safety already uses to derive destructiveHint for MCP tool calls (a DELETE, or a PATCH/POST that moves an Entitlement to disabled/revoked, removes an Organization member, or requests a Tenant tier change), applied as a hard gate instead of a client-side confirmation hint. One rule, two consumers: an MCP client reads destructiveHint to decide whether to prompt a human before calling a tool; the API itself enforces the identical rule server-side when the calling key has restrict_destructive: true. A key with this flag set can still read everything its permissions allow and can still call every non-destructive write (PATCH updates that don’t disable/revoke/remove anything) — it’s scoped out of the specific small set of actions where an LLM’s tool-call decision, not the platform’s permission check, is the weaker link in the chain.
This is why an agent integration is safer by default without being less capable by default: a read-only or toggle-only assistant never needs restrict_destructive: false at all, and the one that genuinely does (an assistant whose whole job is disabling compromised accounts, say) sets it explicitly, which is itself worth an Audit Event (api_key.restrict_destructive_disabled) precisely because it’s the deliberate exception, not the default.
15. Platform subscription billing processor
Decision #4 resolved who processes payment for a developer’s own end users (the developer, never this API, never even in the marketplace phase). This is the other direction: who processes payment for the platform subscription itself — the Starter/Team/Enterprise charge an Application owner pays Substratal.
Resolved: Stripe Billing. Chosen specifically because adopting it changes nothing already decided — Compliance scope already keeps PCI scope minimal by design (card data never touches this API directly either way), and a subscription-billing product used only for Substratal’s own direct customer relationship is a far smaller integration than the marketplace-payment-processor role Decision #4 explicitly ruled out. The implementation — object mapping, the tenants schema additions, webhook event handling — is specified in root DEPLOYMENT.md → Stripe Billing, not here, since it’s a build/ops concern rather than part of the API Applications and their developers call — see that doc’s own framing of what lives in the project root versus this site.