A SaaS product serves many customers from one system. Each customer, a tenant, expects to see only its own data, to get steady performance whatever the others are doing, and to have its data restored or deleted without affecting anyone else. Multi-tenant architecture is the set of decisions that makes those promises true.
Most of those decisions are cheap at the start and expensive later. This post covers the main isolation models, the database features that back them up, the noisy neighbour problem, the parts of the system people forget, and how to change course once the product is live.
What multi-tenancy means in practice
A tenant is the unit that owns data and pays: usually a company, sometimes a team or a workspace. One tenant has many users, and one user may belong to several tenants, as with a consultant who works in several client workspaces.
Isolation has three sides:
- Data isolation. No request, report, export or search ever returns another tenant's records.
- Performance isolation. One tenant's heavy use does not slow everyone else down.
- Operational isolation. You can back up, restore, migrate, move or delete one tenant on its own.
The isolation models below trade these three against cost and operational effort.
Three ways to separate tenant data
Shared schema with a tenant id
Every tenant's rows live in the same tables, and each row carries a tenant_id column. Every query filters by it.
This is the cheapest model to run and the easiest to scale to many small tenants. One migration updates everyone, and reporting across tenants is simple. The risk is also simple: one missing filter leaks data. Isolation depends on every query in the codebase being right, so it must be enforced by structure, not by care.
Schema per tenant
Each tenant gets its own schema (a namespace of tables) inside one database. Queries run against the tenant's schema, so a missing filter cannot reach another tenant's tables.
Isolation is stronger, and restoring or exporting one tenant is easier. The costs arrive with numbers. Migrations must run once per schema, and a migration that fails halfway leaves tenants on different versions. Thousands of schemas strain connection pools and database catalogues. It suits products with tens or hundreds of sizeable tenants better than thousands of small ones.
Database per tenant
Each tenant gets its own database, sometimes its own server. This gives the strongest isolation on all three sides: separate performance, separate backups, the option of placing a tenant's data in a specific region, and a clean answer to a customer's security review.
It is also the most expensive to run. Every database needs provisioning, monitoring, backups and migrations, so this model only works with thorough automation. It fits a small number of large or regulated customers, and many products offer it as a premium tier on top of a shared pool.
Row-level security as a safety net
In the shared schema model, the database itself can enforce the tenant filter. PostgreSQL's row-level security lets you attach a policy to a table so that a session only sees rows matching its current tenant, typically read from a session setting that the application sets at the start of each request.
This turns a forgotten WHERE clause from a data leak into an empty result. A few rules make it hold:
- The application connects as a role that cannot bypass the policies. Table owners and superusers skip them by default.
- The tenant setting is set per transaction and cleared afterwards, so a pooled connection never carries one tenant's context into the next request.
- Policies cover every tenant-owned table, and a test checks that no such table lacks one.
Row-level security is a second line, not the only one. The data access layer should still scope every query by tenant, so both have to fail before anything leaks.
Noisy neighbours
In any shared model, tenants compete for the same CPU, memory, connections and disk. A single tenant running a large import or an expensive report can slow the product for everyone.
The defences are layered:
- Rate limits and quotas per tenant on API calls, background jobs and storage, tied to their plan.
- Fair queues for background work, so one tenant's backlog of jobs cannot starve the rest.
- Query budgets: timeouts, pagination limits and indexes that start with the tenant id.
- Per-tenant metrics, so you can see who is using what, rather than only that the system is slow.
- A way out: the ability to move a heavy tenant to its own database or pool without changing code.