A multi-tenant LLM product puts many customers' documents, conversations and tool credentials behind one model endpoint. The model itself has no idea who is asking. It completes whatever text lands in its context window, so if a retrieval query, a cache lookup or a tool call fetches tenant B's data while serving tenant A, the model will quote it fluently and nothing in the response marks it as a leak.

Our tenant isolation architecture overview covers the threat model and the menu of isolation strengths. This article is the implementation companion: how to carry tenant identity through the stack, how to make the database enforce it rather than trusting every query author, how to key caches and scope tools, and how to prove isolation with a test suite that runs on every deploy. Code is Python and PostgreSQL with pgvector, but the patterns carry to any stack.

Advertisement

Inventory every surface that holds tenant data

Isolation fails at the surface nobody listed. Before writing code, inventory every place tenant data is stored or derived, because each one needs its own enforcement point. A typical retrieval-augmented assistant has more than teams expect:

SurfaceHolds tenant data asEnforcement point
Document and chunk storeRaw text, embeddings, metadataRow-level security or per-tenant schema
Vector indexEmbeddings, which can be inverted toward source textSame table as chunks, or a namespace per tenant
Conversation historyPrompts and completionsPrimary key includes tenant id
Response and semantic cachesCompletions keyed by promptTenant id inside the cache key
Inference prefix cacheKV blocks for shared prefixesPer-tenant salt, or no cross-tenant sharing
Tool credentialsOAuth tokens, API keysVault path per tenant, fetched by the tool runtime
Logs, traces, eval setsFull prompts and outputsTenant tag plus partitioned retention
Fine-tuned adaptersTraining data baked into weightsOne adapter per tenant, never pooled

The last two rows are the offline paths, and they leak without any attacker: an engineer copies production traces into an eval set, or a pooled fine-tune memorises one customer's contract terms and recites them to another.

Derive tenant context once and carry it everywhere

The tenant id must come from a verified credential, never from a request body, header or anything the model produced. Derive it once at the edge, freeze it in an immutable object, and make it impossible to call a shared subsystem without it. In Python, a ContextVar carries it across async calls without threading it through every signature, and a missing context fails closed:

ClientJWT with tid claimGatewayverify, derive tenantOrchestratorContextVar ctxRetrievalPostgres RLSCacheskey = tenant + hashToolsscoped credentialsModel servershared GPU, salted prefixLogs and tracestenant-tagged, partitionedCanary test suiteprobes every surface1 request2 signed ctx34 every write tagged5 plant and probeTenant identity is derived once, carried everywhere, and enforced by the store, not by the prompt
Tenant context flow: the gateway derives the tenant from a verified credential, every shared subsystem receives it as a mandatory argument and enforces it itself, and a canary suite continuously probes each surface for cross-tenant reads.
from contextvars import ContextVar
from dataclasses import dataclass
from uuid import UUID

@dataclass(frozen=True)
class TenantContext:
    tenant_id: UUID
    user_id: str
    tier: str            # "shared" or "dedicated"

_ctx: ContextVar[TenantContext] = ContextVar("tenant_ctx")

async def auth_middleware(request, call_next):
    claims = verify_jwt(request.headers["authorization"])   # signature, issuer, audience, expiry
    token = _ctx.set(TenantContext(UUID(claims["tid"]), claims["sub"],
                                   claims.get("tier", "shared")))
    try:
        return await call_next(request)
    finally:
        _ctx.reset(token)

def current_tenant() -> TenantContext:
    try:
        return _ctx.get()
    except LookupError:
        raise RuntimeError("no tenant context: refusing to touch shared state")

Background jobs are where this breaks. A queue worker has no HTTP request, so the producer must serialise the tenant id into the job, and the worker must set the context from a signed job envelope before doing anything. Signing matters: an unsigned envelope lets anyone who can write to the queue claim any tenant. A lint rule that flags direct database handles in application code keeps the pattern from eroding.

Advertisement

Make the database enforce it: row-level security with pgvector

Application-level WHERE tenant_id = ? clauses fail the first time someone writes a new query and forgets one. Row-level security moves the filter into the database, so a query that omits it simply returns the caller's rows. With pgvector the chunks and their embeddings live in one table, which means one policy covers both keyword and vector search:

CREATE TABLE chunks (
  id         bigserial PRIMARY KEY,
  tenant_id  uuid   NOT NULL,
  doc_id     uuid   NOT NULL,
  body       text   NOT NULL,
  embedding  vector(1024) NOT NULL
);
CREATE INDEX ON chunks (tenant_id);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);

ALTER TABLE chunks ENABLE ROW LEVEL SECURITY;
ALTER TABLE chunks FORCE ROW LEVEL SECURITY;      -- owners obey the policy too

CREATE POLICY tenant_rows ON chunks
  USING      (tenant_id = current_setting('app.tenant_id')::uuid)
  WITH CHECK (tenant_id = current_setting('app.tenant_id')::uuid);

-- the application role: no BYPASSRLS, not the table owner, not a superuser
CREATE ROLE app_rw LOGIN PASSWORD '...';
GRANT SELECT, INSERT, UPDATE, DELETE ON chunks TO app_rw;

Three details decide whether this actually holds. First, table owners bypass RLS unless the table has FORCE ROW LEVEL SECURITY, and superusers and roles with BYPASSRLS always bypass it, so the application must connect as a dedicated role. Second, the tenant setting must be transaction-local. With a connection pool, a plain session-level SET survives into the next request that borrows the connection, which is a cross-tenant leak waiting for a race. Use set_config(name, value, true), which also accepts a bind parameter, inside the same transaction as the query:

def search_chunks(pool, query_vec, k=8):
    ctx = current_tenant()
    with pool.connection() as conn, conn.transaction():
        conn.execute("SELECT set_config('app.tenant_id', %s, true)", (str(ctx.tenant_id),))
        return conn.execute(
            "SELECT id, doc_id, body FROM chunks "
            "ORDER BY embedding <=> %s::vector LIMIT %s",
            (query_vec, k),
        ).fetchall()

Third, if the setting is missing or empty, the ::uuid cast raises an error. That is the behaviour you want: no context means no rows, loudly.

Filtering has a cost that surprises people. An HNSW scan returns its nearest candidates first and the policy filters them afterwards, so a small tenant in a large shared table can receive fewer than k results, or none. pgvector 0.8 and later can keep scanning until enough rows pass the filter (SET hnsw.iterative_scan = relaxed_order), at extra latency. Large tenants are often better served by a partition or a dedicated table, which also makes deletion provable: drop the partition.

Caches: response, semantic and prefix

A response cache keyed on the prompt alone is a cross-tenant data pipe: tenant A asks about "our refund policy", the answer is cached, and tenant B asking the same words gets A's policy. The fix is mechanical. Put the tenant id, and for user-private data the user id, inside the key, and hash a canonical form so no field can collide with another:

import hashlib, json

def cache_key(prompt: str, model: str, scope: str = "tenant") -> str:
    ctx = current_tenant()
    principal = str(ctx.tenant_id) if scope == "tenant" else f"{ctx.tenant_id}/{ctx.user_id}"
    material = json.dumps({"p": principal, "m": model, "q": prompt},
                          sort_keys=True, separators=(",", ":"))
    return "resp:" + hashlib.sha256(material.encode()).hexdigest()

Semantic caches need the same treatment plus one more rule: the similarity search must run inside the tenant's namespace, not across all cached prompts. The inference engine's prefix cache is subtler. Sharing KV blocks for an identical system prompt is harmless, but sharing blocks for a tenant's document prefix creates a timing side channel: a probe that hits the cache returns its first token measurably faster, which reveals that someone else sent that prefix. Mitigate by mixing a per-tenant salt into prefix-block hashes where your engine supports it, check its documentation, because option names differ between engines and versions, or by disabling cross-tenant prefix sharing for sensitive tiers.

Tools and agents: take tenancy away from the model

Agents turn the model into a confused deputy. If a tool accepts tenant_id or an account number as an argument, a prompt injection hidden in a retrieved document can ask for someone else's records, and the tool will comply with the platform's privileges. The defence is to remove tenancy from the model's control entirely. Tool schemas exposed to the model never contain tenant identifiers, and the runtime binds them from the context:

class TenantTools:
    # Built per request; the model sees only the method arguments.

    def __init__(self, docs_api, vault):
        self.ctx = current_tenant()
        self.docs = docs_api
        self.creds = vault.read(f"tenants/{self.ctx.tenant_id}/crm")  # per-tenant secret path

    def get_document(self, doc_id: str) -> dict:
        doc = self.docs.get(doc_id, tenant_id=self.ctx.tenant_id)   # tenant from ctx, never args
        if doc is None:
            raise ToolError("document not found")    # same error for missing and foreign ids
        return doc

Return the same error for a missing id and a foreign id, otherwise the tool becomes an oracle that confirms which ids exist elsewhere.

Logs, evals and fine-tuning: the offline paths

Online paths get reviewed; offline paths get copied. Tag every log line, trace span and stored completion with the tenant id at write time, and partition storage by tenant so retention and deletion requests can act on one partition.

Evaluation sets built from production traffic must record their source tenants and be excluded from any model or prompt shared across tenants. Fine-tuning is the hardest case because data in weights cannot be filtered at query time. Pooled fine-tunes on customer data are a contractual and technical risk. Per-tenant adapters, such as LoRA weights loaded per request, keep training data inside one tenant's boundary and make deletion a file removal rather than a retraining project. For more on what models memorise and how to measure it, see PII leakage and memorisation.

Proving isolation with a canary test suite

An isolation design is a hypothesis until a test tries to break it. The most effective technique is the canary: plant a unique, unguessable string in tenant A's data on every surface, then probe as tenant B through every read path and assert the string never appears. Run it in CI against a real database and cache, and in production against synthetic tenants:

import secrets, pytest

SURFACES = [DocumentSearch(), ConversationHistory(), ResponseCache(),
            SemanticCache(), AgentTools(), ExportApi()]

@pytest.fixture
def tenants(make_tenant):
    return make_tenant("a"), make_tenant("b")

@pytest.mark.parametrize("surface", SURFACES, ids=lambda s: type(s).__name__)
def test_no_cross_tenant_read(surface, tenants):
    a, b = tenants
    canary = f"CANARY-{secrets.token_hex(8)}"
    surface.plant(as_tenant=a, text=f"Internal code word is {canary}.")
    for probe in surface.probes(hint="internal code word"):
        out = surface.read(as_tenant=b, query=probe)
        assert canary not in out, f"{type(surface).__name__} leaked across tenants"
    # positive control: the owner can still see it, so the test is not vacuous
    assert any(canary in surface.read(as_tenant=a, query=q) for q in surface.probes("internal code word"))

The positive control matters: without it, a broken plant step makes every isolation test pass. Add three more classes of test. A no-context test calls each store with no tenant set and expects an error, not an empty result. A privilege test asserts the application role lacks BYPASSRLS and does not own the tables. A prompt-injection test plants a document instructing the agent to fetch another tenant's record and asserts the tool layer refuses. For injection defences on the retrieval side, see defending RAG pipelines.

Worked example: auditing a support-assistant platform

Consider a support-assistant platform with 400 tenants on one Postgres cluster and one vLLM-style serving fleet. An audit found three gaps. The semantic cache was keyed on embedding similarity alone, so near-identical questions from different tenants could share answers. A nightly re-indexing job connected as the table owner, which silently bypassed the new RLS policy. And the CRM tool took account_id as a model-visible argument.

The team fixed them in order of blast radius. They added the tenant id to the cache key and namespace, flushed the cache, and watched the hit rate fall from 31 percent to 22 percent. They applied FORCE ROW LEVEL SECURITY and moved the job to a role that sets the tenant per batch. They rewrote the CRM tool to resolve accounts within the tenant's own scope. The canary suite, added last, failed on its first run: the CSV export endpoint used a reporting replica with RLS disabled. That fifth surface had never been in the inventory, which is exactly why the suite exists.

Failure modes and trade-offs

  • Session-level settings on pooled connections. The next borrower inherits the previous tenant. Use transaction-local settings and reset connections on return.
  • Owner and superuser roles in jobs. RLS silently does nothing for them unless forced; audit every role that connects.
  • Tenant id from the client. Any header or body field is attacker-controlled; derive it from the verified credential only.
  • Filtered ANN starvation. Small tenants get empty retrieval and the model hallucinates; monitor result counts per tenant.
  • Unlisted replicas and exports. Read replicas, analytics copies and backups need the same controls or must hold no raw tenant data.

The trade-offs are real. Shared tables with RLS are cheap and simple but carry a small blast radius from any policy bug and filtered-search recall loss. Per-tenant schemas or databases isolate strongly and make deletion trivial, but multiply migrations, connections and index memory. Most platforms run a hybrid: pooled for the long tail, dedicated for regulated or very large tenants, behind the same context interface so the application code does not change. The system prompt is a surface too; see system prompt leakage for what not to put there.

What to do next

  1. Write the surface inventory for your product, including replicas, exports, eval sets and adapters, and name an enforcement point for each.
  2. Make tenant context come only from verified credentials, carry it in a context object, and fail closed when it is missing.
  3. Enable and force row-level security on every shared table; connect as a role without BYPASSRLS and set the tenant with transaction-local set_config.
  4. Put the tenant id in every cache key and semantic-cache namespace; decide explicitly whether the inference prefix cache may be shared.
  5. Strip tenant identifiers from all model-visible tool schemas and read third-party credentials from per-tenant secret paths.
  6. Build the canary suite with a positive control, run it in CI and against synthetic production tenants, and add a surface to it whenever you add a store.
  7. Track per-tenant retrieval result counts and cache hit rates so isolation costs are visible rather than discovered.
Key takeaway: The model cannot enforce tenant boundaries, so every store, cache and tool must. Derive tenant identity only from verified credentials, carry it in a context object that fails closed, and let the database enforce it with forced row-level security and transaction-local settings. Key caches by tenant, keep tenant ids out of model-visible tool arguments, and treat logs, evals and adapters as tenant data. Then prove it continuously with canary tests that include a positive control.