Teams that build on hosted models often feel that their prompt is the product. A system prompt that took months of iteration encodes tone, domain rules, edge cases and formatting that a competitor would like to copy. Then someone posts it on a forum, extracted with a single clever question, and the team learns two things at once: the prompt was never secret in the way they assumed, and most of the copies came from somewhere other than the chat box.
This article is about protecting prompt intellectual property as an engineering problem. It does not cover the attack catalogue, which the site's guide to system prompt stealing treats in depth. Instead it covers what is actually worth protecting, the five routes by which prompts escape, an architecture that keeps the valuable parts out of reach, output-side leak detection with code, provenance techniques for proving a copy is yours, and the realistic limits of all of it.
What is actually worth protecting
Start by being honest about where the value lies. The literal text of a system prompt is the most copyable part of an LLM product and usually the least defensible. A competitor who obtains it gets your wording, but not the evaluation set that told you which wording works, not the retrieval corpus and tools that the prompt orchestrates, not your fine-tuning data, and not the feedback loop that keeps improving it. A prompt is typically a few thousand tokens that took a long time to discover but take seconds to read.
That leads to a design principle: assume any text placed in the model's context can be read by a determined user, and move value out of that text. What remains in the prompt should be instructions you can live with being public. The genuinely sensitive material, such as pricing logic, scoring rules, partner names, internal thresholds and detection heuristics, belongs in code, tools and data that the model calls but never sees in full.
Classify prompt content explicitly. Persona, tone and formatting rules are low sensitivity. Domain procedures and worked examples are medium. Business rules, thresholds and anything that would help someone game your system are high. Credentials are never acceptable in a prompt.
Five routes by which prompts escape
The extraction attacks get the headlines, but in incident reviews prompts usually escape through mundane paths. The diagram shows the five routes and the table maps each to its control.
| Route | How it happens | Control |
|---|---|---|
| Client bundle | Prompt shipped in a mobile app, browser JavaScript or a desktop binary that calls the model directly | Never call the model from the client with your own prompt; proxy through a server |
| Logs and traces | Full request bodies, including the system prompt, written to application logs or tracing spans | Log a prompt version ID and hash, not the text; redact at the logger |
| Vendors and evals | Third-party observability, eval or annotation tools receive full prompts | Contractual terms, version-only logging, self-hosted options for high-sensitivity prompts |
| Repos and insiders | Prompts in widely readable repositories, tickets, chat threads and slide decks | A registry with access control and audit; prompts out of general source control |
| Model output | Users ask the model to reveal its instructions, directly or indirectly | Minimise what is in context, output leak detection, monitoring |
Architecture: registry, assembler, and rules behind tools
The reference architecture has four parts. A prompt registry stores templates as versioned artefacts with an owner, a sensitivity class and an access policy; reads are audited. A server-side assembler fetches the template at runtime, fills it with per-request data, attaches tools, and sends it to the model. Clients only ever send user input and receive answers. An output filter checks responses for leaked prompt content before they leave. Logging records the template version and a hash, which is enough to reproduce a request from the registry without duplicating the text into every log line.
The most effective single move is splitting. Take the high-sensitivity rules out of the prompt and put them behind a tool. Instead of a prompt that says to offer a discount of a certain percentage when certain conditions hold, the prompt says to call a quote tool for prices and discounts, and the tool runs your real logic in code. An extraction attack against that model reveals that a quote tool exists, which is harmless. A secondary benefit is that logic in code is testable and deterministic, while logic in a prompt is neither.
# Assemble prompts server-side; the client never sees templates.
import hashlib
class PromptRegistry:
def __init__(self, store, audit):
self.store, self.audit = store, audit # store: encrypted KV; audit: append-only log
def get(self, name, version, caller):
rec = self.store.read(f"{name}@{version}") # {text, owner, sensitivity, readers}
if caller.service not in rec["readers"]:
raise PermissionError(f"{caller.service} may not read {name}")
self.audit.write(name=name, version=version, caller=caller.service)
return rec["text"]
def build_request(registry, user_msg, ctx):
tmpl = registry.get("support_agent", ctx.prompt_version, ctx.caller)
system = tmpl.format(product=ctx.product, locale=ctx.locale)
fp = hashlib.sha256(system.encode()).hexdigest()[:16]
ctx.log.info("llm_request", prompt="support_agent", version=ctx.prompt_version, sha=fp)
return {"system": system,
"messages": [{"role": "user", "content": user_msg}],
"tools": ctx.tools} # business rules live behind tools
Detecting leaks in model output
Instructions such as never revealing the system prompt reduce casual leaks but do not hold against persistent users, so add a check on the output. The simple and effective approach is to measure overlap between the response and the protected prompt: break both into word shingles, sequences of n consecutive words, and compute what fraction of the prompt's distinctive shingles appear in the answer. Exclude shingles that are common in normal answers, such as greeting templates, by computing them from a sample of benign traffic. Then block, truncate or flag responses above a threshold.
import re
def shingles(text, n=6):
words = re.findall(r"[a-z0-9']+", text.lower())
return {" ".join(words[i:i + n]) for i in range(len(words) - n + 1)}
class LeakDetector:
def __init__(self, protected_texts, benign_sample, n=6):
self.n = n
common = set().union(*(shingles(t, n) for t in benign_sample))
self.protected = [shingles(t, n) - common for t in protected_texts]
def score(self, response):
got = shingles(response, self.n)
# max over protected texts: fraction of each one's shingles present
return max((len(p & got) / len(p) for p in self.protected if p), default=0.0)
det = LeakDetector([system_prompt, *tool_descriptions], benign_answers)
s = det.score(model_output)
if s > 0.15:
action = "block" # replace with a refusal and alert
elif s > 0.05:
action = "flag" # deliver, but log for reviewOverlap detection catches verbatim and near-verbatim leaks, which are most of what gets posted. It does not catch paraphrase (a model asked to summarise its instructions), translation, or encoding tricks such as asking for the prompt one word per line or in base64. A second cheap layer covers some of those: normalise the output by joining lines and decoding obvious encodings before shingling. A third, more expensive, is a classifier or LLM judge asked whether a response describes the system's own instructions. Use it on a sample, or only on conversations already flagged by rate-based signals such as many meta-questions in a short session. Pair this with rate limiting, because extraction is iterative and slowing it down raises its cost.
Provenance: canaries and fingerprints
When a copy appears, you will want to show it came from you. Two techniques help. A canary is a unique, meaningless string or instruction placed in the prompt, for example an internal reference code, that has no reason to exist elsewhere. Its appearance in a competitor's product or a public dump is strong evidence of copying, and you can search for it routinely. The site's article on canary tokens covers designs that also alert on use. A fingerprint is subtler: deliberate, harmless idiosyncrasies in wording or examples that vary per deployment or per tenant, so a leaked copy also tells you which copy leaked. Keep a record of which variant went where, and date-stamp registry versions so you can show when each text existed.
Two cautions. Canaries reveal themselves once the prompt is public, so rotate them after a leak. And a determined copier can paraphrase away both canaries and fingerprints. Provenance helps with lazy copying and insider leaks, which in practice are the common cases.
The legal footing, briefly
Whether a prompt is protectable as intellectual property is unsettled and depends on jurisdiction, and this is not legal advice. Copyright protection for short functional instructions is uncertain. Trade-secret protection, where it applies, generally requires that the owner took reasonable measures to keep the information secret. That is where engineering matters: access-controlled storage, audit logs, confidentiality terms with vendors and staff, and avoiding shipping the prompt to clients are the kind of evidence such a claim relies on. Publishing a prompt in a mobile app bundle undercuts that argument. The site's guide to AI and trade secrets maps legal expectations to controls; involve counsel before relying on any of this.
Worked example: a contract-review assistant
A legal-tech startup ships a contract-review assistant. Its system prompt is about 3,000 tokens: persona and formatting, a clause taxonomy, a scoring rubric with weights, and fifteen worked examples. A review finds three problems: the desktop app calls the model directly with the prompt embedded in the binary; full request bodies go to a third-party tracing service; and the prompt lives in the main repository, readable by 140 people.
The fix takes two sprints. All model calls move behind the company's API, and the desktop app sends only documents and questions. The rubric and its weights move into a scoring tool implemented in code; the prompt now tells the model to extract clause features and call the tool, which returns scores. The worked examples are reduced to five generic ones, and the specialised examples move to retrieval, fetched per clause type. Prompts move into a registry readable by the serving service and six editors, with audited reads. Tracing records version IDs and hashes. A leak detector built from the remaining prompt and tool descriptions runs on every answer, and a canary reference code sits in the prompt.
Six months later an extraction attempt succeeds in getting the model to paraphrase its instructions. What leaks is a persona and an instruction to call a scoring tool. The rubric, the weights and the specialised examples, the parts that took the longest to build, were never in context. (The scenario is illustrative.)
Failure modes
| Failure mode | What happens | Mitigation |
|---|---|---|
| Secrets in prompts | API keys or internal URLs leak with the prompt | Secret scanning on the registry; credentials only in tools |
| Detector too strict | Legitimate answers that quote policy text are blocked | Subtract benign shingles; flag before you block |
| Detector too narrow | Paraphrased or encoded leaks pass | Normalisation, sampled LLM judge, session-level signals |
| Tool descriptions overlooked | Tool schemas reveal the logic you moved out | Treat tool descriptions as part of the prompt; keep them generic |
| Registry bypass | Copies of prompts drift back into repos and tickets | Pre-commit scanning for registry hashes and canaries |
| False confidence | Team treats instructions not to reveal the prompt as protection | Assume context is readable; design accordingly |
Trade-offs
Every control has a cost. Moving logic into tools adds latency and a round trip per call, and the model can misuse tools in ways it would not misuse inline text. Version-only logging makes debugging slower, because engineers must fetch the template to see what the model saw; give them a privileged tool for that instead of putting text back in logs. Output filtering adds a few milliseconds and some false positives, and on streaming responses it must run on a rolling buffer, which delays the first visible tokens slightly. Weigh these against what a leak would actually cost you. For a hobby project with an ordinary persona prompt, the right answer may be to publish the prompt and compete on everything else.
What to do next
- Classify each section of your production prompts as low, medium or high sensitivity, and remove any credentials immediately.
- Verify that no client, whether web, mobile or desktop, sends your system prompt to a model provider; route every call through your server.
- Grep logs, traces and third-party tools for prompt text; switch to logging version IDs and hashes.
- Move high-sensitivity rules and thresholds behind tools implemented in code, and review tool descriptions for leaks.
- Put prompts in a registry with owners, access control and audited reads, and scan repositories for stray copies.
- Deploy the shingle-overlap leak detector on outputs in flag-only mode for two weeks, tune the threshold, then enable blocking above it.
- Add a canary to each high-value prompt, set up a recurring search for it, and record which variant went to which deployment.
- Ask counsel what reasonable measures mean for your jurisdiction and keep evidence of the controls you run.