A trade secret is the odd one out among intellectual property rights. There is no registration, no examiner and no expiry date. A formula, a manufacturing tolerance, a pricing model or a customer list is protected for exactly as long as it stays secret and its owner keeps taking sensible steps to keep it that way. Disclose it carelessly once and the protection can be gone for good.
Generative AI puts pressure on both halves of that bargain. Employees paste source code and process notes into chat tools, retrieval systems index documents that used to sit in locked folders, agents copy data between systems, and the models you build become valuable secrets that can leak through an API. This article explains the legal test in engineering terms, maps the three ways trade secrets meet AI, shows how to build a gateway that classifies and routes sensitive content with an evidence trail, and walks through a worked rollout. It is not legal advice. Statutory references are to United States federal law and the EU directive as of October 2026; local law differs, and a lawyer should review your programme.
What makes information a trade secret
Under the US Defend Trade Secrets Act (DTSA), information qualifies as a trade secret when two conditions hold: the owner has taken reasonable measures to keep it secret, and it derives independent economic value from not being generally known or readily ascertainable by others who could profit from it (18 U.S.C. 1839(3)). The EU Trade Secrets Directive (2016/943) uses a closely parallel test: the information is secret, has commercial value because it is secret, and has been subject to reasonable steps to keep it secret. Most US states apply a similar definition through their own statutes.
Two consequences matter for engineers. First, the test is about conduct, not paperwork. A court asks what the owner actually did: who had access, whether documents were marked, whether staff were told, whether systems enforced the rules. A policy nobody follows is weak evidence. Second, the law protects against misappropriation by improper means, such as theft, bribery or breach of a duty of confidentiality. The DTSA states expressly that reverse engineering and independent derivation are not improper means (18 U.S.C. 1839(6)(B)). If a competitor can reconstruct your secret from what you expose publicly, trade secret law will usually not help you, and you must rely on contracts or technical barriers instead.
A chat tool, a retrieval index or a model API is therefore a new channel out of the circle of people bound to confidentiality, and each needs a control you can later show to a court.
Three flows where secrets meet AI
It helps to separate three flows, because each has a different owner and a different control.
- Outbound: your secrets into AI systems. Prompts, pasted code, uploaded files, retrieval corpora, fine-tuning datasets and agent tool calls all move information to a model and often to a third party that runs it. The question is whether that transfer happens under terms that preserve confidentiality, and whether you can prove it.
- Your AI assets as secrets. Model weights, training data recipes, data-cleaning pipelines, evaluation sets, system prompts and tool schemas can all be valuable because competitors do not have them. Some of these are exposed every time the product answers a question.
- Inbound: other people's secrets into your models. A new hire who brings a previous employer's code, a dataset scraped from a partner portal, or a vendor document fed into a retrieval index can put someone else's trade secret into your product. That creates liability for you, and it is hard to remove once it is baked into weights.
The first flow is usually an everyday productivity shortcut, not a malicious act, so controls have to make the safe path the easy one.
Architecture of a secret-aware AI stack
The architecture has four parts. A secret registry lists what the organisation treats as a trade secret, who owns each item, and a sensitivity tier. An AI gateway sits between users or agents and every model endpoint; it classifies content, matches fingerprints of registered secrets, and routes the request to a destination whose terms are acceptable for that tier. An evidence log records decisions without storing the secret itself. Finally, separate controls protect the AI assets you build and screen inbound data.
Routing beats blocking: people keep the productivity gain, and you keep the evidence.
Reasonable measures, translated into controls
Courts look at the whole picture, so no single control is required, but the following translation from legal language to engineering work is a sound starting point.
| "Reasonable measures" element | Engineering control for AI | Evidence it produces |
|---|---|---|
| Identify what is secret | Secret registry with owners and tiers | Versioned registry, review history |
| Limit access to need-to-know | Per-tier model routing; RAG index ACLs mirror source ACLs | Gateway decisions, index permission audits |
| Bind recipients to confidentiality | Only vendors with no-training and confidentiality terms | Contract register linked to each endpoint |
| Tell people the rules | Inline gateway messages at the moment of use | Acknowledgement and training records |
| Monitor and respond | Fingerprint alerts, anomaly detection on bulk export | Alerts, tickets, incident timelines |
| Mark documents | Classification labels read by the gateway and indexer | Label coverage reports |
Contracts are part of the measure set. Employment and contractor agreements should contain confidentiality clauses that cover AI tools explicitly. In the US, an employer that wants the full DTSA remedies against an employee must include the whistleblower immunity notice described in 18 U.S.C. 1833(b) in agreements governing confidential information, so review templates when you add AI language. On the vendor side, read whether your inputs may be used for training, how long they are retained, and who can review them; the companion article on governance explains how to keep that register current.
A fingerprinting gateway in code
The core of the gateway is a classifier that combines labels with fingerprints. Labels catch documents that were marked; fingerprints catch secrets that were copied out of their marked home into a chat box. The example below hashes overlapping word shingles of each registered secret and stores only the hashes, so the gateway never needs a plaintext copy of the crown jewels.
import hashlib, re
from dataclasses import dataclass
SHINGLE = 8 # words per shingle; tune per corpus
def shingles(text: str):
words = re.findall(r"[A-Za-z0-9_]+", text.lower())
for i in range(max(0, len(words) - SHINGLE + 1)):
chunk = " ".join(words[i:i + SHINGLE])
yield hashlib.blake2b(chunk.encode(), digest_size=8, key=b"rotate-me").hexdigest()
@dataclass
class Secret:
secret_id: str
owner: str
tier: int # 3 = crown jewel, 2 = confidential, 1 = internal
hashes: frozenset
def register(secret_id, owner, tier, text):
return Secret(secret_id, owner, tier, frozenset(shingles(text)))
ROUTES = {3: "selfhosted", 2: "vendor_contracted", 1: "vendor_contracted", 0: "any_approved"}
def decide(prompt: str, label_tier: int, registry: list, min_hits: int = 3):
found = set(shingles(prompt))
hits = [(s.secret_id, s.tier, len(found & s.hashes)) for s in registry]
matched = [h for h in hits if h[2] >= min_hits]
tier = max([label_tier] + [t for _, t, _ in matched])
return {
"route": ROUTES[tier],
"tier": tier,
"matched": [sid for sid, _, _ in matched],
"evidence": {"prompt_sha256": hashlib.sha256(prompt.encode()).hexdigest(),
"shingles": len(found)},
}Three design choices are deliberate. The shingle hash is keyed, so someone who obtains the hash list cannot test guesses offline. The evidence record stores a digest of the prompt rather than the prompt, so the log does not become a second copy of the secret. And the routing table maps tiers to destinations, so changing a vendor is a configuration change, not a code change.
Fingerprints only catch near-verbatim copies. A paraphrase of a process recipe will slip past them, which is why the gateway also honours labels, why retrieval indexes inherit source permissions, and why staff training matters. Treat the fingerprint layer as a tripwire for the most valuable items, not as a complete data loss prevention system.
Protecting the AI assets you build
The models and pipelines you build can themselves be trade secrets, but only if you treat them that way. Weights should live in storage with the same access control as source code for your most sensitive systems: named owners, short-lived credentials, download alerts and no public buckets. Training recipes, data mixtures and evaluation sets usually carry more competitive value than any one checkpoint, and they are often scattered across notebooks and shared drives. Register them.
Exposed products leak information by design. Every answer reveals something about the model and its prompt. Because reverse engineering is a lawful means under the DTSA, a competitor who reconstructs behaviour by querying your public API may not be misappropriating anything under trade secret law; your protection there comes from terms of service, rate limits and technical barriers. Practical consequences follow:
- Assume a system prompt can be extracted. Keep business logic, credentials and secret thresholds in code behind the tool layer, not in prompt text.
- Rate-limit and monitor for extraction patterns: very high query volume, systematic coverage of an input space, or requests for probabilities and token-level detail you do not need to expose.
- Return only what the product needs. Log-probabilities, raw retrieval chunks and verbose tool traces are convenient for debugging and generous to an attacker.
- Put use restrictions in your customer terms and keep evidence that users accepted them.
Keeping other people's secrets out
Inbound contamination is the flow teams forget. If a model is fine-tuned on code a new hire brought from a competitor, the competitor's secret is now in your weights and potentially in your outputs. Removing it may mean retraining. Controls that work in practice:
- Record provenance for every dataset and document collection: source, licence or agreement, who approved it, and the date. Refuse data with no provenance.
- Ask new staff to certify, as part of onboarding, that they have not brought prior employers' confidential material, and keep their early commits under normal code review.
- Screen partner and customer documents before indexing them for retrieval, and keep them in tenant-scoped indexes so one customer's material cannot surface in another's answers.
Worked example: a manufacturer's rollout
Consider a contract manufacturer with about 2,000 staff. Its most valuable assets are furnace profiles and inspection thresholds for a set of alloys; its code base is ordinary. It wants a coding assistant for engineers and a retrieval assistant over process documentation.
The team starts by registering secrets: 40 process recipes (tier 3), customer drawings and pricing (tier 2), and general engineering documentation (tier 1). The recipe owners supply the documents once; the gateway stores keyed shingle hashes. The coding assistant routes to a contracted vendor whose terms exclude training on inputs; the process assistant runs on a self-hosted model inside the plant network, with a retrieval index whose permissions are copied from the document system nightly.
In the first month the gateway records 61 fingerprint matches. Fifty-five are engineers pasting recipe fragments into the coding assistant while writing control software; these are rerouted to the self-hosted model automatically, with a short explanation. Six are bulk pastes from a single account in one evening, which triggers an incident review and turns out to be a contractor preparing a handover document. Nothing left the boundary, and the company now has dated evidence of its measures, which is precisely what it would need if a recipe ever surfaced at a rival.
Failure modes
- Block-only gateways. Staff switch to personal accounts on personal phones, and the organisation loses both the control and the evidence.
- Logs that copy the secret. Full prompt logging in a third-party observability tool can be a bigger disclosure than the original prompt. Log digests and decisions, and keep any full-text capture inside the boundary with short retention.
- Retrieval indexes that flatten permissions. A single index built with a service account lets anyone who can ask a question read documents they could never open directly.
- Secrets in prompts. Pricing rules or partner credentials embedded in a system prompt will eventually be extracted by a user.
- A registry nobody maintains. A list compiled once for a policy launch drifts until it covers last year's products. Assign owners and review dates.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Self-hosted model for crown jewels | Data never leaves your boundary | GPU spend, weaker models, patching duty |
| Vendor with no-training terms | Strong models, low operations work | Reliance on contract and vendor security |
| Fingerprint matching | Catches verbatim copies cheaply | Misses paraphrase; needs owner effort |
| Full-text prompt logging | Rich forensics | Creates another copy of the secret |
| Strict extraction limits on public APIs | Harder reverse engineering | Friction for legitimate heavy users |
What to do next
- Write a secret registry this month: top 20 items, owners, tiers and review dates.
- Put every model endpoint behind one gateway and route by tier instead of only blocking.
- Add keyed fingerprints for tier-3 items and alert on bulk or repeated matches.
- Check each AI vendor contract for training use, retention, human review and confidentiality terms, and record the result next to the endpoint.
- Make retrieval indexes inherit source permissions and test with a low-privilege account.
- Move secrets and business rules out of system prompts and into the tool layer.
- Add AI language and the required whistleblower notice to employment and contractor templates.
- Record provenance for every dataset and run an inbound certification for new hires.
Related reading on this site: employee AI usage policies, data exfiltration via LLM tools, system prompt leakage, AI and copyright and data governance for AI.