Publishing now sits on both sides of generative AI. Inbound, editors and platforms receive text and images that machines wrote or helped write, sometimes in volumes that swamp human review. Inside, staff use models to triage submissions, summarise manuscripts, suggest edits, translate and check references. Outbound, readers, retailers and regulators increasingly expect to be told when AI was involved. Each stage has security problems that are easy to miss because they look like editorial policy rather than engineering.
This article treats a publisher's or publishing platform's workflow as a system to secure. It covers the threat model, an architecture with a quarantine boundary, rate limits for submission floods, detecting instructions hidden in manuscripts, keeping unpublished work confidential, catching fabricated citations, and attaching disclosure and provenance metadata to what you publish. It is engineering guidance, not legal advice.
The threat model
Start by listing what you protect and who attacks it. The useful assets are the integrity of editorial decisions, the confidentiality of unpublished work, the accuracy of what you print, the capacity of your reviewers, and your standing with retailers and readers.
| Threat | Example | Primary control |
|---|---|---|
| Volume flood | hundreds of generated submissions a week | per-account quotas, triage queue, disclosure |
| Injection via manuscript | hidden text telling an AI reviewer to praise the work | quarantine extraction, no-authority model |
| Confidentiality leak | an unpublished manuscript pasted into a consumer chatbot | approved endpoints, data classes, policy |
| Fabricated content | references to papers that do not exist | automated citation lookup |
| Undisclosed AI use | generated cover art sold as human-made | author attestations, provenance metadata |
| Rights exposure | text reproducing copyrighted passages | similarity checks, contract terms |
An editorial pipeline with a quarantine boundary
The architecture has one rule: manuscripts are untrusted input, and no model that reads them gets authority. Files go through a quarantine extractor that renders them, records anything invisible to a human reader, and passes on only visible text. The assist model is reached through an approved endpoint, has no tools and no write access, and returns structured notes to a human who makes every decision.
Submission floods
Floods are the most visible problem. In February 2023 the science-fiction magazine Clarkesworld closed submissions after a surge of machine-written stories; its editor reported roughly 500 such submissions against about 700 legitimate ones that month. Amazon's Kindle Direct Publishing responded to the same pressure on the self-publishing side with a limit of three new titles per day per account (adjustable, with exceptions on request) and separate disclosure questions for AI-generated text, images and translations.
The lesson for any intake is that detection alone does not scale. AI-text detectors give unreliable verdicts on short texts and have been shown to flag non-native English writers more often, so they cannot be the gate. Rate controls can: quotas per account and per payment instrument, lower limits for new accounts, a mandatory disclosure question whose false answer is a terms violation, and a triage queue that lets editors work by reputation. These controls cost honest authors almost nothing and make volume abuse expensive.
Instructions hidden in manuscripts
In July 2025 Nikkei reported finding instructions hidden in 17 arXiv preprints from 14 institutions in eight countries, written in white or tiny text and telling any AI reviewer to give a positive review; Nature separately identified 18 preprints with similar text. It is a textbook indirect prompt injection: the attacker controls a document that a model will read on someone else's behalf.
Defend in two layers. First, extract what a human sees and record what they do not. The detector below reads PyMuPDF span data and flags text that is tiny, matches the background colour, is transparent, or sits off the page. Treat any finding as a fact for the editor, not an automatic rejection: authors sometimes leave invisible artefacts from templates, and the content of the hidden text tells you which case you have. For Word and LaTeX sources, apply the same idea to hidden runs, comments and tiny fonts, or render to PDF first and use one detector.
import fitz # PyMuPDF
def luminance(rgb_int):
r, g, b = (rgb_int >> 16) & 255, (rgb_int >> 8) & 255, rgb_int & 255
return (0.2126 * r + 0.7152 * g + 0.0722 * b) / 255
def hidden_spans(path, min_size=4.0, bg_lum=1.0):
"""Flag text a human reader will not see. Assumes white pages; render
and sample the background per span if your inputs have coloured pages."""
findings = []
with fitz.open(path) as doc:
for pno, page in enumerate(doc, start=1):
area = page.rect
for block in page.get_text("dict")["blocks"]:
for line in block.get("lines", []):
for span in line["spans"]:
t = span["text"].strip()
if not t:
continue
why = []
if span["size"] < min_size:
why.append("tiny font")
if abs(luminance(span["color"]) - bg_lum) < 0.05:
why.append("same colour as background")
if span.get("alpha", 255) < 30:
why.append("transparent")
if not fitz.Rect(span["bbox"]).intersects(area):
why.append("off page")
if why:
findings.append({"page": pno, "why": why, "text": t[:200]})
return findingsSecond, assume extraction will miss something, because it will: text in images, metadata fields, and tricks with fonts that map glyphs to different characters. The model that reads the manuscript must therefore be unable to act on what it reads, which is the next section.
Keeping the model's authority small
An injected instruction is only dangerous in proportion to what the model can do. Keep that small. The assist model should produce structured notes, such as a summary, a list of claims needing checks and a list of unclear passages, and never a score, an accept or reject recommendation, or an email. It should have no tools, no retrieval over other manuscripts, and no ability to write to the submission system. Put the manuscript in a clearly delimited data field, tell the model its contents are material to analyse, not instructions, and validate the output against a schema so a hijacked response that drifts into praise or commands is rejected by code.
Then test it. Keep a small red-team set of manuscripts with hidden and visible injections, run every prompt or model change against it, and check that the structured notes neither change tone nor mention the injected request. A reviewer-facing UI that shows the hidden-text report next to the model's notes makes successful injections visible to the person who decides.
Confidentiality of unpublished work
Unpublished manuscripts, peer reviews and grant proposals are confidential by contract and by norm. The United States National Institutes of Health made this explicit in June 2023 (notice NOT-OD-23-149): its peer reviewers may not use generative AI to analyse or write critiques, and uploading application content to such tools breaches review confidentiality. Many journals have similar reviewer policies.
The engineering response is to give staff a sanctioned path so they do not improvise. Classify content (public, embargoed, confidential review material), route each class only to endpoints whose contract excludes training on your data and states retention, log which documents went to which model, and block consumer AI domains for the confidential class at the network or browser level. For external reviewers you cannot control, state the policy in the invitation and in the review form, and ask for an attestation.
Fabricated citations
Language models produce plausible references that do not exist, and authors who draft with them sometimes submit those references unchecked. Checking every reference by hand is slow; checking them automatically against a bibliographic index is cheap. The sketch below queries Crossref's public works API with the free-text reference and tests whether the best match's title appears in it. A not-found result is a prompt for the copy editor, not proof of fabrication: books, preprints, reports and non-English works are often missing from Crossref, so route them to other indexes or to a person.
import requests
from difflib import SequenceMatcher
API = "https://api.crossref.org/works"
def check_reference(ref_text, mailto="editorial@example.org", threshold=0.85):
"""Look up a free-text reference; return the best match or a flag."""
r = requests.get(API, params={"query.bibliographic": ref_text, "rows": 3,
"mailto": mailto}, timeout=20)
r.raise_for_status()
best = None
for item in r.json()["message"]["items"]:
title = (item.get("title") or [""])[0]
score = SequenceMatcher(None, title.lower(), ref_text.lower()).find_longest_match(
0, len(title), 0, len(ref_text)).size / max(1, len(title))
if best is None or score > best[0]:
best = (score, title, item.get("DOI"))
if best is None or best[0] < threshold:
return {"status": "not_found", "ref": ref_text}
return {"status": "found", "doi": best[2], "title": best[1], "score": round(best[0], 2)}
Disclosure and provenance on the way out
What you publish should say how it was made, in a form machines can read. For images, the IPTC digital source type vocabulary has a term, trainedAlgorithmicMedia, for content created by a trained model, and C2PA manifests can carry it in their actions, signed and bound to the file; C2PA for AI-generated content covers the mechanics. For text there is no comparably robust carrier: metadata in EPUB or on a product page is easy to strip, and text watermarks are fragile under editing. Record disclosure in your own catalogue as the system of record and emit it everywhere you can.
Retail platforms already ask. KDP's disclosure questions distinguish AI-generated content from AI-assisted editing, and the answer is required from the author, so your contracts should require authors to tell you. In the European Union, the AI Act's Article 50 transparency duties began to apply in August 2026, including marking of synthetic content and labelling of AI-generated text published to inform the public on matters of public interest; check the current text and any transitional provisions with counsel, since they affect generator providers and deployers differently. Rights questions about training data are a separate topic, covered in copyright and AI training.
Worked example: a journal's first month
A mid-sized journal adds an AI assist step that summarises each new submission for the handling editor. In the first month the quarantine extractor flags 41 of 1,900 PDFs. Thirty-five are template artefacts such as white page numbers and hidden form fields. Six contain sentences addressed to reviewers or language models, two of them asking for a positive assessment. Because the assist model only returns a schema-checked summary and claim list, none of the six changed its output, and the editors see the hidden-text report next to each summary. The journal writes to the authors, applies its misconduct policy, and adds the six files to its red-team set.
The citation checker, run at acceptance, flags 3 per cent of references. Most are books and conference talks missing from Crossref; eleven references across four manuscripts cannot be found anywhere, and the authors withdraw or correct them before publication. The figures are illustrative; the point is that each control produced evidence a person could act on, and none of them made a decision on its own.
Failure modes
- Detector as gate. Rejecting on an AI-text detector score punishes honest and non-native authors and is easy to evade. Use rate limits and disclosure instead.
- Extraction only in the happy path. The hidden-text check runs on PDFs, but editors paste Word files straight into the assist tool. Route every format through it.
- Authority creep. A summary tool gains a recommend-decision field, then an auto-reject threshold, and injected text now has a target.
- Shadow AI. With no sanctioned endpoint, staff paste confidential manuscripts into whatever chatbot is open.
- False certainty from lookups. Not-found citations are treated as fabricated; authors of legitimate but unindexed works are accused wrongly.
- Metadata loss. Disclosure is captured at intake and dropped by the production pipeline or the retailer feed.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Quotas vs open intake | flood resistance | friction for prolific honest authors; needs an exception path |
| Assist LLM vs none | faster triage and summaries | injection surface and confidentiality risk to manage |
| Strict quarantine extraction | hidden instructions surfaced | false positives from templates; editor time |
| Hosted vs self-hosted model | quality and no ops burden | contract and residency review; vendor dependency |
| Automated citation check | catches fabricated references early | unindexed works need manual follow-up |
What to do next
- Write the threat table for your own workflow: assets, attackers, current controls.
- Add per-account submission quotas, a stricter limit for new accounts and a disclosure question.
- Route every incoming file through a quarantine extractor and show its hidden-text report to editors.
- Restrict the assist model to schema-checked notes with no tools, scores or decisions.
- Publish an AI use policy for staff and reviewers and provide an approved endpoint.
- Run automated citation lookups at acceptance and send not-found results to a copy editor.
- Store AI-use disclosure in your catalogue and emit it in IPTC, C2PA and retailer metadata.
- Keep a red-team set of injected manuscripts and test every model or prompt change against it.