Large language models make persuasive, fluent, multilingual text nearly free to produce. For anyone running an influence operation, that removes what used to be a real cost: writers who could pass as locals in several languages, producing enough variants to fill hundreds of personas. The risk is real, but the obvious defence of detecting AI-written posts and removing them mostly fails, and the way it fails explains where defenders should spend their effort.
This article is for defenders: platform trust and safety engineers, provider integrity teams and researchers. It covers what LLMs change, why judging authorship is weak evidence, how to detect coordination instead (with runnable code), where provenance fits, and how to measure impact without harming legitimate users. Attacker techniques appear only as far as detection needs. Provider-side account abuse scoring has its own article, LLM Abuse Detection Architecture, and is referenced here rather than repeated.
What LLMs change about influence operations
An influence operation has a pipeline: build or buy personas, produce content, distribute it, gain engagement, and ideally break out into authentic conversation, so real people and media repeat the message. LLMs act almost entirely on the second stage.
| Stage | What LLMs change | What stays hard for the operator |
|---|---|---|
| Personas | Bios, backstories and consistent voices on demand | Accounts still need phone numbers, history and plausible activity |
| Content | Volume, fluency, translation, endless paraphrase | Knowing what will resonate with a real audience |
| Distribution | Comment replies at scale, tailored per thread | Posting still leaves timing, network and infrastructure traces |
| Engagement | Fake replies can simulate a conversation | Authentic engagement cannot be generated |
| Breakout | Little direct effect | Getting real people and outlets to carry the narrative |
Public evidence so far supports this split. In May 2024 OpenAI reported disrupting five covert influence operations that had used its models: two Russian (one it named Bad Grammar, plus the long-running Doppelganger), the Chinese network known as Spamouflage, the Iranian International Union of Virtual Media, and a commercial Israeli operation it called Zero Zeno. The models were used for writing comments and articles, translating, and in some cases debugging code for posting bots. OpenAI's assessment was that none of them gained significant authentic audience engagement from it. That is one company's view of the operations it saw, not a guarantee, but the pattern matters: generation got cheaper, and audience did not.
The risks worth planning for are therefore volume (flooding a hashtag, a comment section, a public consultation or a review site), language reach (operations that previously stayed in one language), and targeted micro-persuasion in direct conversation, which is harder to observe than public posting.
What LLMs change about influence operations
An influence operation has a pipeline: build or buy personas, produce content, distribute it, gain engagement, and ideally break out into authentic conversation, so real people and media repeat the message. LLMs act almost entirely on the second stage.
| Stage | What LLMs change | What stays hard for the operator |
|---|---|---|
| Personas | Bios, backstories and consistent voices on demand | Accounts still need phone numbers, history and plausible activity |
| Content | Volume, fluency, translation, endless paraphrase | Knowing what will resonate with a real audience |
| Distribution | Comment replies at scale, tailored per thread | Posting still leaves timing, network and infrastructure traces |
| Engagement | Fake replies can simulate a conversation | Authentic engagement cannot be generated |
| Breakout | Little direct effect | Getting real people and outlets to carry the narrative |
Public evidence so far supports this split. In May 2024 OpenAI reported disrupting five covert influence operations that had used its models: two Russian (one it named Bad Grammar, plus the long-running Doppelganger), the Chinese network known as Spamouflage, the Iranian International Union of Virtual Media, and a commercial Israeli operation it called Zero Zeno. The models were used for writing comments and articles, translating, and in some cases debugging code for posting bots. OpenAI's assessment was that none of them gained significant authentic audience engagement from it. That is one company's view of the operations it saw, not a guarantee, but the pattern matters: generation got cheaper, and audience did not.
The risks worth planning for are therefore volume (flooding a hashtag, a comment section, a public consultation or a review site), language reach (operations that previously stayed in one language), and targeted micro-persuasion in direct conversation, which is harder to observe than public posting.
Why detecting AI-written text is the wrong target
The tempting control is a classifier that labels a post as human or AI and removes the AI ones. It fails on three counts.
- Accuracy. OpenAI launched an AI text classifier in January 2023 and withdrew it on 20 July 2023, citing its low rate of accuracy. Short texts, which is what social posts are, give a classifier the least to work with.
- Bias. Liang and colleagues (2023) found that several GPT detectors flagged essays by non-native English writers as AI-generated far more often than essays by native speakers. Enforcement driven by such a detector falls hardest on people writing in a second language.
- Irrelevance. Authorship is not the harm. Plenty of legitimate users draft posts with an assistant, and plenty of operations are entirely human-written. A perfect detector would still mostly flag honest people.
Platform policies on coordinated inauthentic behaviour reflect this: they target deception about who is speaking and how many people are really speaking, regardless of what tool wrote the words. Model-generated text is at most a weak feature, and an occasional accidental one: investigators have repeatedly found operation accounts posting a model's refusal message or meta-commentary verbatim, because nobody read the output before the bot posted it. That artefact is worth matching, but it is a lucky break, not a strategy.
Detecting coordination: the pipeline
The signal that survives cheap text is coordination: many accounts behaving as one. The diagram above shows the shape of a detection pipeline built on it. Four families of features feed an account graph:
- Content similarity: posts from different accounts that say the same thing. Exact and near-duplicate matching catches lazy operations; embedding similarity and narrative clustering catch paraphrased ones.
- Temporal synchrony: accounts that post about the same thing within minutes of each other, repeatedly. One coincidence means nothing; the tenth one does.
- Account and infrastructure signals: creation dates in bursts, shared devices or network ranges, identical client fingerprints, the same profile image edited slightly. These stay internal to the platform, and their use is constrained by privacy law and policy.
- Outside intelligence: indicators shared by model providers, other platforms and researchers. Each party sees one slice of an operation that spans several services.
Edges between accounts carry weights from each family; clusters of strongly connected accounts become cases for human analysts. The analyst, not the model, decides whether the cluster is deceptive, because the same graph shape also describes a union's members sharing a press release.
Worked example: linking accounts by similar text and timing
Here is a minimal, runnable version of the first two feature families. It links two accounts when they post similar text within a time window, counts how often that happens, and groups accounts that are linked repeatedly. The similarity function is word-shingle Jaccard to keep it dependency-free; the next section explains why production systems replace it.
import re
from collections import defaultdict
def shingles(text, k=3):
w = re.findall(r"[a-z0-9']+", text.lower())
return {" ".join(w[i:i+k]) for i in range(max(1, len(w) - k + 1))}
def jaccard(a, b):
return len(a & b) / len(a | b) if a | b else 0.0
def coordinated_pairs(posts, sim=0.3, window_s=600):
"""posts: dicts with account, ts (seconds), text -> {(acct_a, acct_b): co-posting count}"""
sh = [shingles(p["text"]) for p in posts]
edges = defaultdict(int)
order = sorted(range(len(posts)), key=lambda i: posts[i]["ts"])
for x, i in enumerate(order):
for j in order[x+1:]:
if posts[j]["ts"] - posts[i]["ts"] > window_s:
break # sorted by time: stop early
a, b = posts[i]["account"], posts[j]["account"]
if a != b and jaccard(sh[i], sh[j]) >= sim:
edges[tuple(sorted((a, b)))] += 1
return edges
def clusters(edges, min_weight=2, min_size=3):
parent = {}
def find(x):
parent.setdefault(x, x)
while parent[x] != x:
parent[x] = parent[parent[x]]
x = parent[x]
return x
for (a, b), w in edges.items():
if w >= min_weight: # repeated coincidence only
parent[find(a)] = find(b)
groups = defaultdict(set)
for x in list(parent):
groups[find(x)].add(x)
return [sorted(g) for g in groups.values() if len(g) >= min_size]Worked example: eight posts about a fictional city transit levy. Accounts u1, u2 and u3 each post a reworded version of one claim within 95 seconds, and an hour later each posts a reworded version of a second claim within 100 seconds. Account u4 discusses the levy independently, hours apart, and u5 complains about a late bus in the same minute as the second burst. The output is {('u1','u2'): 2, ('u1','u3'): 2, ('u2','u3'): 2} and one cluster, ['u1', 'u2', 'u3']. u5 is not linked, despite posting at the same moment, because its text is unrelated; u4 is not linked, despite discussing the same topic, because its timing is independent. Requiring a weight of 2 means a single shared moment never forms a cluster.
At scale, a viral moment puts thousands of posts in one window, so production systems bucket by topic first, use MinHash or approximate nearest-neighbour search, and build the graph in batch over days.
When every post is a fresh paraphrase
Shingle overlap works on copy-paste operations, and many still are copy-paste. An operator who asks a model for a fresh paraphrase per account defeats it: the posts share meaning but few word triples. Three adjustments keep the approach working:
- Compare meaning, not wording. Replace Jaccard with cosine similarity between sentence embeddings, calibrated on labelled pairs from your own platform, because thresholds do not transfer between embedding models or languages. Multilingual embeddings also link the same narrative across translations.
- Cluster narratives, then accounts. Group posts into claims (topic clustering over embeddings, then human naming), and score accounts by how many distinct claims they push in synchrony. Ordinary users who agree with one claim rarely move across the same five claims on the same schedule.
- Lean on behaviour. Paraphrase changes text and leaves everything else: scheduling regularity, sleep cycles inconsistent with the claimed location, account creation bursts, reply patterns that only ever target one set of accounts. These features are the reason coordination detection still works when text is free.
Expect adaptation. Operators add random delays and mix in unrelated filler posts. That raises their cost, which is the realistic goal: detection that forces operations to behave more like real users also forces them to be slower and smaller.
Where provenance and watermarks fit
Provenance asks a different question: where did this specific piece of media come from? For images, audio and video, the C2PA standard attaches signed manifests that record which tool produced or edited the file, and some generators also embed invisible watermarks. Both are covered in Content Authentication, in depth. They help with deepfake evidence and with labelling, but they work in one direction: a valid manifest is evidence, while a missing one proves nothing, because platforms often strip metadata and operators simply use tools that add none.
For text, provenance is weaker still. Statistical watermarks that bias token choices have been published and deployed by some providers, but they only cover that provider's outputs, they weaken under paraphrase and translation, and detection needs the provider's key or detector. Treat a positive watermark result as a lead for a cross-provider investigation, never as grounds for enforcement on its own. More useful for text is provenance of claims: link a viral claim back to the earliest accounts and sites that published it, which is the same timeline reconstruction described in AI Forensics, in depth.
Measuring whether a campaign mattered
Volume is easy to count and misleading to report. A network of a thousand accounts that only reply to each other has done almost nothing. Measure the following instead:
- Authentic engagement: replies, shares and follows from accounts outside the cluster.
- Breakout: whether the narrative crossed into other platforms, communities, or mainstream coverage. Ben Nimmo's Breakout Scale (Brookings, 2020) grades operations in six categories, from a single platform and community up to amplification by prominent people or a call for policy response, and is widely used to keep reports honest.
- Time to detection and time to action: how long the network operated before you saw it.
- Recidivism: how quickly removed operators return, and whether the same infrastructure shows up again.
Publishing takedown reports with these measures and shareable indicators helps the next platform recognise the operation faster.
Operating it without harming legitimate users
The hardest part of this work is not catching operations but avoiding harm to legitimate coordination. Fan communities, political campaigns, advocacy groups and newsrooms all post similar messages at the same time, by design and in the open. Build the process around that:
- Act on deception, not on coordination: fake identities, concealed sponsors, account networks presented as independent people. Write that into the review rubric.
- Give analysts an evidence bundle per cluster (shared texts, timing plots, infrastructure overlaps) and require two independent signal families before a network action.
- Prefer graduated responses: reduce distribution or add labels before removal, and keep an appeal path with human review. The moderation architecture article covers escalation design.
- Red-team the detector with a synthetic campaign in staging (paraphrase, jittered timing, filler) and measure recall.
- Recalibrate around elections and crises, when organic synchrony spikes and quiet-week thresholds misfire.
Failure modes
| Failure mode | What it looks like | Mitigation |
|---|---|---|
| Authorship-based enforcement | Second-language writers and assistant users flagged | Drop AI-text scores as an enforcement trigger |
| Shingle-only similarity | Paraphrased networks invisible | Embedding similarity plus behavioural features |
| Single-signal action | Activist or fan groups removed | Require two signal families and human review |
| Volume-based reporting | Small networks hyped, large impact missed | Report authentic engagement and breakout |
| Siloed data | Each platform sees a fragment | Share indicators through established industry channels |
| Static thresholds | False-positive spikes during news events | Recalibrate per event; monitor review overturn rates |
What to do next
To start, or improve, a coordination-detection capability:
- Write down your policy target as deceptive coordinated behaviour, and explicitly exclude authorship of text as an enforcement criterion.
- Run the code above over a day of posts on one topic. Inspect the clusters by hand and note which are legitimate coordination; those notes become your first labelled set.
- Swap Jaccard for multilingual sentence embeddings, and calibrate the threshold on pairs your analysts have labelled.
- Add two behavioural features (posting-time regularity and account creation bursts are good starts) and require agreement across signal families.
- Build the analyst evidence bundle and an appeal path before turning on any automatic action.
- Run a synthetic paraphrased campaign through staging and measure recall.
- Define your impact metrics (authentic engagement, breakout category, time to action) and report them for every takedown. Add the provenance context from LLM output provenance where your own products generate content.