"AI trading" covers three different systems that are often discussed as one: statistical models that turn market data into signals, language models that read news, filings and transcripts and turn them into features or recommendations, and agents that are allowed to call an order-entry tool. Each adds risk on top of ordinary algorithmic trading, and the last one adds the most, because a model that can be talked into something is now connected to money that moves in milliseconds.

This article is about securing that order path. It assumes the general controls for LLMs in finance, which are covered in the finance LLM threat model, and the regulatory map in AI finance regulation. Here the focus is narrower: where untrusted text enters a trading decision, how a typed order intent and a deterministic pre-trade gateway contain a model that has been fooled or has simply gone wrong, how to stop the whole thing fast, and how to test it before the market tests it for you. Manipulation techniques are discussed only at the level a defender needs to model them.

Three systems called AI trading

Start by deciding which kind of system you are actually building, because the attack surface follows the inputs and the authority.

SystemUntrusted inputAuthorityMain risk
Signal model on market dataPrices and order-book events anyone can influence by tradingFeeds a strategy that sizes ordersAdversarial or anomalous data moves the signal; overfitting; runaway feedback
LLM feature extractorNews, filings, social posts, transcriptsProduces sentiment or event featuresPrompt injection and fabricated stories become features
LLM agent with an order toolAll of the above plus chat and tool outputProposes or places ordersExcessive agency: a fooled model acts directly on the market

The rule that follows is simple and it is the backbone of everything below: the more untrusted text a component reads, the less authority it may hold. A model that reads social media may emit a score; it may not bypass, change or waive a limit.

Architecture: authority lives outside the model

The model proposes, a deterministic gateway disposesUntrusted inputsnews, filings, social, chatMarket dataquotes, trades, order bookModel / agentsignals, LLM reasoningOrder intenttyped, schema-validatedproposesPre-trade risk gatewaycollars, value, size, rate,position, credit, restricted listowned by risk, not by the modelBroker / exchangeorders that passedacceptReject + alertreason codes loggedrejectKill switchindependent of model hostFills and positions feed back into both the model and the gateway's limits
Figure 1. A trading system that uses ML or an LLM keeps authority in a deterministic risk gateway the model cannot configure. Untrusted text and market data influence proposals; only typed intents that pass every check become orders.

Figure 1 shows the shape that survives real incidents. Market data is untrusted because any participant can move it by trading; text is untrusted in a stronger sense, because anyone can publish it, it can carry instructions, and it can be fabricated outright.

The model, whether a gradient-boosted signal or an LLM agent, produces a proposal. The proposal is converted into a typed order intent with a fixed schema. The intent goes to a pre-trade risk gateway that is a separate service, written in boring deterministic code, configured by the risk function, and impossible for the model to reconfigure. Only intents that pass every check become orders. A kill switch sits outside the model's host so it works when that host is the thing misbehaving.

This is not a new idea invented for LLMs. In the United States, SEC Rule 15c3-5, the Market Access Rule, has since 2010 required broker-dealers with market access to keep risk controls reasonably designed to stop orders that exceed pre-set credit or capital thresholds and erroneous orders that break price or size parameters or duplicate earlier ones, and those controls must be under the broker-dealer's direct and exclusive control. In the EU, the RTS 6 rules on algorithmic trading under MiFID II require pre-trade controls including price collars, maximum order values, maximum order volumes and maximum message limits, plus a kill functionality. AI relaxes none of this.

Threat model

ThreatHow it reaches the order pathPrimary control
Indirect prompt injectionInstructions hidden in a press release, PDF filing or web page the LLM readsExtraction to a fixed schema; no tool access in the reading step
Fabricated or spoofed newsA fake headline or a hijacked account reports a merger or defaultSource allowlist, provenance, corroboration before size
Adversarial market activityOther participants place and cancel orders to provoke a model they have profiledSignal sanity bounds, rate and position limits, surveillance
Runaway behaviourA bug, a stale flag or a feedback loop sends orders repeatedlyMessage limits, duplicate checks, kill switch
Training data poisoningManipulated historical data or labels shift what the model learnsData provenance, holdout replay, change review
Strategy and data leakagePrompts, logs or a vendor model expose positions or material non-public informationData minimisation, private deployment, log controls

Two entries deserve emphasis. Injection matters because the LLM's job is to read exactly the documents an attacker can write; the background is in indirect prompt injection. Adversarial market activity matters because it needs no access to your systems at all: if a model's reaction to a pattern is predictable, someone can produce the pattern.

The order intent: a narrow output

The first containment layer is the shape of what the model is allowed to say. Free text is never parsed into an order. The model, or the code around it, fills in a narrow structure, and anything that does not validate is dropped and logged.

from dataclasses import dataclass
from decimal import Decimal
from enum import Enum

class Side(Enum):
    BUY = "BUY"
    SELL = "SELL"

ALLOWED_SYMBOLS = {"AAPL", "MSFT", "NVDA"}   # loaded from config owned by risk

@dataclass(frozen=True)
class OrderIntent:
    strategy_id: str
    symbol: str
    side: Side
    quantity: int
    limit_price: Decimal          # market orders are not representable
    rationale_id: str             # pointer to the stored model output, for audit

    def __post_init__(self):
        if self.symbol not in ALLOWED_SYMBOLS:
            raise ValueError(f"symbol not tradable by this strategy: {self.symbol}")
        if not 0 < self.quantity <= 10_000:
            raise ValueError("quantity outside schema bounds")
        if self.limit_price <= 0:
            raise ValueError("limit price must be positive")

Several choices here are deliberate. There is no market order type, so a model cannot ask to trade at any price. Quantities have a schema ceiling that is far above normal size but well below disaster size. The rationale is stored elsewhere and referenced by ID, so the audit trail keeps the model's reasoning without letting that text travel into the execution path. An agent's order tool should accept exactly this structure and only enqueue it, never talk to a broker.

The pre-trade risk gateway

The gateway is where authority actually lives. It holds limits, current positions and reference prices, and it evaluates every intent against all of them. It is ordinary code with ordinary tests, and that is the point: it behaves the same way whether the model upstream is brilliant, broken or compromised.

import time
from collections import deque

class RiskGateway:
    def __init__(self, limits, positions, ref_prices, restricted):
        self.limits = limits            # per strategy, owned by the risk team
        self.positions = positions      # updated from fills, not from the model
        self.ref = ref_prices           # independent market data source
        self.restricted = restricted    # restricted / watch list
        self.msgs = deque()
        self.recent = deque(maxlen=200)

    def check(self, intent):
        lim = self.limits[intent.strategy_id]
        if lim["halted"]:
            return "REJECT_STRATEGY_HALTED"
        if intent.symbol in self.restricted:
            return "REJECT_RESTRICTED"
        ref = self.ref[intent.symbol]
        if abs(intent.limit_price - ref) / ref > lim["price_collar"]:
            return "REJECT_PRICE_COLLAR"
        notional = intent.limit_price * intent.quantity
        if notional > lim["max_order_value"]:
            return "REJECT_ORDER_VALUE"
        if intent.quantity > lim["max_order_qty"]:
            return "REJECT_ORDER_QTY"
        signed = intent.quantity if intent.side.value == "BUY" else -intent.quantity
        if abs(self.positions.get(intent.symbol, 0) + signed) > lim["max_position"]:
            return "REJECT_POSITION"
        now = time.monotonic()
        while self.msgs and now - self.msgs[0] > 1.0:
            self.msgs.popleft()
        if len(self.msgs) >= lim["max_msgs_per_sec"]:
            return "REJECT_MESSAGE_RATE"
        key = (intent.symbol, intent.side, intent.quantity, intent.limit_price)
        if key in self.recent:
            return "REJECT_DUPLICATE"
        self.msgs.append(now)
        self.recent.append(key)
        return "ACCEPT"

Production gateways add credit usage, aggregate notional, venue and short-sale rules, and two-person limit changes. Three properties must not erode as it grows. Reference prices come from a feed the model does not see or control. Positions come from fills reported by the broker, not from what the model believes it did. Every reject has a reason code, so a wave of rejects is visible as a signal in itself.

Reading untrusted text safely

For the LLM that reads news, the defence is to make the reading step incapable of acting and narrow in what it outputs. A practical pipeline has four stages.

  1. Ingest with provenance. Each document carries its source, retrieval time and, where possible, a verified publisher identity. Sources are on an allowlist with a trust tier; a newswire and an anonymous post are not the same input.
  2. Extract to a fixed schema. The model returns fields such as entity, event type from a closed list, direction and confidence. It has no tools in this step, so an embedded instruction to "buy now" has nothing to call.
  3. Corroborate before size. A high-impact event from a single source can move a signal only within a small cap. Full weight needs a second independent source or a structured feed such as an exchange or regulator announcement.
  4. Quarantine the strange. Text that tries to address the model, uses hidden characters, or produces an event far outside historical frequency goes to a queue for review instead of into features.

Scanners help at the edges but are not the control; see prompt injection scanners.

Stopping fast

Stopping must be faster and more reliable than starting. Build stops at three levels: a strategy halt that the gateway enforces by flag, a firm-level stop that cancels resting orders and blocks new ones at the broker connection, and automatic triggers that fire without a human, such as reject rate, loss over a window, message rate or position drift between what the model believes and what the broker reports. The kill path must not depend on the model's host, its queue or its credentials. The general design is in agent kill switches; for trading, add one rule: after a kill, restarting needs a person who did not trigger it to confirm that positions reconcile.

Testing against history and adversaries

A trading model is tested against history and against adversaries before it touches money. Replay recorded market data and news through the full pipeline, including the gateway, and compare intents with what a previous version produced. Then run scenario tests that a backtest will never contain on its own:

  • A fabricated headline from a low-trust source, with and without a corroborating source.
  • A filing with an embedded instruction aimed at the extraction model.
  • A burst of order-book activity that mimics a pattern the signal reacts to strongly.
  • A broker disconnect while orders are resting, followed by a reconnect with partial fills.

Run these in a simulator or paper account on every model or config change, and keep the pass criteria in terms of gateway outcomes: which rejects fired, maximum position reached, and time to halt.

Worked example: the fake takeover headline

Consider an LLM agent that trades a handful of large-cap stocks around company news. At 14:02 a convincing headline appears on a social account that impersonates a wire service: a large company has received a takeover offer at a big premium. The page linked from the post also contains, in white text, an instruction telling any AI reader to buy immediately and ignore risk limits.

Walk it through the architecture. Ingest tags the source as tier three, unverified. Extraction has no tools, so the hidden instruction becomes, at most, odd text that the quarantine check flags. The extracted event, takeover offer with positive direction, is single-source and low-trust, so the feature is capped at a small fraction of its normal weight. The agent proposes a modest buy. The gateway sees that the limit price the agent chose sits well above the reference price and rejects it on the price collar; the agent tries again at a lower price, which passes, and buys a size bounded by its order value and position limits. Minutes later the story is denied; the loss is small and the reason codes explain it.

Remove the cap, the collar, the position limit and the message limit, and each absence multiplies the loss. That is why the controls are layered.

Failure modes

  • Dead code and reused flags. Knight Capital lost roughly 460 million dollars in about 45 minutes on 1 August 2012 after a deployment left old code active on one server and a repurposed flag switched it on. The SEC later charged the firm under the Market Access Rule and it paid a 12 million dollar penalty. Model and config rollouts need the same rigour.
  • Positions from belief, not fills. After a disconnect, the model thinks it holds nothing while the broker shows a large position.
  • Unbounded retries. An agent that treats a reject as a problem to solve keeps probing the limits; reject loops should halt the strategy.
  • Backtest-only validation. History contains no attacks on your model, so a clean backtest says nothing about adversarial robustness.

Trade-offs

Tight collars and small caps on single-source news cost alpha on genuine scoops; loose ones cost money on fakes. Choose per strategy and record the choice. Human approval above a size threshold adds latency that some strategies cannot afford, so those strategies should run with smaller limits rather than no approval. A separate gateway adds a network hop and a few microseconds to milliseconds of latency, which matters for high-frequency trading and almost never for news-driven strategies.

What to do next

  1. Draw your order path and mark every point where untrusted text or market data can change a decision.
  2. Make the model emit a typed, schema-validated order intent with limit prices only, and remove any direct broker access from model hosts.
  3. Stand up or audit a pre-trade gateway with collars, value, quantity, position, message-rate and duplicate checks, owned and configured by risk.
  4. Add provenance tiers, schema-only extraction and corroboration caps to the news pipeline.
  5. Wire automatic halt triggers and a kill path that works with the model host down, and rehearse it.
  6. Build a scenario suite of fake headlines, injected filings and disconnects, and run it on every model or config change.
  7. Review where prompts, traces and positions are logged, and who can read them.
Key takeaway: In AI trading the model should propose and never dispose. Keep untrusted text out of the authority path, make the model emit a narrow typed intent, enforce collars, size, position, rate and duplicate limits in a deterministic gateway owned by risk, cap single-source news, stop fast through a path that does not depend on the model, and test against fake news and broken connections, not only against history.