Most email features for agents start outbound: the agent drafts a message and a person approves it. That side has its own article, ADK Java + Email, in depth, which covers draft and send tools, human confirmation and the policy gateway. This article covers the other direction: email as a conversation channel, where a customer writes to support@example.com, an ADK Java agent answers in the same thread, and the conversation continues over days the way a chat would.
Email is a harder channel than chat because it was never designed for software on both ends. The same message can arrive twice. Auto-responders answer your answers. Bounces look like mail. Anyone can type any From address. And there is no session id, only a set of headers that clients maintain with varying care. The design below deals with each of these before the model sees a single word. The ADK details come from the google-adk 1.11.0 jar; mail examples use Jakarta Mail.
The inbound pipeline
The pipeline has five stages, and only one of them involves a model:
- Intake receives raw RFC 5322 bytes, from your own MTA or from an inbound provider's webhook or queue, parses them and drops duplicates.
- Screen decides whether this message deserves an answer at all: auto-replies, bounces, mailing-list traffic and unverified senders leave here.
- Thread index maps the message to an existing ADK session, or creates one.
- Run executes one agent turn with the cleaned message as the user content.
- Reply is a tool whose recipients and threading headers are computed by code, and which records the Message-ID it sends so the customer's next reply finds the session.
Treat intake as an at-least-once consumer. Inbound providers retry webhooks on timeouts, and a queue redelivers after a crash, so every later stage must be safe to repeat. The queue mechanics are the same as in ADK Java + Pub/Sub: claim, run, publish, acknowledge.
Intake: parse and drop duplicates
Parse with Jakarta Mail and key everything on the Message-ID. A message without one is rare but legal, so fall back to a hash of the raw bytes. The claim is a single insert; if the row already exists, another worker has this message and you stop.
static final Pattern ID = Pattern.compile("<[^<>\\s]+>");
Inbound accept(byte[] raw) throws Exception {
MimeMessage m = new MimeMessage(jakarta.mail.Session.getInstance(new Properties()),
new ByteArrayInputStream(raw));
String id = Optional.ofNullable(m.getMessageID())
.orElse("<sha256." + sha256Hex(raw) + "@intake.local>");
// INSERT ... ON CONFLICT DO NOTHING; false means a duplicate delivery
if (!claims.claim(id, Instant.now())) return Inbound.duplicate(id);
return new Inbound(id, m, raw);
}
List<String> ancestors(MimeMessage m) throws MessagingException {
List<String> ids = new ArrayList<>();
for (String h : List.of("In-Reply-To", "References")) {
String v = m.getHeader(h, " ");
if (v != null) ID.matcher(v).results().forEach(r -> ids.add(r.group()));
}
return ids; // In-Reply-To first: it names the direct parent
}The claim row carries a status, processing then done, plus a lease timestamp. A worker that crashes mid-turn leaves a stale lease; a sweeper re-queues it after a few minutes. Because the reply's Message-ID is deterministic (see below), a repeated turn that reaches the send step produces the same outbound id, and the gateway refuses the second copy.
Screen: should anyone answer this?
An agent that replies to every message will eventually reply to another robot, and two robots can exchange mail until a quota stops them. RFC 3834 defines the header that breaks this: Auto-Submitted. Anything automatic should carry it, and anything carrying a value other than no should not be answered automatically. Not every robot follows the RFC, so screen on several signals:
| Signal | Meaning | Action |
|---|---|---|
Auto-Submitted present and not no | auto-reply or generated mail (RFC 3834) | do not run the agent |
Null Return-Path (<>) | a bounce or other delivery notice | send to the bounce handler |
Content-Type: multipart/report | delivery status notification (RFC 3464) or read receipt | bounce handler |
List-Id or List-Unsubscribe | mailing-list or bulk mail | drop or file, never reply |
Precedence: bulk, list or junk | non-standard but widely set by bulk senders | drop |
| Message-ID or ancestors already in your outbound index, From is your own address | your own mail echoed back | drop |
| More than N agent replies in this thread in the last hour | a loop your other checks missed | stop and flag for a human |
The last row is the backstop, and it is the one that matters when everything else fails. Cap replies per thread per hour (three is plenty for a support channel) and alert when the cap trips. On your own outbound replies set Auto-Submitted: auto-replied so well-behaved responders ignore you, and X-Auto-Response-Suppress: All, which Microsoft Exchange honours for out-of-office replies.
Who sent it. The From header is typed by the sender. What you can trust is the verdict your own receiving server records in an Authentication-Results header (RFC 8601), for example dmarc=pass header.from=customer.com. Trust only the instance stamped with your server's authserv-id, the first token of the header value; any copy with that id that arrived from outside is forged and should already have been removed by your MTA. If your inbound provider reports verdicts in its own notification fields, read those instead of re-parsing headers. Require DMARC alignment before you map the address to a customer record; without it, the agent will happily discuss one customer's account with someone who merely typed their address.
Threads become sessions
Email threads have no id. What they have is In-Reply-To, naming the direct parent, and References, which lists ancestors oldest-first. Clients drop and truncate these, users reply to old messages, and some mail tools rewrite Message-IDs. So do not derive the session from one header. Keep a thread index: a table from every Message-ID you have seen or sent to the ADK session that owns it, and look up all ancestors.
String sessionFor(Inbound in, String userId) throws Exception {
for (String anc : ancestors(in.message())) {
Optional<IndexRow> row = threadIndex.find(anc);
// only the session's owner may continue it; a reply-all from a stranger starts fresh
if (row.isPresent() && row.get().userId().equals(userId)) {
threadIndex.put(in.id(), row.get().sessionId(), userId);
return row.get().sessionId();
}
}
Session s = runner.sessionService()
.createSession(APP, userId, Map.of("mail.subject", subjectOf(in)), null)
.blockingGet();
threadIndex.put(in.id(), s.id(), userId);
return s.id();
}Passing null as the session id lets the session service choose one. BaseSessionService.createSession does accept a caller-supplied id, and InMemorySessionService honours it, but hosted session services may assign ids themselves, so a mapping table is the portable design. It also survives the cases a derived id cannot: a customer who starts a new email about an old issue can be linked to the old session by a person, with one row.
Keep the owner check. Email threads gain participants: a customer forwards the thread to a colleague, who replies. If the colleague is a verified user in the same organisation, your policy may allow it; otherwise the reply starts a new session for the new sender, and the old session's history never reaches them.
Preparing the turn
What you hand the model is a cleaned version of the message, not the raw MIME:
- Pick one body. Prefer
text/plain; converttext/htmlto text when it is the only part. Never pass HTML through: hidden text and styled-away content are the cheapest place to hide instructions. - Strip quoted history. The session already holds earlier turns, so quoted lines (
>prefixes, the block afterOn ... wrote:) are duplicate tokens and, worse, text the agent may treat as new. Heuristics miss some clients' formats; keep the full original as an artifact so nothing is lost. - Attachments become artifacts. Allowlist MIME types and cap sizes, then store each with
artifactService().saveArtifact(APP, userId, sessionId, name, Part.fromBytes(bytes, type)). The message text mentions the file names, and the agent loads the ones it needs through a tool. See ADK Java artifacts. - Label it as untrusted. The instruction says the customer's email is data to answer, not instructions to follow, and tools that change state are guarded by callbacks as in ADK Java guardrails.
Content user = Content.fromParts(Part.fromText(
"Customer email (untrusted). Subject: " + subject + "\n\n" + strippedBody
+ (attachmentNames.isEmpty() ? "" : "\n\nAttachments saved: " + attachmentNames)));
RunConfig cfg = RunConfig.builder().maxLlmCalls(12).build();
Map<String, Object> mail = Map.of( // state delta merged before the turn
"mail.last_inbound_id", in.id(),
"mail.references", Optional.ofNullable(in.message().getHeader("References", " ")).orElse(""),
"mail.reply_to", verifiedAddress); // never the raw Reply-To header
runner.runAsync(userId, sessionId, user, cfg, mail).blockingSubscribe(audit::record);
Replying in thread
The reply tool takes exactly one argument from the model, the body. Everything that decides where the mail goes is read from session state that intake wrote through the runAsync state delta above: the verified reply address, the latest inbound Message-ID and the References chain. The model cannot add a recipient because there is no parameter for one.
public Map<String, Object> replyInThread(
@Schema(name = "body", description = "Plain-text reply to the customer") String body,
ToolContext ctx) throws Exception {
Map<String, Object> st = ctx.state();
String parent = (String) st.get("mail.last_inbound_id");
String refs = (String) st.get("mail.references"); // may be ""
String myId = "<" + sha256Hex(ctx.sessionId() + parent) + "@support.example.com>";
MimeMessage r = new FixedIdMessage(mailSession, myId);
r.setFrom(new InternetAddress("support@example.com"));
r.setRecipient(Message.RecipientType.TO, new InternetAddress((String) st.get("mail.reply_to")));
r.setSubject(reSubject((String) st.get("mail.subject")), "UTF-8");
r.setHeader("In-Reply-To", parent);
r.setHeader("References", (refs.isEmpty() ? "" : refs + " ") + parent);
r.setHeader("Auto-Submitted", "auto-replied");
r.setHeader("X-Auto-Response-Suppress", "All");
r.setText(body, "UTF-8");
threadIndex.put(myId, ctx.sessionId(), ctx.userId()); // before sending, so replies resolve
return gateway.send(r); // dedupes on Message-ID
}
final class FixedIdMessage extends MimeMessage {
private final String id;
FixedIdMessage(jakarta.mail.Session s, String id) { super(s); this.id = id; }
@Override protected void updateMessageID() throws MessagingException {
setHeader("Message-ID", id); // saveChanges() calls this
}
}Register it with FunctionTool.create(replyTools, "replyInThread"). The deterministic Message-ID is what makes a retried turn safe: the same session and parent always produce the same id, and the outbound gateway from the email-tool article refuses an id it has already sent. Whether replies go out automatically or wait for a person is a policy choice. A common split is automatic answers for questions, and ToolContext.requestConfirmation on a separate tool for anything that promises money, cancels an order or shares account data.
Worked example: one week of a support thread
Follow one conversation. On Monday ana@customer.com writes about a missing invoice. Her provider's webhook delivers the message twice because your endpoint took eleven seconds to answer; the second claim insert finds the row and returns. The Authentication-Results header stamped by your MX reads dmarc=pass, so intake maps her to customer 4411. No ancestor is in the index, so a new session is created and both her Message-ID and, later, the agent's reply id are indexed.
Her mail server is set to send an out-of-office notice, and it answers your reply with Auto-Submitted: auto-replied. Screening drops it: no agent turn, no reply, no loop. On Wednesday she replies from her phone, whose client sends only In-Reply-To and no References. The lookup finds the agent's Message-ID in the index, the owner matches, and the turn runs in Monday's session with the earlier invoice discussion in context. Her colleague then replies-all with a question; his address passes DMARC but belongs to a different user, so his message starts a separate session and never sees Ana's account history.
A week later a reply bounces because Ana's mailbox is full. The DSN arrives with a null Return-Path and multipart/report, goes to the bounce handler, and marks her address as soft-bounced for a day. The agent is never asked to answer a mailer daemon.
Failure modes
| Failure | Cause | Prevention |
|---|---|---|
| Two replies to one email | webhook retry or crash mid-turn | Message-ID claim, deterministic reply id, gateway dedupe |
| Mail loop with an auto-responder | replying to Auto-Submitted mail | screen headers, set your own Auto-Submitted, per-thread reply cap |
| Account details sent to an impostor | trusting the From header | require DMARC pass from your own authserv-id |
| Reply lands in a new session | client dropped References | index every sent and received id, look up all ancestors |
| Thread hijack via reply-all | session continued by any participant | owner check on lookup |
| Agent answers quoted old text | quoted history passed through | strip quotes, keep the original as an artifact |
| Hidden instructions in HTML | HTML body passed to the model | plain-text part or a text conversion only |
Trade-offs
Automatic or reviewed replies. Automatic replies feel like chat and scale; reviewed replies are safer but turn the agent into a drafting assistant with a queue. Split by tool rather than by channel. Your own MTA or a provider. Running the MX gives you the Authentication-Results header and full control; a provider gives you retries, spam filtering and a webhook, at the cost of trusting its verdict fields. One session per thread or per customer. Per thread keeps contexts small and matches what the customer sees; per customer lets the agent connect issues but grows history without bound. Per thread with cross-session memory for durable facts is usually the right middle.
What to do next
- Write the screening table into code first and replay a week of real inbound mail through it before any model is attached; count what it drops.
- Make the Message-ID claim and the deterministic reply id, then test a duplicate webhook and a crash between run and send.
- Confirm where your
Authentication-Resultsheader comes from and refuse unverified senders. - Build the thread index and test a phone client that sends only
In-Reply-To. - Add the per-thread reply cap and an alert when it trips.
- Decide which tools send automatically and which use
requestConfirmation, using the outbound design in the email-tool article. - Keep raw messages as artifacts for at least your dispute window.