An agent that can only read text asks users to describe the photo of the damaged car, transcribe the voicemail and paste the invoice. Gemini models accept images, audio, video and PDFs directly, and ADK Java passes them through without a separate API: a user message is a Content made of Part objects, and a part can hold bytes or a file reference as easily as text. The hard work is everything around that one line: deciding when bytes travel inline, validating what users upload, keeping large files out of the session history, and letting tools read and produce media.

This page builds that path end to end. You will construct multimodal messages, choose between inline data and file references, validate uploads, use the runner's option that turns uploads into artifacts (and see exactly what the model then receives), write a tool that works with stored files, trace a worked insurance-claim example, and finish with the failure modes. Basic model wiring is in ADK Java + Gemini and assumed here.

A message is a list of Parts

Every message in ADK Java is a com.google.genai.types.Content: a role and an ordered list of parts. A Part holds one kind of payload. The ones that matter here are text, inlineData (a Blob with a MIME type and raw bytes) and fileData (a URI plus a MIME type). The Java GenAI SDK gives you static factories for each:

import com.google.genai.types.Content;
import com.google.genai.types.Part;
import java.nio.file.Files;
import java.nio.file.Path;

byte[] photo = Files.readAllBytes(Path.of("claim-4411/front.jpg"));

Content message = Content.fromParts(
    Part.fromText("Here is the damage to the front bumper."),
    Part.fromBytes(photo, "image/jpeg"),                       // inline bytes
    // gs:// works on Vertex AI; on the Gemini Developer API, register the
    // object with the Files API or pass a Files API or pre-signed URI.
    Part.fromUri("gs://claims-bucket/4411/policy.pdf",         // reference
                 "application/pdf"),
    Part.fromText("Is this covered, and what should the adjuster check?"));

The model reads the parts in order, so order carries meaning. Put a short label before each file ("Photo 1: front bumper") when there are several, and put the question last, after the evidence. If you need to refer to a file later in the conversation, the label is what the model will remember, not the bytes.

Two optional fields on Part are worth knowing exist: videoMetadata, used to point the model at a clip within a video, and mediaResolution, which trades detail for token cost on media. Which models honour them changes over time, so check the model documentation before relying on either.

Inline bytes or a file reference

Inline bytes are simplest: no storage, no URI, no expiry. They are also sent on every request that includes them, base64-encoded in the JSON body, which makes them roughly a third larger on the wire than on disk. File references keep the request small and let the same file be reused across turns and agents, at the cost of somewhere to put the file and credentials to read it.

SituationPreferReason
One small image, used onceInlineNo storage round trip, no lifecycle to manage
Large PDF, long audio or any videoFile referenceKeeps request size and latency down; avoids resending
Same file used across many turnsFile reference or artifactInline bytes in history are resent every turn
Files already in your cloud storageFile referenceNo download-then-forward through your service
Strict data residencyWhichever keeps bytes in the approved regionCheck where the model backend reads the file from

Size limits differ by backend, so make the threshold configuration, not a constant. In January 2026 Google raised the Gemini Developer API's maximum inline payload from 20 MB to 100 MB and added inputs from external HTTPS and pre-signed cloud storage URLs, plus registering existing Cloud Storage objects with the Files API. Vertex AI publishes its own limits, and some MIME types have their own caps. Because the limit applies to the encoded request, a 90 MB file is already over it after base64. A conservative default of a few megabytes for inline data keeps latency predictable and leaves room for history.

Validating uploads before the model sees them

An upload endpoint receives a filename, a client-declared content type and bytes, and all three are attacker-controlled. Never forward the declared MIME type to the model without checking the bytes. A mislabelled file wastes a model call at best, and at worst lets someone push a file type your downstream tools will parse unsafely. Check size first, then sniff the leading bytes against an allow-list:

import java.util.Arrays;
import java.util.Optional;

final class UploadValidator {
  private final long maxInlineBytes;   // from configuration, not a constant

  UploadValidator(long maxInlineBytes) { this.maxInlineBytes = maxInlineBytes; }

  /** Returns the MIME type detected from content, or empty if not allowed. */
  Optional<String> sniff(byte[] b) {
    if (starts(b, 0xFF, 0xD8, 0xFF)) return Optional.of("image/jpeg");
    if (starts(b, 0x89, 'P', 'N', 'G')) return Optional.of("image/png");
    if (starts(b, '%', 'P', 'D', 'F')) return Optional.of("application/pdf");
    if (b.length > 12 && starts(b, 'R', 'I', 'F', 'F')
        && b[8] == 'W' && b[9] == 'A' && b[10] == 'V' && b[11] == 'E')
      return Optional.of("audio/wav");
    return Optional.empty();
  }

  boolean fitsInline(byte[] b) {
    long encoded = (b.length + 2) / 3 * 4;   // base64 size on the wire
    return encoded <= maxInlineBytes;
  }

  private static boolean starts(byte[] b, int... sig) {
    if (b.length < sig.length) return false;
    for (int i = 0; i < sig.length; i++) if ((b[i] & 0xFF) != sig[i]) return false;
    return true;
  }
}

Use the sniffed type, not the declared one, when building the part, and reject uploads where they disagree. Strip metadata you do not need, such as EXIF location from photos, before the file leaves your service: the model does not need the claimant's GPS coordinates to assess a dent. Record the hash of what you forwarded so an investigation can later prove which bytes the model saw.

Keeping bytes out of session history

Sessions store every event, and every turn replays the history to the model. An inline image in turn one is therefore stored in the session database and, unless something removes it, resent with every later request. Ten turns with a 4 MB photo means 40 MB of the same photo uploaded.

ADK Java has an option for this. Set saveInputBlobsAsArtifacts(true) on the RunConfig and the runner, before appending the user message to the session, saves each inline part to the artifact service and replaces it in the stored event with a text part. In the current adk-java source the name is artifact_<invocationId>_<index> and the replacement reads "Uploaded file: <name>. It has been saved to the artifacts". The detail that surprises people: on that turn the model receives the placeholder text, not the image. To let it see the file, give the agent LoadArtifactsTool.INSTANCE, which tells the model which artifacts exist and returns their content when it asks for one.

import com.google.adk.agents.LlmAgent;
import com.google.adk.agents.RunConfig;
import com.google.adk.runner.InMemoryRunner;
import com.google.adk.tools.LoadArtifactsTool;

LlmAgent agent = LlmAgent.builder()
    .name("claims_assistant")
    .model("gemini-2.5-flash")
    .instruction("You assess insurance claims. Uploaded files are stored as "
        + "artifacts; load the ones you need before answering.")
    .tools(LoadArtifactsTool.INSTANCE)
    .build();

InMemoryRunner runner = new InMemoryRunner(agent);
RunConfig cfg = RunConfig.builder()
    .saveInputBlobsAsArtifacts(true)
    .build();

runner.runAsync(userId, sessionId, message, cfg)
    .blockingForEach(event -> event.content().ifPresent(System.out::println));

The trade is one extra model round trip (the model calls the load tool, then answers) in exchange for a history that stays small and files that are versioned and retrievable later. Verify the behaviour on your ADK version with a test that inspects the stored event, because this is the kind of detail that changes between releases. Artifact storage itself, including versioning and scoping, is covered in ADK Java artifacts.

Tools that read and write media

Tools often need the file the user sent, or produce one: a cropped region, a rendered chart, a transcript. ToolContext extends CallbackContext, which exposes loadArtifact(String filename) returning Maybe<Part>, saveArtifact(String filename, Part artifact) returning Completable and listArtifacts(). Scoping to the current app, user and session comes from the context, so a tool cannot read another tenant's files by guessing a name.

import com.google.adk.tools.Annotations.Schema;
import com.google.adk.tools.ToolContext;
import com.google.genai.types.Part;
import java.util.Map;

public final class ClaimTools {
  /** Checks an uploaded file before the model relies on it. */
  @Schema(description = "Report the type and size of an uploaded file.")
  public static Map<String, Object> describeUpload(
      @Schema(name = "filename") String filename, ToolContext ctx) {
    Part part = ctx.loadArtifact(filename).blockingGet();   // null if missing
    if (part == null || part.inlineData().isEmpty()) {
      return Map.of("status", "error", "reason", "no such file: " + filename);
    }
    var blob = part.inlineData().get();
    int size = blob.data().map(d -> d.length).orElse(0);
    return Map.of("status", "ok",
        "mimeType", blob.mimeType().orElse("unknown"), "bytes", size);
  }
}

Register it with FunctionTool.create(ClaimTools.class, "describeUpload"). The blocking calls are acceptable in a synchronous tool on a virtual thread; in a reactive tool, compose the Maybe instead. Return a short structured result rather than the bytes: the model works from the file it loaded, and the tool's job is facts the model cannot see, such as size, page count or a checksum match.

Worked example: an insurance claim with a photo and a policy PDF

Client uploadphoto, PDF, audioUpload validatorsize, magic bytesPart builderinline or fileUriRunner.runAsyncContent of PartsSession serviceevents, placeholdersArtifact serviceversioned bytesLlmAgent + toolsLoadArtifactsToolGemini modelreads Parts in orderObject storageFiles API / GCS / URLeventinline blob savedload on requestLlmRequestlarge: uploadSmall files may travel inline; large or reused files go by reference; session history should hold names, not bytes.
A multimodal request path: validate, choose inline or reference, run, and keep only names in session history.

Trace one claim. The claimant uploads a 3.1 MB JPEG and a 7 MB PDF policy. The validator sniffs both, strips EXIF from the photo and finds the encoded photo under the 4 MB inline threshold; the PDF exceeds it and is written to the claims bucket. The service builds a message of four parts: a label, the photo inline, the PDF as a gs:// reference (this agent runs on Vertex AI; on the Developer API it would be registered with the Files API first), and the question.

With saveInputBlobsAsArtifacts on, the runner stores the photo as an artifact and appends an event whose photo part is now placeholder text; the PDF reference, already small, is kept as-is. The model sees the label, the placeholder, the PDF and the question, calls the load tool for the photo, receives it, and answers: front bumper and grille damage, policy section 4.2 covers collision with a deductible. On turn two the adjuster asks about the headlight. History holds a file name, not 3 MB, and the model loads the photo again only if it needs it.

Audio and video: files versus live streams

Audio and video files go through the same Part path as images: an MP3 voicemail or an MP4 dashcam clip is a blob or a file reference with the right MIME type, and the model transcribes or describes it as part of the turn. Video is the case where references stop being optional, because even short clips are large.

Live conversation is different. Streaming microphone audio to the model and speaking back uses bidirectional streaming configured through RunConfig (response modalities and speech settings), not a file upload, and is covered in Gemini 2.5 features in ADK Java. Decide early which you need: a voicemail-triage agent should process files in batch, while a phone agent needs the live path and its latency budget.

Failure modes

The common ways this breaks:

  • Trusting the declared MIME type. A file labelled image/png that is really a PDF or HTML gets forwarded, rejected late or parsed by a tool. Sniff the bytes.
  • Expecting the model to see stored uploads. With saveInputBlobsAsArtifacts on and no load tool, the model answers from a placeholder and may invent what the image shows. Add the tool and test for it.
  • History bloat. Inline bytes in session events grow storage and resend cost every turn. Watch request size per turn in your metrics.
  • Over-limit requests. A file under the limit on disk exceeds it after base64, or several files together do. Measure the encoded total.
  • Expired or unreadable references. Files API uploads expire, signed URLs time out and bucket permissions change. Keep your own copy and re-upload on failure.
  • Leaking metadata. EXIF location, document authors and embedded comments reach the model and its logs. Strip what is not needed.
  • Prompt injection in media. Text in an image or a PDF is model input like any other, and can carry instructions. Treat file content as untrusted and keep tool permissions narrow.

Trade-offs

ChoiceGainCost
Inline bytesSimplest, no storageRequest size, resend on every turn, size caps
File referencesSmall requests, reuse, large filesStorage, credentials, expiry, residency questions
saveInputBlobsAsArtifacts + load toolSmall history, versioned filesOne extra model round trip per file use
Pre-processing (resize, extract text)Fewer tokens, predictable costLoses detail the model might have used

A reasonable default for most agents: validate everything, send small images inline on the first turn, send everything else by reference, and turn on artifact storage once conversations run past a few turns. Callbacks are a natural place to enforce these rules centrally; see ADK Java callbacks.

What to do next

Before you ship multimodal input:

  1. List the MIME types you accept and write the magic-byte check for each; reject anything else.
  2. Make the inline threshold configuration, measured on the base64-encoded size, and document the backend limit you set it against.
  3. Strip metadata from images and documents before they leave your service, and log a hash of what you forwarded.
  4. Decide inline versus reference per file type, and keep your own copy of every referenced file.
  5. Turn on saveInputBlobsAsArtifacts for multi-turn agents, add LoadArtifactsTool, and write a test that inspects the stored event.
  6. Give tools artifact access through ToolContext, and return facts about files, not the files themselves.
  7. Track request size and model latency per turn so history bloat shows up on a dashboard, not a bill.
  8. Treat text inside media as untrusted input in your guardrails.
Key takeaway: In ADK Java a multimodal message is a Content of ordered Parts, and a Part holds text, inline bytes or a file reference. Validate every upload by its bytes, choose inline or reference against a configurable limit measured after base64, and keep bytes out of session history with saveInputBlobsAsArtifacts plus LoadArtifactsTool, remembering that the model then sees a placeholder until it loads the file. Give tools artifact access through ToolContext and treat any text inside media as untrusted input.