ADK Java hands Gemini a list of Content objects, and each Content is a list of Parts. Getting a photo, a screen recording or a voice note into that list is the easy half. The site's ADK Java multimodal article covers it: inline bytes versus file references, upload validation, keeping bytes out of session history. This article covers the other half, which is what Gemini does with media once it arrives and how you control that from Java. That means how each modality is turned into tokens, the knobs that change the count (media resolution, video clipping and frame rate), how to read per-modality usage back from ADK events, how to get structured answers out of images and video, and the failures that only show up in production.
API names below were checked on 2026-10-03 against the java-genai javadoc (Part, VideoMetadata, GenerateContentConfig, usage metadata), the ADK Java source for LlmAgent and Event, and Google's Gemini API pages on image, video and file handling. Token figures are the documented planning numbers at that date. They differ between model generations, so treat them as estimates and confirm with token counting, as shown later.
How Gemini sees a multimodal request
Gemini models accept text, images, audio, video and documents in the same request, interleaved in any order, and return text. Some models can also return images or audio. In the Java types a Part carries exactly one kind of payload: text(), inlineData() for bytes with a MIME type, fileData() for a URI, or a function call or response. A Part can also carry videoMetadata() and a mediaResolution() setting. The factory methods you will use most are Part.fromText(String), Part.fromBytes(byte[], String mimeType) and Part.fromUri(String uri, String mimeType).
Three facts about this model shape the rest of the article. First, every modality ends up as tokens in the same context window, and you pay for them as prompt tokens, so a single video can cost more than a long conversation. Second, order is meaningful. Google's prompting guidance suggests putting a single image before the text that asks about it, and when several media parts appear, labelling each with a short text part ("Screenshot 1:") lets the instruction refer to them unambiguously. Third, ADK does not touch the media. LlmAgent forwards your parts and its GenerateContentConfig to the model, so the controls below are Gemini controls set through ADK.
How each modality becomes tokens
Plan budgets per modality, not per request. The figures below come from Google's Gemini API documentation as checked on 2026-10-03:
| Modality | Documented cost | Main lever |
|---|---|---|
| Image | 258 tokens if both sides are at most 384 px; larger images are cut into 768×768 tiles of 258 tokens each (this is the rule for the models the page describes; newer models document per-resolution caps instead) | Resize before sending; media resolution |
| Video | Sampled at 1 frame per second by default; about 100 tokens per second at default or low media resolution, about 300 at high | Clip with start and end offsets; lower fps; media resolution |
| Audio | 32 tokens per second | Trim silence; send only the relevant segment |
| Pages are processed visually as well as by extracted text, so cost grows with page count | Send only relevant pages |
Use the tiling rule on a common case. A 1920×1080 screenshot is larger than 384 px on both sides. The documented crop unit is about the smaller side divided by 1.5, which is 720 px. That gives 3 tiles across and 2 down, 6 tiles in all, or roughly 1,548 tokens. The same screenshot downscaled to 1280×720 gives a 480 px unit and 3×2 tiles again, so resizing does not always help. Resizing to fit within 384×384 gets you to 258 tokens, but you lose the small text the model needed to read. That is why media resolution exists: you set a cap on tokens per image or frame, and the model's own tokeniser trades detail for tokens within it, instead of you guessing with an image library.
Video adds up quickly. Four minutes at default resolution is about 24,000 video tokens. If you count the audio track separately at 32 tokens per second, add another 7,680. Thirty such uploads a minute is a large bill. The documented ceilings are generous: a model with a 1M-token context can take about three hours of video at low resolution, or about an hour at high. Usually the binding constraint is cost, not capacity.
The controls: media resolution, clipping and frame rate
Three controls change what the model sees. Set them on the agent, or on a single part.
Media resolution, per agent. GenerateContentConfig.Builder has mediaResolution(MediaResolution.Known) as well as overloads that take a MediaResolution or a string. ADK's LlmAgent.Builder.generateContentConfig passes it on with every model call the agent makes:
LlmAgent triage = LlmAgent.builder()
.name("bug_triage")
.model("gemini-2.5-flash")
.instruction("You triage bug reports from a screenshot and a screen recording. "
+ "Describe only what is visible. If text is unreadable, say so.")
.generateContentConfig(GenerateContentConfig.builder()
.mediaResolution(MediaResolution.Known.MEDIA_RESOLUTION_LOW)
.temperature(0.2f)
.build())
.outputSchema(TRIAGE_SCHEMA) // structured answer, see below
.build();Clipping and frame rate, per video part. VideoMetadata has startOffset(Duration), endOffset(Duration) and fps(Double). The javadoc gives the valid fps range as greater than 0 and up to 24. A lower frame rate suits slow screen recordings. A higher one suits fast motion, where 1 fps misses events.
Part clip = Part.builder()
.fileData(FileData.builder()
.fileUri(videoUri) // Files API or gs:// URI
.mimeType("video/mp4")
.build())
.videoMetadata(VideoMetadata.builder()
.startOffset(Duration.ofSeconds(150))
.endOffset(Duration.ofSeconds(210))
.fps(0.5)
.build())
.build();
Content userMsg = Content.fromParts(
Part.fromText("Screenshot 1:"),
Part.fromBytes(screenshotPng, "image/png"),
Part.fromText("Recording, 2:30 to 3:30, where the user reports the freeze:"),
clip,
Part.fromText(userReport));Per-part resolution. Part also exposes a mediaResolution field. Newer models document per-part control, so a dense screenshot can be sent at high resolution while a decorative photo goes at low. Check that your model supports it before relying on it. A setting the model ignores fails silently: the request still succeeds and the bill does not change.
Large or reused media should go by reference, not as base64 in every request. On the Gemini Developer API that means the Files API (client.files.upload(path, UploadFileConfig) in java-genai, stored for 48 hours, up to 2 GB per file and 20 GB per project as documented). On Vertex AI it usually means a gs:// URI. The Gemini pages currently give two different inline thresholds: the image page still shows the older 20 MB figure, while the Files API page says 100 MB (50 MB for PDFs). Make the threshold a configuration value, as the multimodal article recommends. Uploaded videos are processed before they can be used, so wait for the file to report an active state before you reference it.
Worked example: a bug-triage endpoint
Here is a bug-triage endpoint. A user submits a 1920×1080 screenshot, a four-minute screen recording and a paragraph of text. The service wants a JSON verdict: component, severity, reproduction steps and whether the screenshot shows an error dialog.
Naive version. The whole recording at default resolution is about 24,000 video tokens plus audio, and the screenshot about 1,500. That is over 30,000 prompt tokens per report before the instruction and history.
Shaped version. The client already records when the user pressed ‘report’, so the service clips to the minute before that moment. A screen recording changes slowly, so it samples at 0.5 fps, and it sets low media resolution on the agent. The planning figures put this at roughly 3,000 video tokens and 1,920 audio tokens, plus the screenshot. That is about a sixth of the naive cost. Preflight counting then confirms the real number for the model in use:
CountTokensResponse est = genai.models.countTokens(
"gemini-2.5-flash", List.of(userMsg), null);
int prompt = est.totalTokens().orElse(0);
if (prompt > budget.maxPromptTokens()) {
throw new PayloadTooExpensive(prompt); // ask the user to trim, or clip further
}
Flowable<Event> events = runner.runAsync(userId, sessionId, userMsg, RunConfig.builder().build());
events.blockingForEach(ev -> {
ev.usageMetadata().ifPresent(u ->
u.promptTokensDetails().ifPresent(details -> details.forEach(m ->
metrics.record("prompt_tokens",
m.modality().map(Object::toString).orElse("UNKNOWN"),
m.tokenCount().orElse(0)))));
if (ev.finalResponse()) {
ev.content().ifPresent(c -> store.saveVerdict(reportId, c));
}
});Two things make this work in production. Event.usageMetadata() returns the response's usage metadata. Its promptTokensDetails() is a list of per-modality counts, so dashboards can show image, video, audio and text tokens separately and catch the day a client starts sending full-length recordings. And the preflight count runs outside ADK. Your service can reject an expensive payload without creating a session event. Measure the real counts for your model and keep them in your test fixtures. Don't keep the planning figures. The per-second and per-tile numbers have changed between model generations, and a budget built on stale numbers fails without any error.
Structured answers and output modalities
Free-text descriptions of an image are hard to act on. For extraction tasks, give the agent an output schema. LlmAgent.Builder.outputSchema(Schema) exists in ADK Java. Keep such an agent tool-free and use it as a focused extraction step, which also makes its behaviour easy to test:
static final Schema TRIAGE_SCHEMA = Schema.builder()
.type("OBJECT")
.properties(Map.of(
"component", Schema.builder().type("STRING").build(),
"severity", Schema.builder().type("STRING")
.enum_(List.of("low", "medium", "high", "critical")).build(),
"errorDialogVisible", Schema.builder().type("BOOLEAN").build(),
"reproSteps", Schema.builder().type("ARRAY")
.items(Schema.builder().type("STRING").build()).build(),
"unreadable", Schema.builder().type("BOOLEAN").build()))
.required(List.of("component", "severity", "errorDialogVisible", "unreadable"))
.build();The unreadable field matters more than it looks. Vision models describe what they expect to see when the pixels are ambiguous. Low media resolution makes small text ambiguous, so the model reads a blurred error code as a plausible but invented one. An explicit escape hatch, plus an instruction to use it, turns a hallucinated error code into a flag your service can act on, for example by retrying at high resolution.
Output modalities go the other way. Models that support image output take responseModalities("TEXT", "IMAGE") in the same config, and the image comes back as an inlineData part in the event content. Save those bytes as an artifact (see ADK Java artifacts) rather than leaving them in session history, where every later turn would resend them. Only some models support image or audio output, so check the model page before you enable it.
Failure modes
- Silent cost blow-ups. A client update starts sending full-resolution, full-length video and nothing errors. Alert on video tokens per request from
promptTokensDetails, not only on total spend. - History amplification. Media in an early turn is resent on every later turn of the session. Store media as artifacts and reference them, or use a fresh session per extraction.
- Expired references. Files API uploads expire after 48 hours. A session resumed two days later fails on the old URI. Keep your own copy and re-upload when a reference is rejected.
- Using a file before processing ends. Uploaded videos are processed before they become usable. Referencing one too early fails the request, so poll the file state first.
- Wrong MIME type. A HEIC photo labelled
image/jpegfails or is misread. Detect the type from the bytes, not the file name. The documented image types are PNG, JPEG, WEBP, HEIC and HEIF. - Hallucinated reading. Low resolution plus small text produces confident wrong values. Give the schema an ‘unreadable’ flag, and re-run critical reads at high resolution.
- Instructions inside media. Text in an image or spoken in a recording is untrusted input, the same as a web page. An extraction agent with no tools limits the damage.
Trade-offs
Low resolution and sparse frames are cheap and fast, but they miss small text and quick events. High resolution reads fine detail, but it multiplies tokens and latency. Clipping saves the most, but it depends on knowing where to look, which usually means instrumenting the client to record timestamps. Inline bytes keep the request self-contained and stateless. File references avoid resending large payloads, but they add an expiry and a lifecycle to manage. A single multimodal agent is simple. A two-stage design, where a cheap low-resolution pass decides whether a high-resolution pass is needed, often costs less overall, at the price of an extra round trip. Measure both on your own traffic. For counting across providers and models, see token counting across models. For the client and backend setup underneath all this, see ADK Java and Gemini integration.
What to do next
- Log
promptTokensDetailsper modality fromEvent.usageMetadata()today, before changing anything, to get a baseline. - Add a preflight
countTokenscheck with a per-endpoint budget, and reject or trim payloads that exceed it. - Set
mediaResolutionon each media-heavy agent deliberately, and record the choice and the reason in code. - Clip videos with
VideoMetadataoffsets and pick fps per content type: low for screen recordings, higher for motion. - Give extraction agents an
outputSchemawith an explicit ‘unreadable’ or ‘not visible’ field. - Move media out of session history into artifacts, and handle Files API expiry with re-upload.
- Write regression fixtures with real token counts for your model, and re-measure when you change model versions.