Model Garden is Google Cloud's catalog of foundation models: Google's own models such as Gemini and Gemma, models from partners, and open models you can deploy yourself. Until April 2026 it was Vertex AI Model Garden. Google then folded Vertex AI into Gemini Enterprise Agent Platform, and the catalog is now simply Model Garden within that platform. The client libraries still describe the backend as "previously Vertex AI", and the Python module used below is still vertexai.model_garden, so you will see both names in code and documentation for some time.

The catalog itself is the easy part. What matters in production is that a model reaches your application by one of three very different paths, with different billing, latency control, data handling and failure modes. This article explains those paths from first principles, deploys an open model with the SDK and the CLI, works through a cost crossover, and covers governance, operations and the mistakes that cost money. Model lists, machine types and prices change monthly, so this article avoids them on purpose; the SDK calls shown are how you read the current values.

Three ways a model reaches you

Model Garden is a catalog; the model reaches your app by one of three pathsModel Garden catalogmodel cards, licences, deploy optionsGoogle modelsmanaged API, per tokenPartner and managed openMaaS: serverless, per tokenSelf-deployed open modelyour endpoint, per node-hourEndpoint in your projectGPU or TPU nodes, replicasdeploy()Your applicationSDK or REST, service accountgeneratepublisher APIpredictGovernance across all pathsIAM, org policy vertexai.allowedModels, VPC Service Controls, quotas
Figure 1. One catalog, three consumption paths. Only the self-deployed path creates infrastructure in your project; the governance layer applies to all three.

Every model card in the catalog leads to one or more of these paths.

Google models (managed API)Partner and managed open (MaaS)Self-deployed open model
Who runs servingGoogleGoogle, for the partner or open modelYou, on an endpoint in your project
Billing shapePer token or requestPer token, partner terms may applyPer node-hour while deployed, idle or not
SetupEnable the APIEnable the model from its card, accept termsPick hardware, deploy, wait for the model to load
Latency and throughput controlQuotas, provisioned throughput optionsQuotasFull: machine type, replicas, serving container
CustomisationTuning where offeredLimited to what the provider exposesAnything: your weights, your container, fine-tunes
Capacity riskShared quotaShared quotaAccelerator availability in your region

The first two paths are serverless: you call an API and pay for what you use. The third is infrastructure: Model Garden helps you create a prediction endpoint with a prebuilt serving container, then the endpoint is yours, with its own scaling, monitoring and bill. Most teams use more than one path, for example a managed Gemini model for general reasoning and a self-deployed small open model for a high-volume classification task. The platform architecture article places endpoints in the wider stack, and the Gemini API deep dive covers the first path in detail.

Choosing a path

Before reaching for hardware, walk the paths in order and stop at the first that fits. It keeps the default cheap and pushes infrastructure decisions to the cases that need them.

  1. Does a managed Google model meet the quality bar on your evaluation set? If yes, use it; there is nothing to operate. Evaluate on your own examples, not on public leaderboards.
  2. Is the model you need offered as a service? Partner models and many popular open models are available per token. Check the card for regions, quotas and the provider's data terms, then use it.
  3. Do you need your own weights, an unlisted model, or guaranteed capacity? Only then deploy an endpoint, and start from a configuration that list_deploy_options lists.
  4. Do you need custom serving behaviour, such as a specific inference server version, adapters or batching settings? Replace the serving container, or export the weights and serve them on your own cluster, accepting that you now own upgrades and security patches.

Record which step each model stopped at and why. When prices, quotas or the catalog change, that record tells you which decisions to revisit, and it is the evidence an architecture review will ask for.

Deploying an open model

The Python SDK's vertexai.model_garden module wraps the deploy path. Model names use the form publisher/model@version, or the Hugging Face organization/model form for models pulled from Hugging Face. Start by listing what you may deploy, then ask the model for its verified deployment configurations rather than guessing hardware:

import vertexai
from vertexai import model_garden

vertexai.init(project="my-project", location="us-central1")

# Discover deployable models; model_filter narrows the list by name.
for name in model_garden.list_deployable_models(model_filter="gemma"):
    print(name)

model = model_garden.OpenModel("google/MODEL@VERSION")   # copy the exact id from the model card

# The source of truth for machine type, accelerator and container per model.
print(model.list_deploy_options(concise=True))

Then deploy with one of the listed options. The arguments below are taken from the SDK's deploy signature; machine and accelerator values must come from your list_deploy_options output.

endpoint = model.deploy(
    accept_eula=True,                      # records licence acceptance for this project
    machine_type=MACHINE_TYPE,             # from list_deploy_options
    accelerator_type=ACCELERATOR_TYPE,
    accelerator_count=ACCELERATOR_COUNT,
    min_replica_count=1,
    max_replica_count=2,
    endpoint_display_name="summariser-prod",
    use_dedicated_endpoint=True,           # own DNS name instead of the shared one
    deploy_request_timeout=3600,           # large models take a long time to load
)

# Request and response shape depend on the serving container; copy it from the model card.
response = endpoint.predict(instances=[{"prompt": "Summarise: ...", "max_tokens": 256}])
print(response.predictions[0])

The same deploy is available in the CLI, which is easier to script in CI. gcloud ai model-garden models list lists models, gcloud ai model-garden models list-deployment-config shows the supported configurations, and deploy takes the same choices as flags:

gcloud ai model-garden models deploy \
  --model=google/MODEL@VERSION \
  --region=us-central1 \
  --machine-type=MACHINE_TYPE \
  --accelerator-type=ACCELERATOR_TYPE \
  --accelerator-count=1 \
  --endpoint-display-name=summariser-prod \
  --accept-eula

Other deploy arguments worth knowing: spot=True for Spot capacity on interruption-tolerant workloads, reservation affinity to land on capacity you have reserved, hugging_face_access_token for gated Hugging Face models, enable_private_service_connect with a project allow list for private access, and the serving_container_* family for replacing the prebuilt container with your own. OpenModel.export() copies weights to a Cloud Storage path when you would rather serve them on your own GKE cluster, a route described in GKE for ML.

Worked example: the cost crossover

The central decision is serverless per-token pricing against a deployed endpoint billed by the hour. The arithmetic is simple, and doing it explicitly prevents both overspending and premature optimisation. The numbers below are invented for illustration; substitute your own quotes.

Suppose a support-ticket summariser handles 40 million input-plus-output tokens a day. A managed model costs an illustrative 0.50 dollars per million tokens blended, so 20 dollars a day. A self-deployed open model on one GPU node costs an illustrative 4 dollars an hour, and for high availability you keep two replicas: 192 dollars a day, whether or not tickets arrive. The break-even volume is node_cost_per_day / price_per_token: 192 divided by 0.50 per million is 384 million tokens a day. Below that, the managed path is cheaper. Above it, the endpoint wins only if the two replicas can actually serve that many tokens; measure throughput under your prompt lengths before believing the spreadsheet.

Throughput is the number people skip. Measure it against the deployed endpoint with prompts drawn from real traffic, at increasing concurrency, and record tokens per second per replica at the latency you can accept. A sketch:

import time
from concurrent.futures import ThreadPoolExecutor

def one_call(prompt):
    t0 = time.perf_counter()
    endpoint.predict(instances=[{"prompt": prompt, "max_tokens": 256}])
    return time.perf_counter() - t0

for concurrency in (1, 4, 16, 32):
    batch = sample_prompts[: concurrency * 10]          # real prompts, realistic lengths
    start = time.perf_counter()
    with ThreadPoolExecutor(max_workers=concurrency) as pool:
        latencies = sorted(pool.map(one_call, batch))
    wall = time.perf_counter() - start
    p95 = latencies[int(len(latencies) * 0.95) - 1]
    print(f"c={concurrency} req/s={len(batch) / wall:.1f} p95={p95:.2f}s")

Multiply the highest request rate that still meets your p95 target by your average tokens per request; that is the capacity side of the break-even. If the volume you need exceeds what two replicas sustain, add replicas to the cost column and redo the division.

Cost is not the only axis. Self-deployment is chosen for reasons the arithmetic ignores: weights you fine-tuned, a model not offered as a service, strict control of latency, or data handling requirements. And the managed path is chosen for reasons the endpoint ignores: zero idle cost, no capacity planning, and no on-call rotation for a GPU fleet. Write both columns down.

Governance

Because the catalog makes hundreds of models one click away, governance has to be explicit.

  • IAM. Calling and deploying models are separate permissions. Give applications a service account that can call specific endpoints, and keep deploy rights for a pipeline identity. See Google Cloud IAM.
  • Organisation policy. The vertexai.allowedModels constraint controls which Model Garden models users can access, set at organisation, folder or project level. Per the documentation, each model must be listed individually, an explicit deny takes precedence over an allow, a policy holds at most 500 allowed and denied values, and the constraint applies only to Model Garden models, not to models you register yourself in the model registry. Copy the value format from the current documentation.
  • Licences. Open models carry their own licences, and accept_eula records acceptance for the project. Have someone who reads licences approve each model before it enters your allow list.
  • Network perimeter. Put the projects that hold endpoints and sensitive data inside VPC Service Controls, and use private access for endpoints that should not be reachable from the internet.
  • Data path. For self-deployed models, prompts are processed by containers in your project. For partner models, read the partner's data terms for the specific offering; they differ from Google's own.

Operating self-deployed endpoints

A self-deployed endpoint is a small serving fleet, and it needs the same operational care.

  • Accelerator quota first. Deploys fail when the region lacks quota or capacity for the chosen accelerator. Request quota before the launch date, and keep a second listed configuration or region as a fallback.
  • Load time. Large models take many minutes to download and load. Set a deploy timeout, and remember that scaling out is just as slow, so size the minimum replicas for the peak you cannot wait for.
  • Health and saturation. Watch request latency, error rate and accelerator utilisation per replica, and alert on sustained queueing rather than on single slow requests.
  • Versioning. Deploy a new model version to the same endpoint with a traffic split, compare, then shift traffic, instead of replacing in place.
  • Clean up. Billing continues until the model is undeployed. Undeploy and delete experiment endpoints in the same script that created them, and run a weekly report of endpoints with no traffic.
from google.cloud import aiplatform

aiplatform.init(project="my-project", location="us-central1")
for ep in aiplatform.Endpoint.list(filter='display_name="summariser-exp"'):
    ep.undeploy_all()      # stops node billing
    ep.delete()            # removes the empty endpoint

Treat endpoints as code: create them from a pipeline with labels for owner and cost centre, and make the pipeline the only identity allowed to deploy. Then the idle report can name an owner for every endpoint it finds, and an endpoint without labels is itself a finding.

Failure modes

  • Forgotten endpoints. A one-click test deployment runs at full node cost for weeks.
  • Guessed hardware. A machine type that is not in the model's deploy options fails, or deploys and runs out of accelerator memory under real prompt lengths.
  • Wrong request schema. Different serving containers expect different instance formats; copying a payload from another model's card returns errors or silently ignored parameters.
  • Quota surprises. MaaS quotas are per model and region; a launch traffic spike returns rate-limit errors with no infrastructure to scale. Request increases before launch and back off with jitter.
  • Unapproved models in production. Without an allow list, a developer's experiment becomes a dependency with an unreviewed licence.
  • Naming drift. Docs, consoles and SDKs mix Vertex AI and Agent Platform names during the transition; pin SDK versions and test upgrades, rather than assuming renamed parameters behave the same.

What to do next

  1. List the models you use today and mark each with its path: managed API, MaaS or self-deployed.
  2. For each self-deployed model, run list_deploy_options() and confirm your hardware is a listed configuration.
  3. Do the break-even calculation with your real prices and measured throughput, and record the decision.
  4. Set vertexai.allowedModels at the folder or organisation level, starting from the models in step 1.
  5. Separate deploy and invoke permissions between a pipeline identity and application service accounts.
  6. Request accelerator and MaaS quota for launch peaks, and choose a fallback configuration.
  7. Add a scheduled job that reports and undeploys idle endpoints.
Key takeaway: Model Garden is a catalog; what you operate is the path a model takes to your app. Managed APIs and MaaS bill per token with no idle cost, while self-deployed endpoints bill per node-hour and give full control. Read hardware from list_deploy_options, compute the break-even with measured throughput, restrict models with vertexai.allowedModels, separate deploy from invoke permissions and undeploy idle endpoints.