BentoML is an open-source Python framework (Apache 2.0) for turning model code into production inference services. You write a Python class, mark methods as APIs, declare the resources the service needs, and BentoML gives you an HTTP server, a versioned deployable artifact called a Bento, container builds, and service-to-service composition. Its managed platform, BentoCloud, adds autoscaling and scale-to-zero. In February 2026 Modular, the company behind the MAX inference engine and the Mojo language, announced that it had acquired BentoML; the framework remains open source.

For LLM serving the most important point is what BentoML is not. It is not an inference engine. Token generation speed comes from engines such as vLLM, SGLang or TensorRT-LLM, which implement continuous batching, paged KV caches and fused kernels; the trade-offs between them are covered in LLM serving stacks compared. BentoML sits one layer up: it packages an engine with its model and environment, composes it with the CPU and GPU services around it, and scales replicas. This article explains that layer, shows verified code against the current API, works through capacity and autoscaling numbers, and lists the failure modes.

Architecture: services, APIs, models and Bentos

BentoML for LLMs: the framework packages, composes and scales; the engine owns the GPUClientOpenAI SDK, HTTPGateway ServiceCPU, validation, routingdependsEmbed / guard Serviceadaptive batchingHTTPLLM ServiceBentoML proxyvllm serve subprocessGPUweights + KV cacheModel storeHuggingFaceModel, BentoModelbentoml build: Bento (code, image, models)versioned artifactBentoCloud or your Kubernetesreplicas scaled on concurrency per replicadeployContinuous batching, paged KV cache and kernels live in the engine. BentoML adds packaging, composition and scaling.
A typical LLM application: a CPU gateway, a batched embedding or guard model, and an LLM service that runs the engine as a subprocess behind BentoML's proxy.

A Service is a Python class decorated with @bentoml.service. Each one becomes an independently scalable unit with its own resources, so a GPU-hungry LLM and a lightweight preprocessing step do not have to scale together. An API is a method decorated with @bentoml.api; type hints define the request and response schema. Services call each other through bentoml.depends(), which lets you call another service's API methods as if they were local methods; locally they run in one process tree, while deployed they become separate deployments connected over the network.

Models are declared as class attributes, for example bentoml.models.HuggingFaceModel(...) for a Hugging Face repository or bentoml.models.BentoModel(...) for a model in the local BentoML model store, so the build knows which weights belong to which service. The runtime environment is declared in Python with bentoml.images.Image, which pins the Python version, system packages and requirements. bentoml build freezes code, environment and model references into a Bento, and bentoml containerize turns a Bento into an OCI image for your own cluster.

Serving an LLM with vLLM inside BentoML

The current BentoML vLLM example does not wrap the engine in Python calls at all. It declares the model and GPU, and returns the engine's own command line from a __command__ method. BentoML starts that process and proxies HTTP traffic to it, so clients get vLLM's OpenAI-compatible endpoints unchanged. This follows BentoML's documented example; check the version you install, since the API has changed across releases.

# service.py
import bentoml

MODEL_ID = "meta-llama/Meta-Llama-3.1-8B-Instruct"

image = (
    bentoml.images.Image(python_version="3.12")
    .system_packages("curl", "git")
    .requirements_file("requirements.txt")     # pins vllm and its dependencies
)

@bentoml.service(
    name="llama31-8b",
    image=image,
    envs=[{"name": "HF_TOKEN"}],
    resources={"gpu": 1, "gpu_type": "nvidia-h100-80gb"},
    traffic={"timeout": 300, "concurrency": 64},
)
class LLM:
    hf_model = bentoml.models.HuggingFaceModel(MODEL_ID, exclude=[".pth", ".pt", "original/**/*"])

    def __command__(self) -> list[str]:
        return [
            "vllm", "serve", self.hf_model,
            "--served-model-name", MODEL_ID,
            "--max-model-len", "16384",
            "--gpu-memory-utilization", "0.90",
        ]
bentoml serve                                   # local, OpenAI-compatible on port 3000
bentoml deploy --secret huggingface --scaling-min 0 --scaling-max 3   # BentoCloud
bentoml build && bentoml containerize llama31-8b:latest              # your own cluster

The --max-model-len and --gpu-memory-utilization flags are vLLM flags, not BentoML ones; they control the engine's KV cache budget, explained in paged KV cache. The timeout must cover your longest generation, and the concurrency value is what the autoscaler uses, so it is the single most important number in the file.

Calling the service and streaming tokens

Because the LLM service proxies vLLM's OpenAI-compatible server, any OpenAI client works. Point the base URL at the BentoML service and stream the response, so the user sees the first token as soon as prefill finishes rather than waiting for the whole answer.

import time
from openai import OpenAI

client = OpenAI(base_url="http://localhost:3000/v1", api_key="unused-locally")

t0 = time.perf_counter()
first = None
stream = client.chat.completions.create(
    model="meta-llama/Meta-Llama-3.1-8B-Instruct",
    messages=[{"role": "user", "content": "Explain KV cache paging in three sentences."}],
    max_tokens=256,
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content or ""
    if delta and first is None:
        first = time.perf_counter() - t0          # time to first token
    print(delta, end="", flush=True)
print(f"\nTTFT {first:.2f}s, total {time.perf_counter() - t0:.2f}s")

The data flow for one request is: client to the BentoML HTTP proxy, proxy to the engine process on the same container, engine scheduler to the GPU, and tokens streamed back along the same path. The proxy hop is cheap compared with prefill and decode, but it is one more place where timeouts, request size limits and connection limits apply. Measure time to first token and time per output token separately in this client loop, because the first grows with prompt length and queueing while the second grows with concurrency, and they call for different fixes: shorter prompts or more replicas for the first, a lower concurrency setting for the second.

Composition and adaptive batching

Real LLM applications need more than the model: input validation, a safety or PII classifier, embeddings for retrieval, a reranker. These small models are where BentoML's adaptive batching is useful. Marking an API with batchable=True lets the framework group concurrent requests into one call, bounded by max_batch_size and max_latency_ms; if the service cannot keep within the latency bound it returns HTTP 503.

import os
import bentoml
import httpx

@bentoml.service(resources={"gpu": 1}, traffic={"timeout": 30})
class Embedder:
    def __init__(self):
        from sentence_transformers import SentenceTransformer
        self.model = SentenceTransformer("BAAI/bge-small-en-v1.5", device="cuda")

    @bentoml.api(batchable=True, max_batch_size=64, max_latency_ms=50)
    def embed(self, texts: list[str]) -> list[list[float]]:
        return self.model.encode(texts, normalize_embeddings=True).tolist()

@bentoml.service(resources={"cpu": "2"}, traffic={"timeout": 300, "concurrency": 256})
class Gateway:
    embedder = bentoml.depends(Embedder)

    @bentoml.on_startup
    def init_client(self):
        # LLM_URL points at the LLM service's OpenAI-compatible base URL.
        self.http = httpx.AsyncClient(base_url=os.environ["LLM_URL"], timeout=300)

    @bentoml.api
    async def answer(self, question: str) -> str:
        if len(question) > 8000:
            raise ValueError("question too long")
        [vec] = await self.embedder.to_async.embed([question])
        context = retrieve(vec)                     # your vector store lookup
        r = await self.http.post("/v1/chat/completions", json={
            "model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
            "messages": [{"role": "system", "content": context},
                         {"role": "user", "content": question}],
            "max_tokens": 512})
        r.raise_for_status()
        return r.json()["choices"][0]["message"]["content"]

The rule to remember: use adaptive batching for small fixed-shape models, never in front of an LLM engine. The engine already runs continuous batching at the token level, admitting and retiring sequences every step. Grouping whole LLM requests in BentoML first only adds queueing delay and makes every request in a group wait for the slowest one.

Sizing concurrency for autoscaling

Autoscaling in BentoCloud is driven by concurrency, the number of requests a replica processes at the same time. The documented behaviour is that if concurrency is 32 and 100 requests are in flight, the autoscaler targets ceil(100 / 32) = 4 replicas. So the number you choose should be the concurrency at which one replica still meets your latency target, which you get from a load test, bounded above by what the KV cache can hold. Here is the memory side for the 8B example on one 80 GB GPU.

QuantityCalculationResult
Weights (BF16)8.0 B parameters x 2 bytesabout 16 GB
Engine budget80 GB x 0.90 utilisationabout 72 GB
KV cache space72 - 16 - about 4 GB activations and overheadabout 52 GB
KV per token2 (K and V) x 32 layers x 8 KV heads x 128 dims x 2 bytes128 KiB
Tokens that fit52 GB / 128 KiBabout 400,000 tokens
Sequences at 4,000 tokens each400,000 / 4,000about 100

Memory would allow roughly 100 sequences of 4,000 tokens. Latency usually sets a lower limit: as concurrency rises, each decode step processes more sequences and inter-token latency grows. If your load test shows that p95 time per output token stays under your target up to 64 concurrent requests and breaks at 96, set concurrency to about 64, leaving headroom. Capacity planning at fleet level, including how request mix changes these numbers, is covered in GPU serving capacity estimation.

Two scaling settings matter for GPUs. Stabilisation windows, scale_up_stabilization_window and scale_down_stabilization_window, stop the autoscaler from flapping when traffic is bursty; a GPU replica is expensive to start, so scale down slowly. And scale-to-zero, with minimum replicas 0, means the first request waits for a cold start; BentoCloud's external queue holds requests until a replica is ready, and the timeout does not start until then.

Cold starts and scale-to-zero

A cold start for an LLM replica is node provisioning, image pull, weight download, weight load into GPU memory, engine warm-up such as CUDA graph capture, and the readiness check. For an 8B model the weights alone are about 16 GB; at an effective 500 MB/s from object storage that is over 30 seconds before loading begins, and larger models take minutes. Scale-to-zero is therefore right for development, batch jobs and low-traffic internal tools, and usually wrong for interactive products with a latency SLO. Keep at least one warm replica for those, and if you need to cut startup time, pre-cache weights on the node, keep images small, and measure each phase separately so you know which one to attack.

Operating it

  • Health. Point load balancers at the readiness endpoint /readyz; a replica whose engine is still loading should not receive traffic.
  • Timeouts in layers. Client, gateway, the service traffic.timeout and the engine's own limits should increase from inside out, or the client retries while the GPU is still generating the first answer.
  • Streaming. Use the engine's streaming endpoint or an async generator API so time to first token is visible to users and to your metrics.
  • Pin versions. Pin BentoML, the engine and the model revision in the image. Engine upgrades change memory behaviour and flags.
  • Metrics. Scrape engine metrics (queue length, KV cache usage, tokens per second) as well as BentoML request metrics; request counts alone will not show KV saturation.
  • Shutdown. Use @bentoml.on_shutdown to close clients and flush logs, and make sure scale-down drains in-flight streams.

Failure modes

FailureSymptomFix
Concurrency set above KV capacityEngine preempts or queues; latency spikes but autoscaler does not scaleLower concurrency to the load-tested value
Timeout shorter than long generationsLong answers cut off; clients retry and double GPU loadSet timeout from p99 generation time plus margin
Adaptive batching in front of the LLMHigher latency with no throughput gainBatch only small models; let the engine batch tokens
Scale-to-zero on an interactive appFirst users wait tens of seconds or minutesMinimum one warm replica
Unpinned engine versionNew build fails or uses more memoryPin versions in requirements and test before rollout
Gateway and LLM in one serviceCPU work scales GPU replicasSeparate services with separate resources

Trade-offs and alternatives

Choose BentoML when your team writes Python, your application is several models plus business logic, and you want packaging and autoscaling without building them on Kubernetes yourself. Compare the alternatives honestly.

OptionStrengthWeakness
BentoML plus enginePython-native composition, one artifact, managed scaling optionAnother layer to version and debug
Plain vLLM on KubernetesFewest moving partsYou build packaging, autoscaling and composition
Ray ServeStrong multi-node and placement group storyRay cluster operations
NVIDIA TritonMany backends, ensembles, mature metricsHeavier configuration for Python-first teams

What to do next

  1. Run the vLLM example locally with bentoml serve and call it with the OpenAI SDK.
  2. Compute weight and KV memory for your model and context length as in the table above.
  3. Load-test at increasing concurrency and set concurrency just below the point where p95 time per output token misses your target.
  4. Split CPU logic and small models into their own services; use adaptive batching only for those.
  5. Decide minimum replicas from your latency SLO, and measure each cold-start phase.
  6. Pin BentoML, engine and model revisions, and alert on engine KV usage and queue length.
Key takeaway: BentoML is the packaging, composition and scaling layer, not the engine: let vLLM or another engine batch tokens on the GPU, use BentoML to build one versioned artifact, split CPU and small-model work into separately scaled services, and set the concurrency value from a load test bounded by KV cache memory, because that one number drives autoscaling.