BentoML is an open-source Python framework (Apache 2.0) for turning model code into production inference services. You write a Python class, mark methods as APIs, declare the resources the service needs, and BentoML gives you an HTTP server, a versioned deployable artifact called a Bento, container builds, and service-to-service composition. Its managed platform, BentoCloud, adds autoscaling and scale-to-zero. In February 2026 Modular, the company behind the MAX inference engine and the Mojo language, announced that it had acquired BentoML; the framework remains open source.
For LLM serving the most important point is what BentoML is not. It is not an inference engine. Token generation speed comes from engines such as vLLM, SGLang or TensorRT-LLM, which implement continuous batching, paged KV caches and fused kernels; the trade-offs between them are covered in LLM serving stacks compared. BentoML sits one layer up: it packages an engine with its model and environment, composes it with the CPU and GPU services around it, and scales replicas. This article explains that layer, shows verified code against the current API, works through capacity and autoscaling numbers, and lists the failure modes.
Architecture: services, APIs, models and Bentos
A Service is a Python class decorated with @bentoml.service. Each one becomes an independently scalable unit with its own resources, so a GPU-hungry LLM and a lightweight preprocessing step do not have to scale together. An API is a method decorated with @bentoml.api; type hints define the request and response schema. Services call each other through bentoml.depends(), which lets you call another service's API methods as if they were local methods; locally they run in one process tree, while deployed they become separate deployments connected over the network.
Models are declared as class attributes, for example bentoml.models.HuggingFaceModel(...) for a Hugging Face repository or bentoml.models.BentoModel(...) for a model in the local BentoML model store, so the build knows which weights belong to which service. The runtime environment is declared in Python with bentoml.images.Image, which pins the Python version, system packages and requirements. bentoml build freezes code, environment and model references into a Bento, and bentoml containerize turns a Bento into an OCI image for your own cluster.
Serving an LLM with vLLM inside BentoML
The current BentoML vLLM example does not wrap the engine in Python calls at all. It declares the model and GPU, and returns the engine's own command line from a __command__ method. BentoML starts that process and proxies HTTP traffic to it, so clients get vLLM's OpenAI-compatible endpoints unchanged. This follows BentoML's documented example; check the version you install, since the API has changed across releases.
# service.py
import bentoml
MODEL_ID = "meta-llama/Meta-Llama-3.1-8B-Instruct"
image = (
bentoml.images.Image(python_version="3.12")
.system_packages("curl", "git")
.requirements_file("requirements.txt") # pins vllm and its dependencies
)
@bentoml.service(
name="llama31-8b",
image=image,
envs=[{"name": "HF_TOKEN"}],
resources={"gpu": 1, "gpu_type": "nvidia-h100-80gb"},
traffic={"timeout": 300, "concurrency": 64},
)
class LLM:
hf_model = bentoml.models.HuggingFaceModel(MODEL_ID, exclude=[".pth", ".pt", "original/**/*"])
def __command__(self) -> list[str]:
return [
"vllm", "serve", self.hf_model,
"--served-model-name", MODEL_ID,
"--max-model-len", "16384",
"--gpu-memory-utilization", "0.90",
]bentoml serve # local, OpenAI-compatible on port 3000
bentoml deploy --secret huggingface --scaling-min 0 --scaling-max 3 # BentoCloud
bentoml build && bentoml containerize llama31-8b:latest # your own clusterThe --max-model-len and --gpu-memory-utilization flags are vLLM flags, not BentoML ones; they control the engine's KV cache budget, explained in paged KV cache. The timeout must cover your longest generation, and the concurrency value is what the autoscaler uses, so it is the single most important number in the file.
Calling the service and streaming tokens
Because the LLM service proxies vLLM's OpenAI-compatible server, any OpenAI client works. Point the base URL at the BentoML service and stream the response, so the user sees the first token as soon as prefill finishes rather than waiting for the whole answer.
import time
from openai import OpenAI
client = OpenAI(base_url="http://localhost:3000/v1", api_key="unused-locally")
t0 = time.perf_counter()
first = None
stream = client.chat.completions.create(
model="meta-llama/Meta-Llama-3.1-8B-Instruct",
messages=[{"role": "user", "content": "Explain KV cache paging in three sentences."}],
max_tokens=256,
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content or ""
if delta and first is None:
first = time.perf_counter() - t0 # time to first token
print(delta, end="", flush=True)
print(f"\nTTFT {first:.2f}s, total {time.perf_counter() - t0:.2f}s")The data flow for one request is: client to the BentoML HTTP proxy, proxy to the engine process on the same container, engine scheduler to the GPU, and tokens streamed back along the same path. The proxy hop is cheap compared with prefill and decode, but it is one more place where timeouts, request size limits and connection limits apply. Measure time to first token and time per output token separately in this client loop, because the first grows with prompt length and queueing while the second grows with concurrency, and they call for different fixes: shorter prompts or more replicas for the first, a lower concurrency setting for the second.
Composition and adaptive batching
Real LLM applications need more than the model: input validation, a safety or PII classifier, embeddings for retrieval, a reranker. These small models are where BentoML's adaptive batching is useful. Marking an API with batchable=True lets the framework group concurrent requests into one call, bounded by max_batch_size and max_latency_ms; if the service cannot keep within the latency bound it returns HTTP 503.
import os
import bentoml
import httpx
@bentoml.service(resources={"gpu": 1}, traffic={"timeout": 30})
class Embedder:
def __init__(self):
from sentence_transformers import SentenceTransformer
self.model = SentenceTransformer("BAAI/bge-small-en-v1.5", device="cuda")
@bentoml.api(batchable=True, max_batch_size=64, max_latency_ms=50)
def embed(self, texts: list[str]) -> list[list[float]]:
return self.model.encode(texts, normalize_embeddings=True).tolist()
@bentoml.service(resources={"cpu": "2"}, traffic={"timeout": 300, "concurrency": 256})
class Gateway:
embedder = bentoml.depends(Embedder)
@bentoml.on_startup
def init_client(self):
# LLM_URL points at the LLM service's OpenAI-compatible base URL.
self.http = httpx.AsyncClient(base_url=os.environ["LLM_URL"], timeout=300)
@bentoml.api
async def answer(self, question: str) -> str:
if len(question) > 8000:
raise ValueError("question too long")
[vec] = await self.embedder.to_async.embed([question])
context = retrieve(vec) # your vector store lookup
r = await self.http.post("/v1/chat/completions", json={
"model": "meta-llama/Meta-Llama-3.1-8B-Instruct",
"messages": [{"role": "system", "content": context},
{"role": "user", "content": question}],
"max_tokens": 512})
r.raise_for_status()
return r.json()["choices"][0]["message"]["content"]The rule to remember: use adaptive batching for small fixed-shape models, never in front of an LLM engine. The engine already runs continuous batching at the token level, admitting and retiring sequences every step. Grouping whole LLM requests in BentoML first only adds queueing delay and makes every request in a group wait for the slowest one.
Sizing concurrency for autoscaling
Autoscaling in BentoCloud is driven by concurrency, the number of requests a replica processes at the same time. The documented behaviour is that if concurrency is 32 and 100 requests are in flight, the autoscaler targets ceil(100 / 32) = 4 replicas. So the number you choose should be the concurrency at which one replica still meets your latency target, which you get from a load test, bounded above by what the KV cache can hold. Here is the memory side for the 8B example on one 80 GB GPU.
| Quantity | Calculation | Result |
|---|---|---|
| Weights (BF16) | 8.0 B parameters x 2 bytes | about 16 GB |
| Engine budget | 80 GB x 0.90 utilisation | about 72 GB |
| KV cache space | 72 - 16 - about 4 GB activations and overhead | about 52 GB |
| KV per token | 2 (K and V) x 32 layers x 8 KV heads x 128 dims x 2 bytes | 128 KiB |
| Tokens that fit | 52 GB / 128 KiB | about 400,000 tokens |
| Sequences at 4,000 tokens each | 400,000 / 4,000 | about 100 |
Memory would allow roughly 100 sequences of 4,000 tokens. Latency usually sets a lower limit: as concurrency rises, each decode step processes more sequences and inter-token latency grows. If your load test shows that p95 time per output token stays under your target up to 64 concurrent requests and breaks at 96, set concurrency to about 64, leaving headroom. Capacity planning at fleet level, including how request mix changes these numbers, is covered in GPU serving capacity estimation.
Two scaling settings matter for GPUs. Stabilisation windows, scale_up_stabilization_window and scale_down_stabilization_window, stop the autoscaler from flapping when traffic is bursty; a GPU replica is expensive to start, so scale down slowly. And scale-to-zero, with minimum replicas 0, means the first request waits for a cold start; BentoCloud's external queue holds requests until a replica is ready, and the timeout does not start until then.
Cold starts and scale-to-zero
A cold start for an LLM replica is node provisioning, image pull, weight download, weight load into GPU memory, engine warm-up such as CUDA graph capture, and the readiness check. For an 8B model the weights alone are about 16 GB; at an effective 500 MB/s from object storage that is over 30 seconds before loading begins, and larger models take minutes. Scale-to-zero is therefore right for development, batch jobs and low-traffic internal tools, and usually wrong for interactive products with a latency SLO. Keep at least one warm replica for those, and if you need to cut startup time, pre-cache weights on the node, keep images small, and measure each phase separately so you know which one to attack.
Operating it
- Health. Point load balancers at the readiness endpoint
/readyz; a replica whose engine is still loading should not receive traffic. - Timeouts in layers. Client, gateway, the service
traffic.timeoutand the engine's own limits should increase from inside out, or the client retries while the GPU is still generating the first answer. - Streaming. Use the engine's streaming endpoint or an async generator API so time to first token is visible to users and to your metrics.
- Pin versions. Pin BentoML, the engine and the model revision in the image. Engine upgrades change memory behaviour and flags.
- Metrics. Scrape engine metrics (queue length, KV cache usage, tokens per second) as well as BentoML request metrics; request counts alone will not show KV saturation.
- Shutdown. Use
@bentoml.on_shutdownto close clients and flush logs, and make sure scale-down drains in-flight streams.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Concurrency set above KV capacity | Engine preempts or queues; latency spikes but autoscaler does not scale | Lower concurrency to the load-tested value |
| Timeout shorter than long generations | Long answers cut off; clients retry and double GPU load | Set timeout from p99 generation time plus margin |
| Adaptive batching in front of the LLM | Higher latency with no throughput gain | Batch only small models; let the engine batch tokens |
| Scale-to-zero on an interactive app | First users wait tens of seconds or minutes | Minimum one warm replica |
| Unpinned engine version | New build fails or uses more memory | Pin versions in requirements and test before rollout |
| Gateway and LLM in one service | CPU work scales GPU replicas | Separate services with separate resources |
Trade-offs and alternatives
Choose BentoML when your team writes Python, your application is several models plus business logic, and you want packaging and autoscaling without building them on Kubernetes yourself. Compare the alternatives honestly.
| Option | Strength | Weakness |
|---|---|---|
| BentoML plus engine | Python-native composition, one artifact, managed scaling option | Another layer to version and debug |
| Plain vLLM on Kubernetes | Fewest moving parts | You build packaging, autoscaling and composition |
| Ray Serve | Strong multi-node and placement group story | Ray cluster operations |
| NVIDIA Triton | Many backends, ensembles, mature metrics | Heavier configuration for Python-first teams |
What to do next
- Run the vLLM example locally with
bentoml serveand call it with the OpenAI SDK. - Compute weight and KV memory for your model and context length as in the table above.
- Load-test at increasing concurrency and set
concurrencyjust below the point where p95 time per output token misses your target. - Split CPU logic and small models into their own services; use adaptive batching only for those.
- Decide minimum replicas from your latency SLO, and measure each cold-start phase.
- Pin BentoML, engine and model revisions, and alert on engine KV usage and queue length.