LangChain, LlamaIndex and DSPy are usually compared on developer experience: which has the connectors you need, which abstractions read well, which is easiest to debug. Those questions matter, and the framework comparison for RAG tooling covers them. This article asks a different question that matters as soon as the model runs on your own GPUs: what does each framework make the GPU do?
A self-hosted server such as vLLM or SGLang is only as efficient as the traffic it receives. A framework decides how many requests reach it at once, how prompts are laid out, how many hidden calls a single user turn turns into, and whether structured output is enforced on the server or repaired by retries. Those choices can change GPU cost per answer several times over with no change to the model. You will see how all three frameworks connect to an OpenAI-compatible endpoint, and how concurrency maps onto continuous batching. You will learn how to lay out prompts so prefix caching works and how to enforce schemas on the server, then measure the real token bill with a small instrumentation layer.
The request path from framework to GPU
Every framework here ends in the same place: an HTTP request to a chat completions endpoint. The engine behind it does not know or care which library sent the request. It sees a stream of requests, batches whatever is in flight at each decode step, reuses KV blocks for prompt prefixes it has seen before, and returns tokens. The diagram marks the four decisions the framework layer makes on the GPU's behalf.
This also explains why the frameworks feel interchangeable in a notebook and very different in production. One user and one request at a time hides every one of these effects. They appear when you run a 10,000-question evaluation or put a hundred concurrent users on one replica. The lower layers, from PyTorch down to the kernels, are covered in the Python LLM stack overview.
Connecting the three frameworks to your server
All three frameworks reach a self-hosted server through OpenAI-compatible clients. These are the minimal configurations, using the documented class names. Put the model name the server registered, not a Hugging Face path it does not know, and set retries explicitly rather than inheriting a default.
# LangChain: package langchain-openai
from langchain_openai import ChatOpenAI
lc_llm = ChatOpenAI(
model="llama-3.1-8b-instruct",
base_url="http://gpu-host:8000/v1",
api_key="local",
max_retries=1,
timeout=60,
)
# LlamaIndex: package llama-index-llms-openai-like
from llama_index.llms.openai_like import OpenAILike
li_llm = OpenAILike(
model="llama-3.1-8b-instruct",
api_base="http://gpu-host:8000/v1",
api_key="local",
is_chat_model=True, # default is False, which sends completion-style requests
context_window=8192,
)
# DSPy: provider prefix "openai/" routes through an OpenAI-compatible client
import dspy
ds_lm = dspy.LM("openai/llama-3.1-8b-instruct",
api_base="http://gpu-host:8000/v1", api_key="local", max_tokens=512)
dspy.configure(lm=ds_lm)Two details cause most first-day surprises. LlamaIndex's OpenAILike defaults is_chat_model to False. Left that way, it sends raw completion requests, so the server never applies the model's chat template and answer quality quietly drops. DSPy caches LM calls by default, which is good for development cost. It also means a re-run benchmark can report impossibly low latency, because nothing reached the GPU. Turn caching off when you measure the server.
Concurrency is the GPU batch size
A continuous-batching engine runs one decode step for every active sequence together. Reading the weights from memory dominates a decode step, so a batch of 32 costs little more than a batch of 1, and throughput scales almost linearly until compute or KV memory runs out. The continuous batching article explains the scheduler. The consequence for framework code is direct: the number of requests your code keeps in flight is the GPU's batch size.
Little's law gives the arithmetic. Requests in flight equals arrival rate times latency. A plain Python for-loop over 10,000 evaluation questions, each taking 8 seconds, keeps one request in flight and needs about 22 hours. Run with 64 in flight, and assume latency rises to 12 seconds under the larger batch (an illustrative number; measure yours). The same job then takes 10,000 × 12 / 64, about 31 minutes. Same GPU, same model, forty times faster.
# LangChain: batch() and abatch() accept a concurrency cap through the run config
answers = await chain.abatch(questions, config={"max_concurrency": 64})
# Framework-neutral: bound concurrency yourself with a semaphore
import asyncio
async def run_bounded(items, fn, limit=64):
sem = asyncio.Semaphore(limit)
async def one(x):
async with sem:
return await fn(x)
return await asyncio.gather(*(one(x) for x in items))Pick the limit from the server, not from the client. The engine admits as many sequences as its KV cache holds, so it is useful to send a little more than that and let the server queue the rest. Thousands in flight only move the queue to the client, inflate tail latency and trigger timeouts. Start near the engine's maximum batch size and adjust while watching time-to-first-token.
Prompt layout decides the prefix cache hit rate
Engines with prefix caching store the KV cache for prompt blocks and reuse it whenever a new request starts with the same tokens. vLLM's V1 engine enables automatic prefix caching by default. The prefix caching article explains what the GPU skips. The cache matches from the first token forward, so layout decides the hit rate.
Frameworks make this easy to break without noticing. Common culprits are a template that puts the current date or a request id in the first line of the system prompt and a memory component that inserts conversation history before the instructions. Few-shot examples chosen per query and placed ahead of fixed rules do the same, as do tool schemas serialised in a different key order each call. The fix is mechanical. Keep fixed instructions and tool definitions first and in a stable serialisation. Then add per-user and per-request material, with the question last. You can test the layout before touching the GPU by rendering prompts and measuring shared prefixes.
import os
def prefix_report(rendered_prompts):
"""Average characters shared with the previous prompt: a cheap proxy for cache reuse."""
shared = [len(os.path.commonprefix([a, b]))
for a, b in zip(rendered_prompts, rendered_prompts[1:])]
avg = sum(shared) / len(shared)
total = sum(map(len, rendered_prompts)) / len(rendered_prompts)
print(f"shared prefix {avg:.0f} of {total:.0f} chars ({avg / total:.0%})")If the report shows a few dozen shared characters where you expected thousands, print the first 200 characters of two prompts side by side. The culprit is almost always visible in the first line.
Structured output: repair on the client or constrain on the server
Applications want JSON that matches a schema. There are two ways to get it. Client-side, the framework asks politely in the prompt, parses the reply, and retries on failure. Server-side, the engine constrains decoding so that only tokens valid under the schema can be sampled. Client-side repair spends a full extra generation on every failure, and an 8 percent parse failure rate is 8 percent more GPU time plus a latency spike. Server-side constraint costs a little per token for the grammar mask and guarantees well-formed output, though not correct values.
from pydantic import BaseModel
class Triage(BaseModel):
category: str
urgent: bool
# LangChain: request a JSON-schema response format, enforced by servers that support it
triage_llm = lc_llm.with_structured_output(Triage, method="json_schema")
# vLLM-specific fields go in extra_body; in v0.12 and later the key is "structured_outputs"
# (the older guided_json style fields were removed in that release)
choice_llm = ChatOpenAI(model="llama-3.1-8b-instruct", base_url="http://gpu-host:8000/v1",
api_key="local",
extra_body={"structured_outputs": {"choice": ["billing", "bug", "other"]}})Check two things on your server version: that the standard response_format with a JSON schema is actually enforced rather than ignored, and which extra fields it accepts. Engines have renamed these fields between releases. The serving stacks comparison covers grammar support across engines.
Worked example: measuring a support assistant
Here is a support assistant built on LangChain. The numbers are illustrative but typical. Each user question triggers a query-rewrite call (600 prompt tokens, 40 output), retrieval, an answer call (3,500 prompt tokens: 1,200 of fixed rules, 2,000 of retrieved context, 300 of question and history; 350 output) and a guard call that classifies the answer (900 prompt, 10 output). That is 5,000 prefill tokens and 400 decode tokens per question, three requests where a notebook demo suggested one.
In the first version the answer template began with the current timestamp, so prefix reuse was zero. The guard parsed free text with client-side retries and failed about 8 percent of the time, and the evaluation script looped sequentially. Three changes followed. Fixed text moved to the front of all three templates, making about 500 + 1,200 + 800 = 2,500 tokens per question cacheable, half of all prefill. The guard switched to a server-side choice constraint, removing its retries. The evaluation moved to abatch with max_concurrency of 64. Prefill work per question halved, guard retries disappeared, and the nightly evaluation went from most of a day to under an hour. No change to the model, the GPU or the engine.
Do not trust estimates like these; count. The wrapper below records tokens and latency per pipeline step from the usage field that OpenAI-compatible servers return, using the official openai client with retries disabled so every attempt is visible.
import time
from collections import defaultdict
from openai import AsyncOpenAI
client = AsyncOpenAI(base_url="http://gpu-host:8000/v1", api_key="local",
max_retries=0, timeout=60)
LEDGER = defaultdict(lambda: [0, 0, 0, 0.0]) # step -> calls, prompt, completion, seconds
async def call(step, messages, **kw):
t0 = time.perf_counter()
r = await client.chat.completions.create(
model="llama-3.1-8b-instruct", messages=messages, **kw)
row = LEDGER[step]
row[0] += 1
row[1] += r.usage.prompt_tokens
row[2] += r.usage.completion_tokens
row[3] += time.perf_counter() - t0
return r.choices[0].message.content
def report(questions):
for step, (n, pt, ct, sec) in sorted(LEDGER.items()):
print(f"{step:10s} calls/q={n/questions:.2f} prompt/q={pt/questions:.0f} "
f"completion/q={ct/questions:.0f} mean_s={sec/n:.2f}")Run your pipeline through it on 200 representative questions and you get calls per question, prefill and decode tokens per question per step, and the step that dominates. That table is the only reliable input to GPU capacity planning for an application built on a framework.
Failure modes
- Agent loops without a step cap. A tool-calling agent that never reaches a final answer keeps issuing requests with a growing context. Prefill cost grows with every turn. Cap iterations and total tokens per user request.
- Stacked retries. The openai client retries twice by default, a framework may retry parse failures, and a load balancer may retry timeouts. Three layers of retry turn one overloaded moment into a retry storm. Retry in exactly one layer.
- Client timeouts shorter than queue time. The client gives up and resends while the original request is still generating, unless the engine sees the disconnect. Proxies that hold the upstream connection hide it. Set timeouts above the p99 latency under load.
- Chat template mismatch. Completion-mode requests, or a server launched without the right chat template, produce degraded answers that look like a model problem.
- Silent response caching. Framework caches make benchmarks lie. Disable them for load tests.
- Optimizer runs on the serving fleet. A DSPy optimizer or prompt search can issue thousands of calls in a burst. Point it at a separate replica or a rate-limited key, not the production endpoint.
- Context bloat from retrieval defaults. Retrieving more chunks than the answer needs multiplies prefill. Measure answer quality against chunk count and cut to the knee.
Trade-offs
| Choice | GPU-side effect | Cost |
|---|---|---|
| High client concurrency | Large decode batches, high throughput | Higher per-request latency; needs backpressure |
| Stable prompt prefix | Prefill skipped for shared blocks | Template discipline; dynamic data moves later |
| Server-side constrained decoding | No repair retries, valid structure | Engine-specific fields; small per-token overhead |
| Client-side parse and retry | Portable across any endpoint | Extra generations on every failure |
| Framework abstractions | Fast to build multi-step pipelines | Hidden calls you must instrument to see |
| Thin direct client | Every request explicit | You write the orchestration yourself |
What to do next
- List every model call your pipeline makes per user turn, including rewrites, guards and retries, and wrap them with a ledger like the one above.
- Run 200 representative questions and record calls, prompt tokens and completion tokens per step.
- Render 50 prompts per template and run the prefix report; move all dynamic content after the fixed instructions.
- Replace sequential loops in evaluation and batch jobs with bounded concurrency set near the engine's maximum batch size.
- Move schema-shaped outputs to server-side constrained decoding and verify that your engine version enforces it.
- Choose one layer for retries and disable the others; set client timeouts above p99 latency under load.
- Re-run the ledger after each change and keep the before-and-after table with your capacity plan.