Together AI sells access to open-weight models (Llama, Qwen, DeepSeek, gpt-oss, Gemma and many more) running on GPUs it operates. For an engineering team, the useful question is not "is it a good provider" but "which of its products matches which of my workloads, and what does each one cost me in latency, money and operational risk". The answer depends on how GPUs behave under load: a shared fleet amortises idle time across many customers, a reserved replica turns idle time into your bill, and an offline queue lets the provider fill gaps.
This article explains the four products (serverless inference, dedicated endpoints, the batch API and fine-tuning) in those terms, shows working integration code against the OpenAI-compatible API, builds a break-even model you can fill with your own measurements, and lists the failure modes that bite in production. Product details were checked against Together's documentation on 2026-10-06; prices and rate limits change often and are deliberately not quoted, so read the current pricing page before you plug numbers in.
Four products, one GPU fleet
Serverless inference is a shared pool of popular models billed per token. You provision nothing. The provider batches your requests together with other customers' on the same GPUs, which is why per-token prices can be low: decode on a large model is memory-bandwidth bound, and every extra sequence in the batch reuses the weights already streamed from HBM. The cost of that sharing is variance. Your latency depends on everyone else's load, and you get 429 or 503 responses under pressure.
Dedicated endpoints are single-tenant replicas of one model, billed per minute while running. A deployment profile bundles the hardware, the quantisation and, where the configuration enables it, speculative decoding; the profile is chosen at creation and cannot be changed afterwards. Replica bounds are set at deploy time and can be adjusted later, and an inactivity timeout can stop an unused deployment so the meter stops. You get predictable latency and no shared rate limits, and you pay for every idle minute.
Batch accepts a JSONL file of requests and returns results within a best-effort 24-hour window. Documented limits at the time of writing: up to 50,000 requests per batch, 100 MB per input file and 10 MB per line. The discount is "up to 50%" off serverless rates and applies to selected models only, so check yours. Note that this is not the OpenAI-shaped Batch or Files API; code written for OpenAI's batch endpoint will not port unchanged.
Fine-tuning trains either a LoRA adapter (the default) or full weights from a JSONL dataset. Since serverless LoRA inference has been discontinued, the output of a fine-tune is served from a dedicated endpoint, or downloaded and served yourself. That single fact changes the economics of a fine-tune: you now owe the cost of a replica, not just the training run.
Architecture: what sits behind the endpoint
The diagram shows the shape that matters for design. One authenticated API fronts three serving paths, and fine-tuning feeds the dedicated path. Because all three serving paths speak the same chat-completions dialect, you can move a workload between them by changing the model identifier and, for batch, the submission mechanism, without rewriting prompts or parsers.
What runs behind the API on the GPU side is the provider's business, but the levers you see are the ones any LLM server exposes: weight precision (a quantised profile streams fewer bytes per decode step, so more tokens per second on the same GPU, with some quality risk), speculative decoding (a small draft model proposes tokens that the large model verifies in one forward pass, which trades spare compute for fewer bandwidth-bound steps), and batch size (higher throughput per GPU at the cost of per-request latency). Together's research group has published widely on attention kernels and inference speed, but your decision should rest on what you measure on your own prompts, not on headline benchmarks.
Calling the API correctly
The API is OpenAI-compatible at https://api.together.ai/v1, so the official OpenAI Python client works with a changed base URL and key. Model IDs are namespaced (meta-llama/Llama-3.3-70B-Instruct-Turbo, openai/gpt-oss-120b), unlike OpenAI's flat names. A minimal call with explicit retries and usage accounting:
import os, random, time
import openai
def log_usage(model, prompt_toks, completion_toks, cached_toks):
... # your metrics sink: one row per request
client = openai.OpenAI(
api_key=os.environ["TOGETHER_API_KEY"],
base_url="https://api.together.ai/v1",
max_retries=0, # we own the retry policy below
timeout=60,
)
MODEL = "meta-llama/Llama-3.3-70B-Instruct-Turbo" # pin it; see failure modes
def cached_tokens(usage) -> int:
"""Cached prompt tokens live in different places for different models."""
u = usage.model_dump() if hasattr(usage, "model_dump") else dict(usage)
nested = (u.get("prompt_tokens_details") or {}).get("cached_tokens")
return int(nested if nested is not None else u.get("cached_tokens") or 0)
def chat(messages, max_attempts=6, **kw):
for attempt in range(max_attempts):
try:
r = client.chat.completions.create(model=MODEL, messages=messages, **kw)
u = r.usage
log_usage(MODEL, u.prompt_tokens, u.completion_tokens, cached_tokens(u))
return r.choices[0].message.content
except (openai.RateLimitError, openai.InternalServerError,
openai.APITimeoutError, openai.APIConnectionError):
if attempt == max_attempts - 1:
raise
# full jitter: 0.5s, 1s, 2s ... capped at 30s
time.sleep(random.uniform(0, min(30, 0.5 * 2 ** attempt)))Three details in that snippet come straight from the documented differences. First, cached-token counts are nested under usage.prompt_tokens_details for some models and flattened onto usage for others, so a parser that assumes one location silently reports zero cache hits for the other. Second, error objects are OpenAI-shaped but carry Together's own type and code values, so branch on HTTP status (the SDK's exception classes) rather than on OpenAI error codes. Third, the documented guidance for both 429 and 503 is to spread load and back off exponentially; jittered backoff stops a fleet of your workers from retrying in lockstep and causing the next 429.
Serverless or dedicated: a break-even model
Serverless versus dedicated is a utilisation question. Let P be the serverless price per million tokens for your input/output mix, C the dedicated price per replica-minute, and R the tokens per second one replica sustains on your traffic at your latency target. A replica running at utilisation u produces 60·R·u tokens per minute, so its cost per million tokens is C·106 / (60·R·u). Setting that equal to P gives the break-even utilisation:
def breakeven_utilisation(P_per_mtok, C_per_min, R_tok_s):
return C_per_min * 1e6 / (60 * R_tok_s * P_per_mtok)
def monthly(P_per_mtok, C_per_min, R_tok_s, tokens_month, replicas, hours_up):
serverless = tokens_month / 1e6 * P_per_mtok
dedicated = replicas * hours_up * 60 * C_per_min
capacity = replicas * hours_up * 3600 * R_tok_s
return serverless, dedicated, tokens_month / capacity # last = utilisationWorked example, with hypothetical numbers (not Together's prices): P = $0.90 per million tokens, C = $0.06 per replica-minute, and a load test shows R = 2,000 tokens per second per replica at your p95 latency target. Break-even utilisation is 0.06·106 / (60·2,000·0.90) = 0.56. If your traffic keeps one replica 56% busy around the clock, the two options cost the same; below that, serverless is cheaper; above it, dedicated wins and also buys latency isolation.
| Workload (hypothetical) | Tokens per 30-day month | Serverless | Dedicated, 1 replica 24x7 | Utilisation |
|---|---|---|---|---|
| Internal assistant | 300M | $270 | $2,592 | 5.8% |
| Customer chat | 3,000M | $2,700 | $2,592 | 58% |
| Bulk extraction | 4,500M | $4,050 | $2,592 | 87% |
Two corrections usually move the answer. Real traffic is diurnal, so a replica sized for peak sits mostly idle at night; autoscaling between a minimum and maximum replica count recovers part of that, at the price of cold-start time when scaling up. And much bulk work does not need to be online at all: if it tolerates a 24-hour turnaround and its model carries the batch discount, batch undercuts both columns.
Fine-tuning and where the model lands
Fine-tuning follows a check, upload, train, deploy sequence. Datasets are JSONL in one of several documented shapes: conversational ({"messages": [...]}), instruction ({"prompt": ..., "completion": ...}), preference pairs, generic text, tool-calling and reasoning. Validate locally before uploading, because a job that fails on line 40,000 has already cost you queue time:
from together import Together
from together.lib.utils import check_file
tg = Together() # reads TOGETHER_API_KEY
report = check_file("train.jsonl")
assert report["is_check_passed"], report
f = tg.files.upload(file="train.jsonl", purpose="fine-tune", check=True)
job = tg.fine_tuning.create(
training_file=f.id,
model="Qwen/Qwen3.5-27B", # base model id from the fine-tuning catalog
lora=True, # default; False = full fine-tune
# optional: lora_r, lora_alpha, lora_dropout, lora_trainable_modules
)
print(job.id) # poll status in the dashboard or SDK, then deploy the output modelPlan the serving cost before you train. A LoRA adapter is tens to hundreds of megabytes, depending on rank and target modules, but it now needs a dedicated endpoint, so a fine-tune serving 100 requests a day may cost more per month than the same traffic on a larger serverless base model with a better prompt. Fine-tune when you have a measured quality gap that prompting and retrieval cannot close, enough traffic to keep a replica busy, or a latency target that a smaller tuned model meets and a larger general one does not.
Operating it in production
Run a provider like this the way you would run any dependency on the hot path.
- Pin model IDs and watch for deprecations. Catalogues change; a model you depend on can be retired or replaced by a new revision. Keep the ID in config, re-run your eval suite when it changes, and alert on 404 responses for model IDs.
- Measure R yourself. Tokens per second depends on prompt length, output length, concurrency and the deployment profile. Load-test at three concurrency levels with replayed production prompts before committing to a dedicated endpoint.
- Log usage per request: model, prompt tokens, completion tokens, cached tokens, latency to first token, total latency, HTTP status. This is the data that answers "should we move to dedicated" next quarter.
- Keep a fallback. Because the API is OpenAI-shaped, a second provider or a self-hosted vLLM endpoint can sit behind the same client interface; see the routing patterns in multi-provider LLM serving and provider failover.
- Check compliance placement. Placement policies (for example a HIPAA placement option) are documented as a dedicated-endpoint deployment setting; confirm the terms that apply to each product before sending regulated data anywhere.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Retry storm on 429 | Error rate climbs after a burst, then oscillates | Jittered exponential backoff, client-side concurrency cap |
| Cache hits read as zero | Cost model overstates spend for prompt-cached traffic | Parse both cached_tokens locations |
| Ported OpenAI batch code | Calls to batch/files endpoints fail | Use Together's batch flow; it is not OpenAI-shaped |
| Fine-tune with no serving plan | Surprise per-minute bill after training | Price the dedicated endpoint before the job |
| Idle dedicated replica | Steady spend with near-zero traffic | Inactivity timeout, min replicas sized to the trough |
| Profile chosen in a hurry | Quality drop from quantisation, cannot change in place | Eval each profile first; redeploy to switch |
| Silent model swap | Eval scores move with no code change | Pin IDs, run evals on catalogue changes |
Trade-offs
Compared with running your own GPUs, Together removes cluster operations, driver and kernel tuning, and capacity planning for spiky traffic, and gives you per-token billing at low volume. You give up control of the serving stack (engine flags, KV cache policy, custom kernels), data-path visibility, and the lowest possible cost at high steady utilisation, where owned or reserved GPUs win. Compared with closed-model APIs, you gain open weights you can later self-host if prices or policies change, at the cost of choosing and evaluating models yourself. For background on the engines underneath any such service, see LLM serving stacks compared; for why mixture-of-experts models price differently, see MoE inference cost.
What to do next
- Point the OpenAI client at the Together base URL in a staging service and log the usage fields from every response, including both cached-token locations.
- Add jittered exponential backoff for 429/503 and a concurrency cap per worker.
- Replay a day of production prompts against two candidate models and score them on your eval set before picking one.
- Load-test the chosen model to measure R at your latency target, then compute the break-even utilisation with current prices.
- Move offline jobs to batch if their model carries the discount and they tolerate 24 hours.
- Before any fine-tune, write down the dedicated-endpoint cost it implies; then follow LoRA fine-tuning practice for the run itself.
- Put a fallback provider behind the same interface and test the failover path.