Fireworks AI is a hosted inference platform for open-weight models. You send OpenAI-style requests to https://api.fireworks.ai/inference/v1 and it runs them on GPUs it operates, either on a shared serverless pool billed per token or on dedicated on-demand deployments billed per GPU-second. It also serves fine-tuned LoRA adapters and supports structured output, function calling and speculative decoding. Its serving engine is marketed under the FireAttention name; Fireworks does not publish enough detail for an outsider to describe its internals, so this article treats it as a black box and focuses on what you control and can measure.
That focus matters because the decisions that dominate latency and cost on a hosted platform are the same ones you would face on your own cluster: how many bytes of weights and KV cache each decode step must read, how many tokens each forward pass produces, and how much idle capacity you pay for. The article walks through calling the API, creating deployments with firectl, the GPU arithmetic behind precision and accelerator choice, speculation and Predicted Outputs, structured output caveats, LoRA addons, and a measurement harness, then lists failure modes and a checklist. For the serverless-versus-dedicated cost break-even on a comparable platform, see Together AI in depth.
Architecture: one endpoint, three ways to get GPUs
From the client's side there is one endpoint, and the model string decides where the request runs. A serverless model is addressed as accounts/fireworks/models/<name>. A dedicated deployment is addressed as accounts/<account>/deployments/<id>, or as a model string with a # suffix naming the deployment. LoRA addons are loaded onto a base model and addressed by their own model id.
Calling it needs nothing Fireworks-specific beyond the base URL. Pin a timeout, stream long outputs so time-to-first-token is visible, and keep the model string in configuration so moving from serverless to a deployment is a config change.
import os, time
from openai import OpenAI
client = OpenAI(base_url="https://api.fireworks.ai/inference/v1",
api_key=os.environ["FIREWORKS_API_KEY"], timeout=60)
MODEL = os.environ.get("FW_MODEL", "accounts/fireworks/models/<model-name>")
def ask(messages, **kw):
t0, first, chunks = time.perf_counter(), None, []
stream = client.chat.completions.create(model=MODEL, messages=messages,
stream=True, **kw)
for ev in stream:
delta = ev.choices[0].delta.content if ev.choices else None
if delta:
first = first or time.perf_counter()
chunks.append(delta)
end = time.perf_counter()
return "".join(chunks), {"ttft_s": (first or end) - t0, "total_s": end - t0}
Serverless or a deployment, and how to create one
Serverless is right while traffic is light or spiky and the model you want is in the shared catalogue. Move to an on-demand deployment when you need predictable latency, a model that is not served publicly, or sustained volume where paying per GPU-second beats paying per token. Deployments are created with the firectl CLI:
# find shapes validated for this model
firectl deployment-shape-version match --model accounts/fireworks/models/<model-name>
# create a deployment from a shape, pinning hardware, precision and scaling
firectl deployment create accounts/fireworks/models/<model-name> \
--deployment-shape <shape-name> \
--accelerator-type NVIDIA_H100_80GB \
--precision FP8 \
--region US \
--min-replica-count 1 --max-replica-count 4 \
--waitDeployment shapes are pre-configured templates in three families: Fast for low latency, Throughput for the lowest cost per token at scale, and Minimal for the lowest absolute cost. The documentation warns that omitting a shape skips validation and is the most common cause of failed deployment creation, so always pass one. Accelerator types listed at the time of writing include NVIDIA_A100_80GB, NVIDIA_H100_80GB, NVIDIA_H200_141GB, NVIDIA_B200_180GB, NVIDIA_B300_288GB, AMD_MI325X_256GB and AMD_MI350X_288GB; check the current list before scripting it.
Scaling has one sharp edge. With a minimum of zero replicas, a deployment scales to zero after an hour of inactivity, and the next request waits for a replica to load the model. That is fine for evaluation jobs and wrong for a user-facing path. Set --min-replica-count 1 for anything interactive, and treat the cost of that idle replica as the price of a predictable first token.
What the GPU is doing: precision, KV cache and fit
Choosing precision and accelerator is easier with the arithmetic in front of you. Generating one token for one sequence reads every weight once and the whole KV cache of that sequence once; at small batch sizes that memory traffic, not arithmetic, sets the speed. Batching amortises the weight reads across sequences, but KV reads grow with every sequence and every token of context.
Worked example. Take a 70-billion-parameter dense model with a Llama-3-70B-like layout: 80 layers, 8 key-value heads of dimension 128 under grouped query attention. In 16-bit precision the weights are about 140 GB, which does not fit on one 80 GB GPU. In FP8 they are about 70 GB. The KV cache per token is 80 layers x 8 heads x 128 x 2 (key and value) x 2 bytes = 327,680 bytes, about 0.33 MB. A single 32k-token conversation therefore holds about 10.7 GB of KV cache.
| Configuration | Weights | Left for KV | 32k-token sequences that fit |
|---|---|---|---|
| FP16 on 2 x H100 80GB | 140 GB | about 20 GB minus overheads | 1, barely |
| FP8 on 1 x H100 80GB | 70 GB | about 10 GB minus overheads | 0 at 32k; short contexts only |
| FP8 on 1 x H200 141GB (if FP8 is offered on H200; confirm in current docs) | 70 GB | about 70 GB minus overheads | about 6 |
| FP8 on 2 x H100 80GB | 70 GB | about 90 GB minus overheads | about 8 |
The table ignores activation memory and allocator overhead, so real numbers are lower, but the shape of the decision is clear: precision decides whether the model fits, and the leftover memory decides how many long conversations a replica can hold at once, which is what throughput really means for chat workloads. FP8 also halves the bytes read per decode step and runs matrix multiplies on the FP8 tensor cores of Hopper and later GPUs. Fireworks documents FP16 as the default with FP8 as the main quantisation option and estimates a 30 to 50 percent cost reduction, while noting that changed numerics can alter outputs. Its quantisation page also states that quantised deployments are served on H100 GPUs; given the newer accelerators in the list, confirm current support for your target hardware. Whatever the docs say, run your evaluation set on both precisions before switching. More on this trade in inference latency.
Speculative decoding and Predicted Outputs
Speculative decoding attacks the one-token-per-forward-pass limit. A cheap proposer guesses the next k tokens; the target model scores all k in a single forward pass, which costs about the same memory traffic as generating one token; the longest prefix that matches what the target would have produced is accepted. With a per-token acceptance probability a, the expected tokens per verification pass is (1 - ak+1) / (1 - a). At a = 0.8 and k = 4 that is about 3.4 tokens per pass; at a = 0.5 it falls to about 1.9. Acceptance, not k, decides the gain. The mechanism is covered in speculative decoding.
Fireworks exposes two sources of drafts. Dedicated deployments can be configured with model-based speculative decoding. Independently, any request can carry a Predicted Output: text you expect the answer to largely repeat, such as the file being edited in a code-rewrite request. The model verifies your prediction instead of a draft model's, and matching spans are emitted at verification speed.
original = open("handler.py").read()
resp = client.chat.completions.create(
model=MODEL,
temperature=0, # recommended for predictions
max_tokens=4096, # also caps prediction length (default 2048)
messages=[{"role": "user",
"content": "Rename get_user to fetch_user. Return the full file.\n\n"
+ original}],
prediction={"type": "content", "content": original},
)The documentation is explicit about the limits: the prediction is capped by max_tokens and defaults to 2048 tokens, temperature 0 is recommended, and a prediction that diverges substantially from the real output can make generation slower, not faster. On on-demand deployments a rewrite_speculation=True option is described as giving further speed-ups for this pattern. Use predictions for edits and reformatting, never for open-ended answers.
Structured output and LoRA addons
Structured output constrains decoding so that only tokens consistent with a schema or grammar can be sampled. Fireworks accepts response_format={"type": "json_object"} for any valid JSON and {"type": "json_schema", "json_schema": {...}} for a specific schema, plus a grammar mode for non-JSON formats. Three caveats from the docs bite in production. Tell the model in the prompt to answer in JSON and include the schema there too, or it may emit whitespace until it hits the token limit. Check finish_reason == "length", because a truncated response is invalid JSON even though decoding was constrained. And on reasoning models, json_schema disables reasoning output, so if you need both, put the schema in the prompt and validate afterwards.
LoRA addons. Fine-tuned adapters are small, so many can share one base model's weights; Fireworks supports training, uploading and serving them. Its documentation notes that LoRA addons on serverless have higher latency than the base model. If an adapter carries production traffic, load it onto a dedicated deployment of the base model and measure again. How multi-adapter serving batches requests across adapters is explained in multi-LoRA serving.
Measuring and operating it
Treat the platform as a dependency you measure continuously, not a benchmark you ran once. A minimal harness replays a fixed set of real prompts through ask and records time to first token, total time, output tokens and errors per configuration: serverless, each deployment shape, FP16 against FP8, with and without predictions. Report the 50th and 95th percentiles, never the mean, and keep the prompt set versioned so results stay comparable across months.
- Fallback. Route through a thin client that can switch the model string to a second deployment, region or provider on timeouts and 5xx responses, with a circuit breaker so retries do not amplify an outage.
- Capacity. Set the replica maximum from load tests, not hope, and alert when the deployment sits at its maximum for more than a few minutes.
- Version pinning. Record the exact model id, precision and shape with every evaluation result so a silent change shows up as a diff.
- Cost tracking. Divide GPU-seconds by tokens served each day; a rising cost per token is usually idle replicas or falling batch occupancy.
Failure modes
Cold starts on the user path. A deployment allowed to scale to zero makes the first morning request wait for model load. Keep a minimum replica.
Silent quality drift after FP8. Outputs change slightly; most tasks are fine, but extraction and arithmetic can regress. Gate the switch on your own evaluation set.
Bad predictions. Passing a prediction that the output will not follow slows requests. Monitor latency per request type and drop predictions where they do not help.
Truncated JSON. Constrained decoding cannot finish a document cut off by max_tokens. Size the limit and retry on length.
KV exhaustion. Long contexts on a replica with little free memory cut concurrency and push queueing delay into time to first token. The worked example shows how fast this happens.
Unvalidated deployments. Skipping --deployment-shape is the documented top cause of failed creation.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Serverless | No idle cost, instant start | Shared capacity, catalogue models only, per-token price |
| On-demand deployment | Predictable latency, any supported model | Pay for idle GPU-seconds, capacity planning |
| FP8 | Fits on fewer GPUs, more KV room, faster decode | Small numeric changes; must re-evaluate |
| Fast shape | Lower latency | Higher cost per token |
| Throughput shape | Lower cost per token | Higher per-request latency under load |
| Predicted Outputs | Large speed-ups on edits | Slower when the prediction is wrong |
| Hosted platform | No cluster to run | Less visibility into kernels; vendor dependency |
What to do next
- Build a versioned set of 100 to 500 real prompts with expected outputs or graders.
- Run it on serverless through the harness and record p50 and p95 time to first token, total latency and quality.
- Work out weights plus KV memory for your model and context lengths, then pick two candidate shapes and accelerators that leave room for your target concurrency.
- Create both deployments with an explicit
--deployment-shape, compare FP16 and FP8 on quality and latency, and keep the winner. - Try Predicted Outputs on every edit-style request type and keep it only where p50 latency falls.
- Enable
json_schemawhere you parse outputs, add alengthretry, and confirm reasoning models still behave. - Put a fallback router and a per-token cost alert in front of production traffic, and rerun the harness monthly. For the wider serving picture see LLM serving architecture.