Model Studio is Alibaba Cloud's managed service for generative models. It hosts the commercial Qwen models, open-weight Qwen releases, multimodal variants such as Qwen-VL and Qwen-Omni, coding models, and a selection of third-party models, behind an API that was historically branded DashScope and still uses that name in its SDK and API key variable. The console also offers application building and model customisation; this article is about the part most engineers integrate first, the inference API.
The goal is to leave you able to put Qwen behind a production service: pick a region and understand what it binds you to, make a first call through the OpenAI-compatible endpoint, control thinking mode deliberately, run a tool-calling loop, cut cost with the context cache and the Batch API, and recognise the failures before your users do. Model names and prices change often on this platform, so the code uses the stable qwen-plus alias and the article tells you where to pin versions rather than quoting numbers that will be stale next quarter.
Regions, workspaces and keys
Model Studio is deployed per region, and the region is the first decision because it binds three things together: where your prompts and outputs are processed, which models are available, and which endpoint and key you use. The documentation lists deployments in Singapore, US (Virginia) and China (Beijing), among others. An API key is created inside a workspace in one region; use it with that region's endpoint, because a key sent to another region's endpoint fails authentication.
Workspaces group keys, permissions and usage, so one workspace per environment or product makes usage reports map onto teams. They do not isolate rate limits: the documentation states that limits, measured in requests and tokens per minute with separate thresholds per model, are counted across every workspace, RAM user and key under the Alibaba Cloud account. A runaway staging job can therefore throttle production calls to the same model, so cap batch and test traffic on the client side. Keep the key in a secret store and inject it as DASHSCOPE_API_KEY; both the OpenAI SDK examples and the native dashscope SDK read it from there.
The OpenAI-compatible endpoint
Model Studio exposes an OpenAI-compatible chat completions API under a /compatible-mode/v1 base URL, so the official OpenAI client libraries work by changing the key and base URL. The US endpoint is https://dashscope-us.aliyuncs.com/compatible-mode/v1; the Singapore and Beijing endpoints in the current documentation are workspace-scoped hosts of the form https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1. Older examples use a shared host instead, so copy the URL your console shows rather than one from a blog post.
Compatibility means the common fields behave as you expect: messages, temperature, top_p, max_tokens, seed, stream, tools. Qwen-specific options such as thinking control are not OpenAI fields, so the Python client needs them in extra_body. The first call below turns thinking off explicitly for a short classification task; the default differs between models, and relying on it means a model upgrade can silently start generating, and billing, thinking tokens.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DASHSCOPE_API_KEY"],
# copy the base URL for YOUR region and workspace from the console, e.g. US (Virginia):
base_url="https://dashscope-us.aliyuncs.com/compatible-mode/v1",
timeout=60, max_retries=0, # retries are handled by our own policy below
)
resp = client.chat.completions.create(
model="qwen-plus", # alias; pin a dated snapshot for production
messages=[
{"role": "system", "content": "You classify support tickets. Reply with one label."},
{"role": "user", "content": "My invoice shows a charge twice this month."},
],
temperature=0.2,
max_tokens=20,
extra_body={"enable_thinking": False}, # not an OpenAI field: pass it explicitly
)
print(resp.choices[0].message.content, resp.usage)The alias qwen-plus moves forward as new snapshots ship. That is convenient in development and a risk in production, where a quiet model change can shift label distributions. Pin a dated model ID from the models page for production traffic, and move the pin deliberately after running your evaluation set against the new version.
Thinking mode
Recent Qwen models can reason before answering. With enable_thinking set to true the model produces a chain of reasoning first and then the reply; thinking_budget caps how many tokens the reasoning may use, after which the model moves straight to its answer. The reasoning comes back in a separate reasoning_content field, alongside the usual content, so you can log it, measure it, or drop it without parsing text.
Two operational details matter. First, some open-source Qwen models served on the platform support only streaming output, and a non-streaming call to them returns an error, so write your client around streaming from the start. Second, reasoning tokens are output tokens: they are billed and they count toward latency. Budget them per use case. Classification, extraction and routing rarely benefit from thinking; multi-step planning, maths and code review often do. Treat the budget as a tuning knob and measure accuracy against it.
stream = client.chat.completions.create(
model="qwen-plus",
messages=[{"role": "user", "content": "Plan a zero-downtime Postgres major upgrade."}],
stream=True,
stream_options={"include_usage": True}, # final chunk carries token usage
extra_body={"enable_thinking": True, "thinking_budget": 2048},
)
reasoning, answer, usage = [], [], None
for chunk in stream:
if chunk.usage: # the usage-only last chunk
usage = chunk.usage
if not chunk.choices:
continue
delta = chunk.choices[0].delta
if getattr(delta, "reasoning_content", None): # thinking arrives in its own field
reasoning.append(delta.reasoning_content)
if delta.content:
answer.append(delta.content)
log_reasoning_privately("".join(reasoning)) # never show raw reasoning to end users
print("".join(answer)); print(usage)Set stream_options with include_usage so the final chunk reports token usage; without it a streamed call gives you no usage record, and cost attribution breaks.
Tool calling
Tool calling uses the OpenAI tools schema: you describe functions with JSON Schema, the model returns tool_calls with a name and a JSON argument string, you execute them and send the results back as tool messages. The loop is the same as with any compatible provider, and the same rules apply. Validate arguments before executing, because the model can produce malformed or out-of-range values. Cap the number of rounds so a confused model cannot loop forever. Make side-effecting tools idempotent, because a timeout followed by a retry can execute the same call twice. For schema design and reliability patterns see structured output with LLMs.
tools = [{
"type": "function",
"function": {
"name": "get_order",
"description": "Look up an order by id",
"parameters": {"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"]},
},
}]
messages = [{"role": "user", "content": "Where is order A-1042?"}]
for _ in range(4): # hard cap on tool rounds
msg = client.chat.completions.create(model="qwen-plus", messages=messages,
tools=tools).choices[0].message
if not msg.tool_calls:
break
messages.append(msg)
for call in msg.tool_calls:
args = json.loads(call.function.arguments) # validate before executing
result = TOOLS[call.function.name](**args)
messages.append({"role": "tool", "tool_call_id": call.id,
"content": json.dumps(result)})
print(msg.content)
Context cache
Many production prompts share a long prefix: a policy manual, a tool catalogue, few-shot examples. Model Studio can cache the processed prefix so later requests skip recomputing it, which lowers both cost and time to first token. There are two modes. Implicit cache is on automatically: the platform detects repeated prefixes, but a hit is not guaranteed. Explicit cache is requested with a cache_control marker of type ephemeral on a content block, gives deterministic hits within a five-minute validity window that resets on every hit, and allows up to four markers per request.
Both modes need a prefix of at least 1,024 tokens. The documentation gives the pricing as typically 125% of the normal input price to create an explicit cache entry and typically 10% on a hit, and typically 20% for an implicit hit with no creation surcharge; check your model's price page for the exact figures. Responses report cached_tokens, and explicit creation shows up as cache_creation_input_tokens, so you can measure the hit rate instead of assuming it.
The prompt layout decides whether caching works at all. Put everything stable first and in byte-identical form, and everything that varies, such as the user's message, retrieved documents and timestamps, after the marker. A timestamp in the system prompt defeats the cache on every call. For general caching strategy see prompt caching.
messages = [
{"role": "system", "content": [
{"type": "text",
"text": POLICY_MANUAL, # ~6,000 tokens, identical on every call
"cache_control": {"type": "ephemeral"}}, # explicit cache marker
]},
{"role": "user", "content": ticket_text}, # the part that changes goes last
]
resp = client.chat.completions.create(model="qwen-plus", messages=messages)
details = resp.usage.prompt_tokens_details
print("cached:", getattr(details, "cached_tokens", 0))
The Batch API
Work that does not need an answer in seconds, such as backfills, evaluation runs, nightly enrichment and bulk classification, should go through the Batch API. You upload a JSONL file where each line is one request with a custom_id, a method, the /v1/chat/completions URL and a body; create a batch with a completion window between 24 hours and 336 hours; poll; and download the results file and the errors file. The documentation prices batch calls at 50% of real-time calls, and limits each file to 50,000 requests, 500 MB, and 6 MB per line.
Results come back keyed by custom_id, not in input order, so the ID must let you join results back to your records. Design for partial failure: some lines land in the error file and need a retry batch, and a job can expire at the end of its window with work unfinished. For the general pattern see batch inference.
# requests.jsonl - one request per line, each with a unique custom_id
{"custom_id": "t-000001", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "qwen-plus", "messages": [{"role": "user", "content": "..."}]}}
# submit, poll, download
f = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
job = client.batches.create(input_file_id=f.id, endpoint="/v1/chat/completions",
completion_window="24h")
while (job := client.batches.retrieve(job.id)).status not in ("completed", "failed",
"expired", "cancelled"):
time.sleep(60)
if job.output_file_id:
open("results.jsonl", "wb").write(client.files.content(job.output_file_id).content)
if job.error_file_id:
open("errors.jsonl", "wb").write(client.files.content(job.error_file_id).content)
Worked example: triaging a million tickets
A support team wants every historical ticket labelled with a category and a sentiment: one million tickets, averaging 400 tokens each, with a 6,000-token instruction and example prefix, and about 20 output tokens per ticket. Measure cost in units of the normal input-token price, which keeps the arithmetic valid whatever the current price is.
Sent naively in real time, each call carries 6,400 input tokens, so the job is 6.4 billion input tokens, about 94% of it the repeated prefix. With explicit cache hits on the prefix at the documented typical 10%, each call costs 600 + 400 = 1,000 units instead of 6,400, roughly 84% less input spend, paid for by occasional cache writes at 125% whenever the five-minute window lapses. Sent as a batch instead, the whole job costs half the real-time price and no service has to keep up with interactive latency.
Whether cache discounts also apply inside batch jobs is not something the documentation pages cited here state, so do not stack them in a forecast; run a 1,000-ticket pilot both ways and read the usage fields. Keep thinking off for this workload, since a 2,048-token thinking budget would be a hundred times the 20-token answer.
Choosing a model tier
Qwen's commercial line is organised in tiers: Max for the hardest tasks, Plus as the balanced default, and Flash for low-latency, low-cost volume. Start with Plus, build an evaluation set of a few hundred real inputs with expected outputs, then try moving down a tier for easy traffic and up for hard traffic. Routing by difficulty usually beats a single model; see LLM routing. Open-weight Qwen models are also an option if you need to self-host for data control, at the price of running the serving stack yourself.
Failure modes
| Symptom | Likely cause | First response |
|---|---|---|
| 401 or invalid key on a key that works elsewhere | Key from a different region or workspace | Match the key's region to the base URL |
| Error on a non-streaming call to an open-source model | Model supports streaming output only | Use stream=True for that model |
| Empty or cut-off answers with thinking on | Reasoning consumed the output budget | Lower thinking_budget or raise max_tokens |
| Cost jumped after no code change | Alias moved to a new snapshot, or thinking default changed | Pin model IDs; set enable_thinking explicitly |
| cached_tokens always zero | Prefix under 1,024 tokens or not byte-identical | Move variable content after the cache marker |
| 429 responses at peak | Account-wide RPM or TPM limit for that model | Client-side token bucket and backoff with jitter; throttle batch and test traffic |
| Batch finished with missing rows | Lines in the error file, or window expired | Retry from error_file_id; size jobs to the window |
Retry only on rate limits, timeouts and server errors, with exponential backoff and jitter, and never on validation errors. Log the request ID, model ID, token usage and cached tokens for every call; those four fields answer most cost and quality questions later. For how this fits alongside the other clouds' model services see cloud AI platforms.
What to do next
- Choose the region from your data-residency requirement, create a workspace per environment, and store the key as
DASHSCOPE_API_KEY. - Copy the base URL from the console and make one call with the OpenAI SDK, setting
enable_thinkingexplicitly. - Switch the client to streaming with
include_usageand log usage, request ID and model ID on every call. - Restructure long prompts so the stable prefix comes first, add a cache marker, and confirm
cached_tokensis non-zero. - Move every non-interactive workload to the Batch API with join-safe
custom_idvalues and an error-file retry path. - Build an evaluation set, pin a dated model ID for production, and re-run the set before moving the pin.