DingTalk is Alibaba's workplace messaging platform, and for many teams in China it is where people actually read alerts, approve deployments and ask for help. Integrating with it looks trivial: paste a webhook URL into your monitoring tool and post JSON. It stops being trivial when an incident fires two hundred alerts in a minute and the robot goes silent for ten, when an engineer wants to type ack 4411 into the group and have it do something, or when the security team asks why a robot secret is committed to a repository.
This article covers the three integration paths and when each fits, signing and rate limits for custom robots, the enterprise application APIs with token caching, receiving messages through Stream mode or signed HTTP callbacks, and a worked alerting-and-acknowledgement flow. The API details were checked against DingTalk's developer documentation on 4 October 2026. DingTalk runs a China edition and an international edition with different host names, so every host in the code below comes from configuration: copy the URL your own console shows.
Three integration paths
There are three distinct ways to talk to DingTalk, and mixing them up is the most common design mistake.
| Path | Direction | Credential | Fits |
|---|---|---|---|
| Custom robot webhook | Send only, to one group | Webhook URL with access token, plus an optional signing secret | Alerts and notifications into a single group |
| Enterprise app robot via OpenAPI | Send to groups and one-to-one chats | App key and app secret, exchanged for an access token | Products that message many groups or users |
| Receiving messages | Inbound, when users message or @ the robot | Stream mode (outbound WebSocket) or HTTP callback with signature headers | ChatOps commands, assistants, approvals |
A custom robot is created inside a group and can only post to that group. It is ideal for alerting because it needs no application registration. An enterprise application robot belongs to an app registered in the developer console; it can be added to many groups, can message individuals, and can receive messages. Anything conversational needs the application route.
Architecture: one notifier service
Put a small notifier service between your systems and DingTalk instead of letting every tool post directly. The notifier owns the secrets, the rate limits and the message formatting, which means one place to fix each when they go wrong.
If you already run serverless functions, the notifier fits Alibaba Function Compute well for the outbound side, but the Stream client holds a long-lived connection and belongs on a process that stays up, such as a small container. The general patterns for durable delivery, retries and replay are in the webhook delivery platform article.
Signing custom robot messages
Custom robots offer three security settings: custom keywords, where a message is accepted only if it contains one of the configured words; signing; and an IP allowlist. Keywords are weak, since anyone with the URL can include the word. Use signing, and add the IP allowlist when your egress addresses are stable.
With signing enabled, every request carries two extra query parameters. timestamp is the current time in milliseconds. sign is an HMAC-SHA256, keyed with the robot secret, over the string timestamp, newline, secret; the digest is base64-encoded and then URL-encoded. A request without them fails with error 310000.
import base64, hashlib, hmac, time, urllib.parse
import requests
def signed_url(webhook_url: str, secret: str) -> str:
ts = str(round(time.time() * 1000))
string_to_sign = f"{ts}\n{secret}".encode("utf-8")
digest = hmac.new(secret.encode("utf-8"), string_to_sign, hashlib.sha256).digest()
sign = urllib.parse.quote_plus(base64.b64encode(digest))
return f"{webhook_url}×tamp={ts}&sign={sign}"
def send_text(webhook_url: str, secret: str, content: str, at_user_ids=()):
body = {"msgtype": "text", "text": {"content": content},
"at": {"atUserIds": list(at_user_ids)}}
r = requests.post(signed_url(webhook_url, secret), json=body, timeout=5)
r.raise_for_status()
result = r.json()
if result.get("errcode") != 0:
raise RuntimeError(f"dingtalk error {result.get('errcode')}: {result.get('errmsg')}")
return resultTwo details bite people. The HTTP status is 200 even when DingTalk rejects the message, so always check errcode in the body. And the timestamp is part of the signature, so a host with a badly skewed clock produces valid-looking requests that fail; run NTP on anything that signs. Store the webhook URL and the secret together in a secret manager, since the URL alone is enough to post if signing is ever disabled; the cloud secrets article covers rotation and injection.
Rate limits, deduplication and digests
DingTalk documents a limit of 20 messages per minute for each custom robot posting to a group. Exceed it and the robot is throttled for 10 minutes. That is the exact moment you most need it, at the start of an incident. The fix is to do the throttling yourself, below the limit, and turn bursts into digests.
import time, collections, threading
class DigestingSender:
# Send at most `rate` messages per 60 s; fold the overflow into one digest.
def __init__(self, send, rate=15):
self.send, self.rate = send, rate
self.sent = collections.deque()
self.pending = []
self.lock = threading.Lock()
def submit(self, key: str, line: str):
with self.lock:
now = time.monotonic()
while self.sent and now - self.sent[0] > 60:
self.sent.popleft()
if len(self.sent) < self.rate - 1: # keep one slot for the digest
self.sent.append(now)
self.send(line)
else:
self.pending.append((key, line))
def flush(self):
# Call every 60 s from a timer.
with self.lock:
if not self.pending:
return
keys = collections.Counter(k for k, _ in self.pending)
top = ", ".join(f"{k} x{n}" for k, n in keys.most_common(10))
self.send(f"{len(self.pending)} more alerts suppressed in the last minute: {top}")
self.sent.append(time.monotonic())
self.pending.clear()Set the budget below 20, because other tools may post through the same robot and because counting starts on DingTalk's clock, not yours. Deduplicate before the limiter too: the same alert firing every 30 seconds should update one message, not consume the budget. If one group genuinely needs more than this, that is a sign of alert design, not a quota problem.
Enterprise apps: tokens and group messages
An enterprise application exchanges its app key and app secret for an access token with POST /v1.0/oauth2/accessToken on the API host. The token lasts 7,200 seconds. Cache it and refresh a few minutes early; fetching a token per message wastes quota and adds latency. Group messages go to POST /v1.0/robot/groupMessages/send with the token in the x-acs-dingtalk-access-token header. The body names the robot (robotCode), the target group (openConversationId), a message template key such as sampleText or sampleMarkdown, and msgParam, which is a JSON string rather than a nested object, limited to 15,000 bytes.
import json, threading, time
import requests
API = "https://api.dingtalk.com" # from config; the international edition uses another host
class TokenCache:
def __init__(self, app_key, app_secret, margin=300):
self.key, self.secret, self.margin = app_key, app_secret, margin
self.token, self.expires_at = None, 0.0
self.lock = threading.Lock()
def get(self) -> str:
with self.lock:
if self.token and time.time() < self.expires_at - self.margin:
return self.token
r = requests.post(f"{API}/v1.0/oauth2/accessToken",
json={"appKey": self.key, "appSecret": self.secret}, timeout=5)
r.raise_for_status()
data = r.json()
self.token = data["accessToken"]
self.expires_at = time.time() + int(data.get("expireIn", 7200))
return self.token
def send_group_markdown(tokens: TokenCache, robot_code, conversation_id, title, markdown):
param = json.dumps({"title": title, "text": markdown}, ensure_ascii=False)
if len(param.encode("utf-8")) > 15_000:
raise ValueError("msgParam over 15,000 bytes; truncate or link to details")
r = requests.post(f"{API}/v1.0/robot/groupMessages/send",
headers={"x-acs-dingtalk-access-token": tokens.get()},
json={"robotCode": robot_code, "openConversationId": conversation_id,
"msgKey": "sampleMarkdown", "msgParam": param},
timeout=5)
r.raise_for_status()
return r.json()Rate-limit errors from this API come back with codes such as send.too.fast and send.byToken.tooFast. Treat them as retryable with backoff, unlike parameter errors, which a retry will never fix. Confirm the parameter names of any template other than text and markdown in the message-type reference before you use it.
Receiving messages: Stream mode and callbacks
To react when someone @ mentions the robot, you either expose an HTTPS callback or use Stream mode. With Stream mode your service opens an outbound WebSocket to DingTalk and events arrive over it, so you need no public endpoint, no TLS certificate for an inbound domain, and no firewall change. The official Python SDK is dingtalk-stream.
import dingtalk_stream
from dingtalk_stream import AckMessage
ALLOWED = {"ack", "status", "silence"}
class OpsBot(dingtalk_stream.ChatbotHandler):
async def process(self, callback: dingtalk_stream.CallbackMessage):
msg = dingtalk_stream.ChatbotMessage.from_dict(callback.data)
if msg.message_type != "text":
return AckMessage.STATUS_OK, "OK"
words = msg.text.content.strip().split()
if not words or words[0] not in ALLOWED:
self.reply_text("Commands: ack <id>, status <id>, silence <id> <minutes>", msg)
else:
reply = handle_command(words, user=msg.sender_staff_id) # your code
self.reply_text(reply, msg)
return AckMessage.STATUS_OK, "OK"
client = dingtalk_stream.DingTalkStreamClient(dingtalk_stream.Credential(CLIENT_ID, CLIENT_SECRET))
client.register_callback_handler(dingtalk_stream.chatbot.ChatbotMessage.TOPIC, OpsBot())
client.start_forever()If you use HTTP callbacks instead, every request carries timestamp and sign headers. Recompute it with the same key and string to sign as outbound signing, using the application secret, but keep the raw base64 digest without URL-encoding it. Compare it to the header in constant time, and reject requests whose timestamp is more than an hour from your clock, which is DingTalk's own tolerance; a tighter window is reasonable. The callback payload includes a sessionWebhook for replying and a sessionWebhookExpiredTime. The session webhook is short-lived, so reply promptly and use the OpenAPI for anything sent later. Never store it as a permanent channel.
Worked example: alerts and acknowledgements
A payments team routes production alerts to an on-call group. Prometheus Alertmanager posts to the notifier, which keys each alert by name and service, drops repeats within five minutes, and formats a markdown card with severity, a dashboard link and the alert id. During a database failover, 140 alerts arrive in 90 seconds. The limiter sends the first 14 individually, folds the overflow into a digest at the minute mark, then resumes as its window slides. In a simulation of this burst the group sees 27 individual messages and 2 digests, 29 instead of 140, never more than 14 in any 60-second window, so the robot is never throttled.
The on-call engineer types ack 4411 while mentioning the robot. The Stream client receives the event, the handler checks that the sender's staff id is on the current rota, and it calls the alerting system's acknowledge API with an idempotency key built from the message id, so DingTalk's redelivery cannot acknowledge twice. The reply goes back through the session webhook within a second. The pattern for the key is explained in idempotency architecture. The same handler can hand free-text questions to an LLM, for example through Model Studio, but give it read-only tools, because group messages are untrusted input.
Failure modes
- Silent rejections. HTTP 200 with a non-zero
errcodemeans nothing was delivered. Alert on the error code, not on the status. - Throttled mid-incident. Twenty a minute disappears fast. Limit and digest locally, and keep one robot per purpose so that a noisy CI pipeline cannot starve production alerts.
- Clock skew. Signatures include the timestamp; skewed hosts fail outbound and wrongly reject or accept inbound callbacks.
- Leaked webhook URLs. URLs end up in logs, screenshots and repositories. Enable signing, redact query strings in logs, and rotate by recreating the robot.
- Expired tokens and session webhooks. Cache the access token with a margin and refresh on an authentication error; do not reuse a session webhook after its expiry.
- Stream disconnects. The SDK reconnects, but a process that is down receives nothing. Run it under a supervisor, export a heartbeat metric, and alert when no events or pings have arrived for a while.
- Commands from anyone. Any member of a group can mention the robot. Authorize each command by staff id, not by group membership.
Operating the integration
Keep a small inventory: which robots exist, in which groups, owned by whom, with which secrets and allowlisted IPs. Export metrics for messages sent, digests sent, error codes by type, token refreshes and Stream reconnects. Test the full path weekly with a synthetic alert that must be acknowledged, so a broken integration is found on a quiet Tuesday rather than during an outage. When choosing between the custom robot and the application route, start with the custom robot for one-way alerts and move to an application once you need more than one group, direct messages or commands.
What to do next
- List every tool that posts to DingTalk today and route them through one notifier service.
- Turn on signing for every custom robot and move webhook URLs and secrets into a secret manager.
- Add a local limiter set below 20 messages per minute, with deduplication and per-minute digests.
- Check
errcodeon every response and alert on non-zero codes. - For commands, register an application, use Stream mode, and authorize each command by sender staff id.
- Make command handlers idempotent using the message id, and reply through the session webhook promptly.
- Run a weekly synthetic alert that must round-trip through the group and be acknowledged.