Microsoft Sentinel is Microsoft's cloud-native SIEM and SOAR service. Underneath the product name it is a Log Analytics workspace with security content on top: connectors that land logs in tables, analytics rules that run Kusto Query Language (KQL) on a schedule and raise alerts, incidents that group alerts for analysts, and automation rules and Logic Apps playbooks that act on them. Each cost, latency or detection problem belongs to one stage of that pipeline.
One platform change frames everything in this article. Microsoft Learn states that after March 31, 2027 Sentinel will no longer be supported in the Azure portal and will be available only in the Microsoft Defender portal; the date was moved there from July 1, 2026. Onboarding to the Defender portal changes who creates incidents and removes a couple of rule options, and those differences are called out where they matter. The limits quoted below were checked against Microsoft Learn on 2026-10-04. Re-check them before designing around them; no prices are quoted.
If you want the vendor-neutral SIEM pipeline first, read SIEM detection pipeline architecture. Posture management and workload protection belong to Microsoft Defender for Cloud, which feeds Sentinel alerts but is a different product.
The pipeline, end to end
Logs arrive through three kinds of entry point. Service connectors pull or receive data from Microsoft and third-party sources such as Entra ID sign-ins, Microsoft 365 audit logs, cloud provider logs and CEF or syslog appliances. The Azure Monitor Agent collects from machines you run. The Logs ingestion API accepts JSON from your own applications. Most modern paths pass through a data collection rule (DCR), which can filter rows, drop or rename columns and route the stream to a table.
Tables live in a storage tier. The analytics tier is the expensive, fast tier that analytics rules and interactive hunting run against; Microsoft Learn describes 90 days of interactive retention by default, extendable up to two years. The data lake tier keeps secondary security data cheaply for longer, and data whose analytics retention has ended remains reachable there. Rules query the analytics tier and write alerts, incidents collect alerts, and automation acts on incidents.
Three consequences follow. Detection quality is bounded by what you ingested into the analytics tier, so tiering is a security decision, not only a finance decision. Latency is the sum of source delay, ingestion delay and rule schedule, which is why the rule types below exist. And cost is set mostly at the DCR and tier stage, long before anyone writes a query.
Getting data in: connectors, DCRs and normalization
Choose connectors from the questions you need to answer, not from the gallery. Identity logs (Entra ID sign-ins and audit logs), endpoint and email alerts from Defender products, cloud control-plane logs and firewall or proxy logs cover most early detections. Each connector lands data in a known table, such as SigninLogs, AuditLogs, SecurityEvent or CommonSecurityLog, and analytics rule templates list the tables they need.
For your own applications, write to a custom table through the Logs ingestion API. You create the table (custom table names end in _CL), a data collection endpoint if your setup needs one, and a DCR that declares the incoming stream and a transformation. The application authenticates with Entra ID and must hold the Monitoring Metrics Publisher role on the DCR. The Python client library looks like this:
from azure.identity import DefaultAzureCredential
from azure.monitor.ingestion import LogsIngestionClient
client = LogsIngestionClient(endpoint=DCE_OR_DCR_ENDPOINT,
credential=DefaultAzureCredential())
events = [{
"TimeGenerated": "2026-10-04T01:12:09Z",
"Actor": "svc-billing",
"Action": "export_customer_data",
"RowCount": 48211,
"SourceIp": "10.20.4.17",
}]
# rule_id is the DCR's immutable id; stream_name is declared in the DCR
client.upload(rule_id=DCR_IMMUTABLE_ID, stream_name="Custom-AppAudit", logs=events)The DCR transformation is a KQL statement over a virtual table called source. Use it to drop what you will never query before you pay to store it, for example source | where Action != "heartbeat" | project-away DebugPayload. Filtering at ingestion is the single biggest cost lever you have. A dropped row can never be detected on.
Normalize before you detect. The Advanced Security Information Model (ASIM) provides parsers that present many sources in one schema; Microsoft recommends writing rules against ASIM parsers rather than native tables, so that a rule written for authentication events covers Entra ID, Windows and third-party sources alike. ASIM unifying parsers such as _Im_Authentication accept time and value filters as parameters.
Detections: scheduled and NRT rules
Sentinel has several analytics rule types. Microsoft security rules, which turned other Microsoft products' alerts into incidents, are disabled automatically when you onboard to the Defender portal. The two types you write yourself are scheduled rules and near-real-time (NRT) rules, and choosing between them is mostly a latency and volume decision.
| Property | Scheduled rule | NRT rule |
|---|---|---|
| How often it runs | you choose, 5 minutes to 14 days | every minute, fixed |
| Lookback | you choose, 5 minutes to 14 days; interval must not exceed it | the preceding minute of ingested events |
| Built-in delay | 5 minutes after the scheduled time | 2 minutes |
| Time column | event time, TimeGenerated | ingestion time |
| Alerts per run | up to 150 | up to 30 |
| Quantity | no per-rule-type cap quoted here | no more than 50 per customer |
| Fits | aggregation, baselines, joins over hours | single high-value events that need fast response |
A scheduled rule runs a KQL query every interval over a lookback window and fires when the result count crosses a threshold. The query must be between 1 and 10,000 characters and cannot contain search * or union *; push long logic into saved functions. Because data arrives late, scheduled rules run five minutes after their nominal time, and a source with longer ingestion delay than that can still slip between windows. The fix is to make the lookback longer than the interval and deduplicate on an ingestion-time filter, or to compare ingestion_time() rather than TimeGenerated for the newest slice.
NRT rules avoid most of that by querying on ingestion time, which is why they run on a two-minute delay. Spend the 50 allowed per customer on events where minutes matter: a break-glass account signing in, a new federation trust, a mass mailbox export.
Both types share alert enhancement. Entity mapping tells Sentinel which columns hold accounts, hosts, IPs and other entities, and it drives incident grouping, investigation graphs and automation conditions. Event grouping chooses between one alert per run and one alert per result row, and alert grouping merges alerts into incidents within a time frame (5 hours by default, adjustable from 5 minutes to 7 days) when entities match. In the Defender portal, the Defender XDR correlation engine names incidents and can override customized alert names, and the option to reopen closed incidents is not available.
Worked example: a password-spray rule
Take one detection from idea to incident. Password spraying tries a few common passwords against many accounts from one address, staying below lockout thresholds. In Entra ID sign-in logs that is many distinct accounts failing from the same IP in a short window, sometimes followed by a success.
let window = 10m;
let failures =
SigninLogs
| where ResultType in ("50126", "50053") // bad password, locked/blocked
| summarize FailedAccounts = dcount(UserPrincipalName),
Attempts = count(),
SampleAccounts = make_set(UserPrincipalName, 20)
by IPAddress, bin(TimeGenerated, window)
| where FailedAccounts >= 15;
failures
| join kind=leftouter (
SigninLogs
| where ResultType == "0" // string, not number
| where IPAddress in ((failures | project IPAddress))
| project SuccessTime = TimeGenerated, IPAddress, SuccessUser = UserPrincipalName
) on IPAddress
| extend SuccessUser = iff(SuccessTime between (TimeGenerated .. (TimeGenerated + 1h)),
SuccessUser, "")
| summarize FailedAccounts = max(FailedAccounts), Attempts = max(Attempts),
Successes = make_set_if(SuccessUser, isnotempty(SuccessUser), 10),
SampleAccounts = take_any(SampleAccounts)
by IPAddress, TimeGeneratedConfigure it as a scheduled rule that runs every 30 minutes with a 1-hour lookback, so each window is seen twice and late rows are not missed. Map IPAddress to an IP entity. Set event grouping to one alert per result row and alert grouping to merge alerts whose IP entity matches within 24 hours, so a spray that lasts all night becomes one incident instead of forty.
Now run the numbers for an illustrative tenant (the figures are invented to show the arithmetic). Over a week, the query returns 11 rows. Seven come from one corporate egress IP where a misconfigured mail client retried expired passwords; four come from two hosting-provider addresses, each with one success. Without tuning, that is 11 alerts and, with the grouping above, three incidents, of which one is noise. Excluding known egress ranges with a watchlist lookup removes the noise incident, and adding the Successes detail lets the analyst start from the one account that matters. The ResultType codes 50126 and 50053 are standard Entra ID sign-in error codes; check that your tenant logs them as you expect before trusting the threshold.
Before enabling, check the wizard's results simulation, which replays the rule's last 50 scheduled runs.
Incidents, automation rules and playbooks
Incidents are the unit of work. Automation rules are lightweight, ordered rules that trigger when an incident is created or updated or when an alert is created, check conditions such as rule name, severity or entity values, and then act: change severity or status, assign an owner, add tags or tasks, or run a playbook.
Playbooks are Logic Apps workflows. They do the work that needs an external call: enrich an IP from threat intelligence, post to a chat channel, open a ticket, disable a user or isolate a device through another API. Give each playbook a managed identity with the narrowest role it needs; a playbook that can disable any account is a privileged identity and deserves the same review as a human admin.
Older deployments may still list playbooks under alert automation (classic), which Microsoft Learn listed for deprecation in March 2026; recreate any that remain as automation rules on the alert-created trigger.
Keep destructive actions behind a human or behind high-confidence conditions. Automatically disabling an account on every spray incident turns an attacker's cheap spray into your outage. Automate enrichment and routing everywhere, containment only for rules with a near-zero measured false-positive rate.
Detections as code, with tests
Treat analytics rules like code. Export rules as templates, keep them in a repository with their KQL in separate files, review changes, and deploy through a pipeline or Sentinel's repositories feature. The part most teams skip is testing. Every rule should have a regression query that is run against a known window and a known expected count, so a schema change in a source or a parser update shows up as a failing check instead of a silent gap. The Azure Monitor query client makes this a few lines:
from datetime import timedelta
from azure.identity import DefaultAzureCredential
from azure.monitor.query import LogsQueryClient, LogsQueryStatus
client = LogsQueryClient(DefaultAzureCredential())
def rule_hits(kql: str, days: int = 7) -> int:
resp = client.query_workspace(WORKSPACE_ID, kql, timespan=timedelta(days=days))
if resp.status != LogsQueryStatus.SUCCESS:
raise RuntimeError(f"partial or failed query: {resp}")
return len(resp.tables[0].rows)
hits = rule_hits(open("rules/password_spray.kql").read())
assert 0 < hits < 50, f"spray rule returned {hits} rows over 7 days"Pair this with a heartbeat check per source: a scheduled query that alerts when a table that normally receives data every few minutes has received nothing for an hour. A detection over an empty table is the most common way a SIEM fails, and it never raises an alert by itself.
Failure modes
The failure modes that matter in practice:
- Silent source loss. A connector's credential expires or an agent stops, the table goes quiet, and every rule over it returns zero. Monitor ingestion per table, not just per workspace.
- Late data between windows. A source with long ingestion delay lands rows after the scheduled rule's window has passed. Overlap lookback with interval, or use ingestion-time logic.
- Schema drift. A vendor renames a field and
projectfails or a filter matches nothing. Usecolumn_ifexists()and test rules on a schedule. - Alert floods. One noisy rule creates thousands of incidents and analysts start bulk-closing, which hides the real one. Cap with thresholds, grouping and suppression (up to 24 hours) and fix the rule.
- Tiering blind spots. Logs routed to cheaper storage are not seen by analytics rules. Before moving a table, list the rules that query it.
- Over-privileged playbooks. A compromised or buggy playbook acts at machine speed. Scope managed identities and log every action.
- Portal migration surprises. After moving to the Defender portal, incident naming and creation follow Defender XDR. Re-check automation conditions that match on incident titles.
Trade-offs
Ingest everything versus ingest deliberately: broad ingestion maximizes hunting options and cost, while deliberate ingestion keeps the analytics tier focused on data that rules use. Revisit the split quarterly using actual query usage.
Scheduled versus NRT: NRT buys minutes of latency at the cost of a small, capped budget of rules and simpler queries. Scheduled rules can aggregate over hours and join across tables. Use NRT for single decisive events and scheduled rules for anything statistical.
Automation versus control: automated containment shortens response time and also widens the blast radius of a false positive. Automate the cheap and reversible actions first.
For identity data, which is where many Sentinel detections start, see Entra ID in depth; for the minimum control set that these detections should be watching, see Cloud Security Baseline.
What to do next
- List the five questions your SOC most needs to answer and the tables that answer them; connect those sources first.
- Write a DCR transformation for each high-volume source that drops fields and rows you will never query, and record what was dropped.
- Decide per table between the analytics tier and the data lake tier, and list the rules that depend on each analytics table.
- Enable a small set of templates, run results simulation on each, and tune thresholds and exclusions before analysts see output.
- Add a per-table heartbeat rule and a regression query per detection that runs in CI.
- Move any classic alert-trigger playbooks to automation rules and scope every playbook's managed identity.
- Plan the move to the Defender portal before March 31, 2027 and re-test automation conditions after onboarding.