Chronicle began as a security analytics platform built on Google infrastructure and is now sold as Google Security Operations, usually shortened to Google SecOps, which combines the SIEM with a SOAR product for cases and playbooks. The documentation still lives under the chronicle path and many teams still call it Chronicle. Its design choices are distinctive: logs are normalised into one schema, the Unified Data Model; detections are written in a purpose-built language, YARA-L 2.0; and contextual data about users, assets and threat indicators sits in an entity graph that rules can join against.
This article explains how data flows through the platform, how to choose ingestion paths, what makes UDM parsing succeed or fail, how to write and test YARA-L rules including multi-event correlation, and how to run the whole thing as an engineering pipeline. For the general architecture of any SIEM see the SIEM detection pipeline; for posture findings inside Google Cloud see Security Command Center.
How data flows
Every log takes the same path. It arrives through an ingestion method, is stored in raw form so it remains searchable even if parsing fails, and is passed to a parser selected by its log type. The parser emits UDM events: structured records whose fields describe who did what to which resource. Separately, context sources such as directory data, asset inventories and threat intelligence populate the entity graph. Rules and searches run over UDM events, optionally joined to entity context, and matches become detections that can raise alerts and open cases in the SOAR layer.
By default Google retains twelve months of data in a SecOps account, and the retention period can be extended up to five years as part of the purchase agreement. That year of hot, searchable history is a large part of the value: a retrohunt can run a new rule over months of past events, which most traditional SIEM deployments cannot afford to keep online.
Architecture
Choosing ingestion paths
| Path | Use it for | Operational notes |
|---|---|---|
| Forwarder | On-prem syslog, files, Windows events | Runs on your hosts; monitor its own health and buffer |
| Bindplane agent | Managed collection from Windows and Linux servers | Telemetry pipeline that can filter and reshape before sending |
| Feeds | Logs already in Cloud Storage or S3, push webhooks, predefined SaaS and EDR API integrations | No agent; credentials and bucket permissions are the usual failure |
| Direct Google Cloud ingestion | Cloud Audit Logs and other Google Cloud telemetry | Configure which log types flow; check volume before enabling everything |
| Ingestion API | Custom applications and pipelines | You own batching, retries and the log type label |
Choose by where the logs already are. If a vendor already lands logs in a bucket, a feed is simpler and more reliable than installing an agent. For Google Cloud itself, start with Cloud Audit Logs (admin activity and data access for sensitive services) and VPC flow or firewall logs for the networks that matter; the routing and filtering side is explained in Cloud Logging. Whatever the path, every source needs the correct log type label, because that label selects the parser.
The Unified Data Model
UDM describes an event with nouns: the principal that acted, the target that was acted on, the src and intermediary where relevant, the observer that reported it, plus metadata such as event type, vendor and product, and a security_result block for actions and verdicts. A login becomes a USER_LOGIN event with the user in target.user and the client address in principal.ip; a blocked API call carries security_result.action set to BLOCK.
metadata.event_type = "USER_LOGIN"
metadata.log_type = "OKTA"
metadata.product_event_type = "user.session.start"
principal.ip = "203.0.113.24"
principal.ip_geo_artifact.location.country_or_region = "Netherlands"
target.user.userid = "a.rivera"
security_result.action = "ALLOW"Google provides default parsers for a long list of products, and parser extensions let you add or override field mappings without forking the whole parser. Most detection gaps trace back here: a vendor changes its log format, a field stops mapping, and every rule that depends on it goes quiet without an error. Treat parser coverage as a measured quantity: for each critical log type, track the share of raw logs that produce UDM events and the share of events with the fields your rules need.
Writing YARA-L rules
A YARA-L 2.0 rule has named sections in a fixed order: meta for descriptive fields, events for the conditions on event variables, an optional match section that groups by placeholder values over a time window, an optional outcome section that computes values such as a risk score, the condition that decides when the rule fires, and optional options. Event variables start with a dollar sign; placeholders bind a field value so that events can be joined on it. The first rule is a straightforward threshold, modelled on Google's public community rules.
rule gcp_permission_denied_burst {
meta:
author = "secops-team"
description = "Many PermissionDenied audit events from one principal across services"
severity = "Low"
events:
$gcp.metadata.log_type = "GCP_CLOUDAUDIT"
$gcp.security_result.action = "BLOCK"
$gcp.principal.user.userid = $user_id
$gcp.target.application = $service
match:
$user_id over 1h
outcome:
$risk_score = max(35)
$event_count = count_distinct($gcp.metadata.id)
$services = array_distinct($gcp.target.application)
$source_ips = array_distinct($gcp.principal.ip)
condition:
#gcp >= 10 and #service > 2
}The match section groups events by user over a one-hour window, and the condition uses counts: #gcp is the number of matching events and #service the number of distinct services. A burst of denials across several services is a classic sign of a stolen credential being used to explore what it can reach. The outcome values travel with the detection and populate the alert, so analysts see the services and addresses without running a search.
Multi-event rules join different kinds of events on a shared placeholder. The next rule looks for many failed logins followed by a success for the same user, and orders them with a timestamp comparison.
rule failed_logins_then_success {
meta:
description = "Ten or more failures then a success for one user within 30 minutes"
severity = "Medium"
events:
$fail.metadata.event_type = "USER_LOGIN"
$fail.security_result.action = "BLOCK"
$fail.target.user.userid = $user
$ok.metadata.event_type = "USER_LOGIN"
$ok.security_result.action = "ALLOW"
$ok.target.user.userid = $user
$fail.metadata.event_timestamp.seconds < $ok.metadata.event_timestamp.seconds
match:
$user over 30m
outcome:
$risk_score = max(60)
$failures = count_distinct($fail.metadata.id)
$fail_ips = array_distinct($fail.principal.ip)
$success_ips = array_distinct($ok.principal.ip)
condition:
#fail >= 10 and $ok
}The third pattern joins events to the entity graph. Threat intelligence ingested as entity context can be matched against event fields, so a connection to a known bad address fires without a hand-maintained list inside the rule.
rule outbound_to_intel_ip {
meta:
description = "Network connection to an IP present in ingested threat intelligence"
severity = "High"
events:
$conn.metadata.event_type = "NETWORK_CONNECTION"
$conn.target.ip = $ip
$ioc.graph.metadata.entity_type = "IP_ADDRESS"
$ioc.graph.metadata.source_type = "ENTITY_CONTEXT"
$ioc.graph.entity.ip = $ip
match:
$ip over 5m
outcome:
$risk_score = max(80)
$hosts = array_distinct($conn.principal.hostname)
condition:
$conn and $ioc
}Field names in the entity examples depend on how your intelligence source is parsed, so confirm them against real entity records before relying on the rule. Match windows have documented limits and rules run at configurable frequencies; check the current documentation for both rather than assuming a value.
Detection as code
Rules are code and deserve the same pipeline as code. Keep them in a repository, review changes, and test each rule in three ways before it can raise alerts: a syntax and field check, a set of sample UDM events that must and must not match, and a test run or retrohunt over recent history to measure how many detections it would have produced. Google publishes a community rules repository with this layout and a Chronicle API for managing rules; the harness below is a local sketch of the first two checks, independent of any particular API version.
import re, json, pathlib
KNOWN_FIELDS = set(json.load(open("udm_fields_in_use.json"))) # exported from real UDM events
def check_rule(path):
src = pathlib.Path(path).read_text()
problems = []
for field in re.findall(r"\$\w+\.((?:[a-z_]+\.)*[a-z_]+)", src):
if field.split(".")[0] in {"graph"}:
continue # entity fields checked separately
if field not in KNOWN_FIELDS:
problems.append(f"field never seen in our data: {field}")
if "severity" not in src:
problems.append("meta.severity missing")
if "$risk_score" not in src:
problems.append("no risk score in outcome")
return problems
for rule in pathlib.Path("rules").glob("*.yaral"):
issues = check_rule(rule)
if issues:
raise SystemExit(f"{rule.name}: {issues}")The field check matters more than it looks. A typo or a field your parsers never populate produces a rule that compiles, runs and never fires, which is indistinguishable from a quiet week. Promote rules from test to alerting only after the retrohunt volume is acceptable, and record the expected weekly detection count so a sudden drop to zero is itself an alert.
Worked example: Google Cloud audit logs
A platform team enables Cloud Audit Logs for IAM, Cloud Storage and Secret Manager, and turns on direct ingestion into SecOps. A week of data shows that the permission-denied burst rule, at a threshold of five events, would fire forty times a day, mostly for a deployment service account that probes for optional permissions on start-up. They raise the count to ten, require at least three services, as in the rule shown above, and add one extra events line excluding that service account. The retrohunt drops to two detections a week, one of which turns out to be a developer key leaked in a public repository and used from an unfamiliar country. The case opens in SOAR with the services and addresses already listed, and the playbook disables the key and notifies the owner. Audit events can also be exported through Pub/Sub to other consumers, but the detection lives in one place.
Operating the platform
- Silent sources. Alert when any critical log type stops arriving or drops sharply in volume; a missing source is the most common reason a real attack goes unseen.
- Parser drift. Track parse success and required-field coverage per log type, and review vendor release notes for format changes.
- Rule health. Track detections per rule per week; investigate rules that drop to zero or spike.
- Tuning. Use reference lists for known-good accounts and assets, and record why each exclusion exists and who approved it.
- Volume and cost. Ingestion volume drives cost under most licensing models; filter verbose debug logs at the forwarder or Bindplane rather than ingesting and ignoring them.
- Access. Restrict who can edit rules and parsers, and log those changes like any production deploy.
Failure modes and trade-offs
Failure modes follow from the design. Wrong log type labels send data to the wrong parser, producing events with empty fields. Parser extensions that override mappings can break default content that relied on the original fields. Multi-event rules with broad placeholders, such as joining on a common IP address behind a proxy, create enormous match groups and noisy alerts. Entity context that is stale, for example a directory export that stopped updating, makes rules that depend on it quietly wrong. Curated detections from Google reduce the rule-writing burden but still depend on your data being present and parsed.
The trade-offs are mostly about control. A single normalised schema makes rules portable across vendors, but you depend on parser quality you do not fully own. Long hot retention makes hunting cheap, while ingestion choices still drive cost. YARA-L is concise and fast for correlation but is another language for the team to learn, and rules written for it are not portable to other SIEMs without translation. Teams with mature Sigma or KQL content should plan the migration effort explicitly; the threat modeling guide helps decide which detections to rebuild first.
What to do next
- Inventory your critical log sources and pick an ingestion path for each by where the logs already live.
- Verify log type labels and measure parse success and required-field coverage for every critical source.
- Set up silent-source and volume-drop alerts before writing any detection rules.
- Put rules in a repository with field checks, sample-event tests and retrohunt volume reviews.
- Start with a few high-value rules: credential misuse, privilege changes, audit logging disabled, and connections to known bad infrastructure.
- Load threat intelligence and asset context into the entity graph and confirm the field names.
- Review detections per rule weekly and tune with documented reference-list exclusions.