All 98 articles, sorted alphabetically
Google App Engine in Depth: Services, Versions, Instances, Scaling Types and When to Move to Cloud Run
How Google App Engine actually runs your code: the application, service, version and instance model, the standard and flexible environments, automatic…
Read article →Artifact Registry, in depth: repository modes, digest-pinned delivery, cleanup policies and supply-chain controls on Google Cloud
A practical guide to Google Cloud Artifact Registry: the resource model, standard, remote and virtual repositories, client authentication, a Cloud Bui…
Read article →BeyondCorp Enterprise, in depth: device signals, access levels in YAML and CEL, a tier ladder, enforcement across IAP, Workspace and Chrome, and a rollout that does not lock people out
How Google's BeyondCorp Enterprise zero-trust stack, now documented as Chrome Enterprise Premium, works in practice: the Endpoint Verification si…
Read article →BigQuery ML, in depth: model types and where they train, data splits, TRANSFORM, a churn model end to end, forecasting, remote models and operations
How BigQuery ML trains and serves models with SQL: built-in, Vertex-trained, imported and remote models; AUTO_SPLIT behaviour and time-based splits; T…
Read article →BigQuery Partitioning + Clustering, in depth: pruning rules, choosing columns, cost arithmetic and changing the layout of a live table
How to lay out BigQuery tables so queries read less: the bytes-scanned cost model, time-unit, ingestion-time and integer-range partitioning, clusterin…
Read article →BigQuery Slots and Reservations, in depth: editions, baseline, autoscaling, idle sharing, assignments and sizing from job history
How BigQuery slot capacity works in practice: on-demand versus editions, commitments, reservations, assignment resolution, autoscaling billing, idle s…
Read article →BigQuery Streaming Inserts, in depth: the REST method versus Storage Write API streams, offsets for exactly-once, pending commits and CDC
Stream rows into BigQuery correctly: the renamed REST insertAll method and its best-effort insertId dedup, Storage Write API default, committed, pendi…
Read article →Bigtable vs Spanner vs BigQuery, in depth: access patterns, storage engines, one workload modelled three ways and how to choose
A practical comparison of Bigtable, Spanner and BigQuery: how each stores and serves data, a fleet telemetry workload modelled in all three with code,…
Read article →Chronicle, in depth: Google Security Operations ingestion paths, UDM parsing, YARA-L detection engineering and running the SIEM as a pipeline
A practical guide to Chronicle, now Google SecOps: how logs flow from forwarders, Bindplane, feeds and APIs into UDM and the entity graph, writing sin…
Read article →Google Cloud Deploy, in depth: releases, rollouts, canaries, automation and deploy policies
How Google Cloud Deploy works: the release and rollout model, render-once promotion, a complete Cloud Run pipeline with a production canary, verify jo…
Read article →Cloud DLP + Sensitive Data Protection, in depth: infoTypes and likelihood, inspection jobs, discovery profiles and reversible tokenisation
How Google Cloud Sensitive Data Protection (formerly Cloud DLP) detects and transforms sensitive data: infoTypes, likelihood, custom detectors and rul…
Read article →Cloud DNS, in depth
Operating Google Cloud DNS: records as code, TTL cutover planning, response policies, Cloud DNS for GKE scopes, query logging and metrics, and a step-…
Read article →Cloud Interconnect, in depth: Dedicated, Partner and Cross-Cloud links, VLAN attachments, BGP and BFD, SLA topologies and capacity planning
A practical guide to Google Cloud Interconnect: connections, VLAN attachments and Cloud Router BGP, Dedicated vs Partner vs Cross-Cloud, MTU and per-f…
Read article →Cloud NAT, in depth: port allocation, timeouts, sizing NAT IPs, GKE egress and diagnosing dropped packets
How Google Cloud NAT works and how to run it: the software-defined data plane and Cloud Router control plane, 64,512 ports per IP and per-destination …
Read article →Cloud Profiler, in depth: how continuous profiling works, per-language setup, reading flame graphs and a worked regression hunt
Google Cloud Profiler from first principles: agent and backend scheduling, profile types per language, documented overhead and retention, Python/Go/Ja…
Read article →Cloud Run, in depth: building and operating a production service and job, from the container contract and shutdown to billing modes, cold starts, private networking, identity and rollouts
A practitioner's guide to running production workloads on Cloud Run: what the container contract requires, graceful shutdown code, request-based …
Read article →Cloud Source Repositories, in depth: operating an end-of-sale service, IAM, push notifications, triggers and migrating to Secure Source Manager
Cloud Source Repositories after its June 2024 end of sale: how repositories, IAM roles, PushBlock, Pub/Sub push notifications, GitHub and Bitbucket mi…
Read article →Cloud SQL Insights, in depth
Cloud SQL Query Insights in depth: normalization, dimensions and sampled plans, enabling it with gcloud and sizing its settings, sqlcommenter tags, a …
Read article →Cloud Storage Classes, in depth: Standard, Nearline, Coldline and Archive as a cost model, with lifecycle rules, Autoclass and the traps in between
How Google Cloud Storage classes really differ: the five cost components, minimum storage durations and early-deletion billing, retrieval and per-oper…
Read article →Cloud Storage FUSE, in depth: how gcsfuse maps files to objects, caching, streaming writes, GKE and where it breaks
A practical guide to Cloud Storage FUSE (gcsfuse): how file system calls become Cloud Storage requests, the POSIX gaps, metadata and file caches, stre…
Read article →GCS Retention Policy and Object Lock, in depth: Bucket Lock, per-object retention, event-based and temporary holds
How write-once retention works in Google Cloud Storage: bucket retention policies and how expiration is computed, irreversible Bucket Lock and its pro…
Read article →Cloud Vision API
Image and PDF analysis via Cloud Vision API: core capabilities including label detection, face detection, text recognition (OCR), and safe search. Aut…
Read article →Google Compute Engine, in depth: machine families, disks, lifecycle, live migration, the metadata server, Spot VMs and managed instance groups
A working engineer's guide to Google Compute Engine: how a VM is placed and built, choosing machine families and disks, instance states and billi…
Read article →Confidential VMs on GCP, in depth
Confidential VMs on Google Cloud in depth: the threat model, AMD SEV vs SEV-SNP vs Intel TDX and confidential GPUs, how memory encryption and bounce b…
Read article →Config Connector, in depth: managing Google Cloud resources as Kubernetes objects
How Google Cloud Config Connector works: CRDs and controllers, Workload Identity, cluster versus namespaced mode, resource references, reconciliation …
Read article →Custom Images + Instance Templates, in depth: golden image builds, families, deprecation, sharing and drift-free fleets
How Compute Engine custom images work with instance templates: Packer and gcloud builds, image families and zonal resolution, deprecation states, cros…
Read article →Dataplex, in depth: the catalog model, data quality scans, lineage and gating pipelines on scan results
Dataplex (now Knowledge Catalog) explained: entries, aspects and entry groups, data quality rule types and thresholds, a worked orders scan, a pipelin…
Read article →Dataproc Metastore and Data Catalog, in depth: a shared Hive metastore, partitions, federation and cataloging after Data Catalog
Dataproc Metastore explained: what a Hive metastore stores, Thrift versus gRPC, MySQL versus Spanner, attaching clusters, partition registration, fede…
Read article →Deployment Manager, in depth: how configs, templates and manifests work, and migrating every deployment to Terraform before the 2027 turndown
How Google Cloud Deployment Manager works, from configs, templates, references and manifests to create and delete policies, and how to inventory deplo…
Read article →Document AI, in depth: processors and versions, the Document model, online versus batch, validation and review loops
How Google Cloud Document AI works and how to run it well: processor families, pinned versions, text anchors, online and batch calls in Python, an inv…
Read article →Filestore, in depth: managed NFS on Google Cloud, tiers, private networking, client tuning, GKE volumes and data protection
How Google Cloud Filestore works and how to run it: the NFS model and consistency, Basic, Zonal, Regional and Enterprise multishare tiers, reserved ra…
Read article →GCE Instance Scheduling, in depth: host maintenance policy, automatic restart, Spot termination, VM time limits and start/stop instance schedules
How to control when Compute Engine VMs run and stop: the per-VM scheduling options for host maintenance, automatic restart and host errors, Spot provi…
Read article →AlloyDB architecture
How AlloyDB really works: disaggregated log-driven storage, the log processing service, read pools, the in-memory columnar engine, PITR, and Cloud SQL…
Read article →BigQuery architecture
Deep-dive on BigQuery internals: Dremel query engine, Capacitor columnar format, Colossus storage, slot allocation, BI Engine, materialized views.
Read article →BigQuery BI Engine
Deep-dive on BigQuery BI Engine: in-memory columnar caching and vectorized execution, automatic query routing and partial acceleration, LRU working-se…
Read article →Cloud Bigtable Architecture: Row Keys, Tablets and Hotspotting
How Google Cloud Bigtable works: the wide-column data model, row key design to avoid hotspotting, tablets and splits, nodes over Colossus, replication…
Read article →Google Cloud Armor: WAF, DDoS Protection, Rate Limiting
How Google Cloud Armor protects apps behind the external HTTP(S) load balancer: security policy rule order, OWASP WAF rules, rate limiting, L7 DDoS de…
Read article →GCP Cloud Build
Cloud Build in depth: the container-per-step model, the shared workspace, triggers and substitutions, service accounts and logging, secrets, and cachi…
Read article →Google Cloud CDN architecture
Deep-dive on Google Cloud CDN: edge PoP caching enabled on the global HTTPS load balancer, cache hits served close to users while misses fill from the…
Read article →GCP Cloud DNS, in depth: public and private zones, the VPC resolution order, hybrid forwarding, DNSSEC and routing policies
How Google Cloud DNS works as both an authoritative DNS service and the resolver for VPC networks: zone types, the five-step VPC name resolution order…
Read article →Cloud Functions -- event-driven serverless functions on GCP
Deep-dive on GCP Cloud Functions: the function-as-a-service model, triggers (HTTP/Pub/Sub/storage/Eventarc), Gen 2 on Cloud Run (container-based), aut…
Read article →Cloud Run architecture
Deep-dive on Cloud Run: GFE ingress, activator and concurrency-based autoscaler, gVisor/microVM sandboxes, immutable revisions, canary traffic splits,…
Read article →Cloud Run Concurrency and Autoscaling: How the Setting Works
How the Cloud Run concurrency setting and autoscaler set instance count, latency and cost: concurrent requests per instance, cold starts, min/max inst…
Read article →Cloud SQL architecture
Deep-dive on Cloud SQL: synchronous regional persistent-disk HA and cold-standby failover, Auth Proxy and IAM-based connections, async read replicas v…
Read article →Google Cloud Tasks: Queues, Retries, Rate Limits
Google Cloud Tasks explained: HTTP push queues, dispatch rate and concurrency limits, retry and backoff policy, scheduled tasks, deduplication, and OI…
Read article →Google Colossus architecture
Deep-dive on Colossus, Google's cluster-level distributed file system and the successor to GFS that underpins BigQuery, Spanner, Bigtable, and Cl…
Read article →Cloud Composer architecture
How Cloud Composer runs Airflow: the environment split, DAG bucket sync, parse loop, scheduler limits, worker OOM, deferrable operators, and IAM.
Read article →Dataflow architecture
Deep-dive on Cloud Dataflow: Beam runner, streaming engine, autoscaler, shuffle service, watermark, Flex Templates, snapshot.
Read article →Google Cloud Dataproc Architecture: Ephemeral Spark Clusters
How GCP Dataproc runs Spark and Hadoop: master and worker nodes, preemptible secondaries, Cloud Storage as the data lake, autoscaling, ephemeral clust…
Read article →Google Cloud Datastream: CDC to BigQuery with Backfill
Google Cloud Datastream explained: serverless change data capture from MySQL, PostgreSQL and Oracle to BigQuery, backfill plus CDC, private connectivi…
Read article →Eventarc architecture
Deep-dive on Google Cloud Eventarc: direct and Cloud Audit Log sources, the Pub/Sub transport backbone, triggers and attribute filtering, CloudEvents …
Read article →Firestore architecture
Firestore architecture in depth: documents and subcollections, Native vs Datastore mode, index-only queries, realtime listeners, transactions, securit…
Read article →Google Cloud Storage internals, in depth: frontends, Spanner metadata, Colossus and how to design for them
Inside Google Cloud Storage: stateless frontends, Spanner-backed metadata, Colossus data, strong consistency, generations and preconditions, uploads, …
Read article →GKE architecture
Deep-dive on GKE: cluster mode, node pools, autoscaler, workload identity, VPC-native, Gateway, Binary Authorization, Config Sync.
Read article →GCP IAM Roles: Basic, Predefined, Custom, and Principals
How GCP IAM roles work: basic vs predefined vs custom roles, principals and service accounts, bindings inherited from org to folder to project, condit…
Read article →GCP Load Balancing, in depth: choosing among Application, proxy Network and passthrough Network load balancers
Google Cloud load balancing explained from first principles: the forwarding rule to backend resource chain, the product and deployment-mode matrix, ho…
Read article →GCP Cloud Logging
GCP Cloud Logging: the log router and sinks, _Required versus _Default buckets, retention and cost control, Log Analytics, and audit log coverage.
Read article →GCP Memorystore -- managed Redis / Memcached
Deep-dive on GCP Memorystore: the fast-cache need, managed Redis/Memcached (no server ops), tiers (basic vs standard/HA), VPC-private access, replicat…
Read article →GCP Cloud Monitoring, in depth: the time-series model, ingestion paths, PromQL, alerting policies, SLOs and cost control
How Google Cloud Monitoring works and how to run it well: metric descriptors, kinds and monitored resources, metrics scopes, ingestion from system met…
Read article →GCP Multi-Region Architecture in Depth, in depth: global load balancing, Spanner placement, data tier RPO and failover
Multi-region on Google Cloud: global external Application Load Balancer, Spanner replica types and quorums, per-service replication semantics, a worke…
Read article →GCP Overview
What GCP is, how projects and folders organize resources, and where GCP fits versus AWS and Azure.
Read article →GCP Pub/Sub architecture, in depth: forwarders and routers, leases, ordering keys, exactly-once and dead letters
How Google Cloud Pub/Sub works inside: the routers and forwarders, durable publish, subscriptions as cursors, ack deadlines and flow control, ordering…
Read article →GCP Secret Manager, in depth: versions, replication, access control, rotation and running it in production
How Google Cloud Secret Manager works and how to run it well: the secret and version model, automatic, user-managed and regional replication, the acce…
Read article →Cloud Spanner Architecture: TrueTime, Paxos, Splits, Reads
How Cloud Spanner works: TrueTime and commit wait, Paxos-replicated splits, two-phase commit, strong vs stale reads, interleaving, hotspots, when to u…
Read article →Spanner TrueTime
Deep-dive on Spanner TrueTime: the interval-returning TrueTime API, GPS/atomic-clock-bounded uncertainty, commit wait for external consistency, Paxos …
Read article →GCP Cloud Trace in Depth: OpenTelemetry Ingestion, Propagation, Log Correlation, Sampling and Limits
How Google Cloud Trace works in practice: sending OpenTelemetry spans through the OTLP Telemetry API or a Collector, automatic traces from Cloud Run, …
Read article →Vertex AI architecture
Deep-dive on Vertex AI: Model Garden and custom training, Model Registry lineage, online endpoints with traffic splits, batch prediction, KFP pipeline…
Read article →GCP VPC
How GCP VPCs differ from AWS: global scope, subnets per region, and shared VPC patterns.
Read article →GCP VPC Service Controls: Perimeters, Ingress and Egress Rules
How Google Cloud VPC Service Controls stop data exfiltration: service perimeters vs IAM, ingress and egress rules, access levels, bridges, dry-run.
Read article →Google Cloud Workflows architecture
Deep-dive on Google Cloud Workflows: a serverless engine that runs YAML/JSON step definitions as a durable state machine, checkpointing state after ev…
Read article →Cloud Storage Autoclass, in depth: per-object access tracking, the real cost model, and when lifecycle rules still win
How Google Cloud Storage Autoclass works and when to use it: the 30, 90 and 365-day transitions and terminal classes, what counts as access, the 128 K…
Read article →Cloud Storage Encryption, in depth: default keys, CMEK, CSEK, key versions, enforcement and failure modes
How Cloud Storage encrypts objects and how to choose and run Google default encryption, CMEK, CSEK or client-side encryption: setup in gcloud, Terrafo…
Read article →Cloud Storage Lifecycle Management, in depth: rule semantics, versioned buckets, Custom-Time retention and the traps that cost money
How Google Cloud Storage Object Lifecycle Management really evaluates rules: the Delete, SetStorageClass and AbortIncompleteMultipartUpload actions, e…
Read article →Cloud Storage Object Versioning, in depth: generations, preconditions, point-in-time restores and the cost of history
How Cloud Storage Object Versioning works: generations and metagenerations, what overwrite and delete do, generation preconditions for compare-and-swa…
Read article →GCS Regional vs Dual-Region vs Multi-Region, in depth: write acknowledgement, replication RPO, outages and cost
How Cloud Storage location types really behave: writes acknowledged in one region across two zones, default replication (99.9% in an hour, 100% in 12 …
Read article →Cloud Storage Signed URLs and Cookies, in depth: V4 signing, keyless signBlob, browser uploads, POST policies and Cloud CDN signed cookies
How Cloud Storage V4 signed URLs work and how to issue them safely: canonical requests, keyless signing on Cloud Run, browser uploads with CORS, resum…
Read article →GCS Storage Transfer Service, in depth: jobs, sync semantics, agents, event-driven replication and a 400 TB migration
How Google Cloud Storage Transfer Service works: jobs and operations, overwrite and delete options, IAM and S3 credentials, agent pools for file syste…
Read article →Gemini API on Vertex + AI Studio, in depth: two backends, one SDK, and what it takes to run Gemini in production
How the Gemini Developer API (AI Studio) and the enterprise path (Vertex AI, renamed Gemini Enterprise Agent Platform) differ in hosts, auth, location…
Read article →GKE for ML, in depth: accelerator pools, capacity, data and serving
How to run ML training and inference on Google Kubernetes Engine: GPU node pools and drivers, capacity with Spot, reservations and flex-start through …
Read article →Global HTTP(S) Load Balancer, in depth: building, testing and debugging Google Cloud's global external Application Load Balancer
An operating manual for Google Cloud's global external Application Load Balancer (formerly the global HTTP(S) Load Balancer): the object chain in…
Read article →Google Global External LB Family, in depth: the two-layer Google Front End path, cross-region traffic algorithms, auto-capacity draining, advanced routing, timeouts and classic migration
How Google's global external load balancers carry a request from an anycast IP to a backend: first- and second-layer Google Front Ends, service l…
Read article →Identity-Aware Proxy (IAP), in depth: request flow, IAM and access levels, verifying the signed header, programmatic access and TCP forwarding
How Google Cloud Identity-Aware Proxy works and how to deploy it safely: what it protects, how authentication and IAM authorization happen, access lev…
Read article →Instance Templates, in depth: immutable VM blueprints, regional vs global, deterministic images and rolling a MIG between versions
How Compute Engine instance templates work and how to operate them: why templates are immutable, what to put in one, regional versus global scope, det…
Read article →Looker, in depth: LookML, SQL generation, symmetric aggregates, datagroups, PDTs and aggregate awareness
How Looker works as a semantic layer over your warehouse: LookML views, explores and measures, how field selections compile to SQL, fanout and symmetr…
Read article →Managed Instance Groups + Autoscaler, in depth: signals, initialization and stabilization, scale-in controls, schedules, predictive scaling and a worked capacity plan
How the Compute Engine autoscaler sizes a managed instance group: CPU, load-balancer and Cloud Monitoring signals, how multiple signals combine, initi…
Read article →Network Connectivity Center, in depth: hubs, spokes, topologies and route control on Google Cloud
How Google Cloud Network Connectivity Center connects VPC networks and on-premises sites: hubs and spoke types, mesh versus star topologies, export an…
Read article →Persistent Disk Types, in depth: pd-standard, pd-balanced, pd-ssd and pd-extreme, the performance formula, sizing by hand and when Hyperdisk wins
How Google Cloud Persistent Disk types differ and how to choose: media and per-GiB rates, the baseline-plus-rate performance formula, VM and regional …
Read article →Preemptible and Spot VMs on GCP, in depth: the preemption lifecycle, a metadata watcher, checkpoint arithmetic, MIGs and GKE Spot pools
Spot and preemptible VMs on Compute Engine: how they differ from standard VMs, the documented preemption sequence (metadata flag, optional 120 s notic…
Read article →Private Google Access + Private Service Connect, in depth: reaching Google APIs and published services without public IPs
How workloads without external IP addresses reach Google APIs and other services privately on Google Cloud: Private Google Access and its VIP domains,…
Read article →GCP Secret Manager, in depth: running it as an organisation-wide platform with layout, Terraform, inventory and migration
Operating Google Cloud Secret Manager at organisation scale: project layout and labels, Terraform without payloads in state (write-only secret_data_wo…
Read article →Security Command Center
GCP Security Command Center: threat detection, vulnerability management, posture assessment, multi-cloud findings from AWS and Azure, compliance dashb…
Read article →Shared VPC, in depth: host projects, subnet-level IAM, GKE and IP planning
How Google Cloud Shared VPC works: host and service projects, the xpnAdmin and networkUser roles, subnet-level grants, service agents for MIGs and GKE…
Read article →Sole-Tenant Nodes on GCP, in depth: templates, affinity, maintenance policies, overcommit and cost
How Compute Engine sole-tenant nodes work: node types, templates and groups, affinity labels, the default, restart-in-place and migrate-within-node-gr…
Read article →Google Text-to-Speech
Synthesize natural speech from text using WaveNet and Neural2 voices. Multilingual, SSML-controlled, async for bulk, custom voice training (regulated)…
Read article →Google TPU v5 + v6 (Trillium), in depth: v5e, v5p and v6e compared, roofline arithmetic, memory fit, slices and JAX sharding
A practical guide to choosing and using Cloud TPU v5e, v5p and v6e (Trillium): the published per-chip specifications, what the chip contains and how s…
Read article →Model Garden (formerly Vertex AI Model Garden), in depth: three consumption paths, deploying open models, cost crossover and governance
How Model Garden on Gemini Enterprise Agent Platform, formerly Vertex AI, delivers models: managed APIs, partner and managed open models as a service,…
Read article →Vertex AI Search, in depth: data stores, parsing and chunking, the search and answer APIs, filters, grounding checks and running it in production
How Google Cloud's Vertex AI Search (being renamed Agent Search) works: data stores, apps and serving configs, digital, OCR and layout parsers, c…
Read article →GCP VPC Peering, in depth: route exchange, non-transitivity, firewalls and DNS, consensus mode and private services access
How Google Cloud VPC Network Peering works: which routes are exchanged, why it is not transitive, IP planning, gcloud setup with firewall and DNS, con…
Read article →