BRIEF #18
September 21, 2026

Platform Pulse: Agent Substrate on GKE, Spanner Queues, and Native BM25 Search

In this 18th edition of the Engineering Brief, we examine the open-sourcing of Agent Substrate on GKE, explore transactional messaging with Spanner queues alongside the removal of DML mutation limits, evaluate native BM25 search in AlloyDB and Cloud SQL, and unpack ADK for Kotlin 1.0 with runtime agent anomaly detection.

🔐 Zero-Trust Security, Identity & Threat Defence

Security frameworks are evolving beyond static boundary policies towards dynamic runtime authorisation, continuous session evaluation, and intent-aware governance for autonomous AI agents.

  1. Build zero-trust AI agents that judge intent, not just syntax: Moving enterprise AI agents from static build-time boundaries to dynamic runtime governance on Gemini Enterprise Agent Platform, combining Model Armor edge filtering, Semantic Governance Policies, and out-of-band behavioural scanning.
  2. Agent Anomaly Detection, now in Private Preview on the Gemini Enterprise Agent Platform: An out-of-band oversight architecture for Gemini Enterprise Agent Platform that inspects OpenTelemetry spans and tool calls to flag multi-turn logic deviations and OWASP Agentic Top 10 exploits without introducing runtime latency to user requests.
  3. Killing the “Keys to the Kingdom”: Automated Just-In-Time Policy Elevator on Google Cloud: An open-source Just-In-Time access broker eliminating standing administrative privileges in Google Cloud environments by leveraging project-scoped IAM Common Expression Language (CEL) conditions, peer review sign-offs, and automated event watchdogs.
  4. Changing the game: How Google uses agentic AI to secure hundreds of millions of lines of code: How Google deploys specialised autonomous security agents across internal development pipelines to review every commit, verify system invariants, and neutralise code vulnerabilities before deployment to production infrastructure.
  5. Introducing new session management tools with native, granular controls: Google Cloud unveils granular session controls natively integrated with Context-Aware Access, enabling platform teams to enforce continuous posture checks, terminate stale credentials, and establish context-specific session lifetimes.
  6. GTIG AI Threat Tracker: From Prompting to Autonomy - The Evolution of Adversarial AI: The Google Threat Intelligence Group assesses emerging adversary playbooks, tracing how attackers are transitioning from simple prompt engineering and phishing to multi-agent command-and-control frameworks and automated exploitation.
  7. Cloud CISO Perspectives: How Google monitors AI threats and advances AI defences: Google Cloud CISO Perspectives analyses three fundamental shifts in the global cyber threat environment and outlines defensive blueprints across identity federation, infrastructure isolation, and AI-accelerated SOC operations.
  8. Mastering fleet-wide governance on GKE using fleet level OPA: Establishing centralised policy enforcement across distributed GKE clusters using GKE Policy Controller, standardising baseline security constraints, container registry allowlists, and resource quotas across entire Kubernetes fleets.
  9. New to Google SecOps: Building Your First Search with SQL Pipes: Querying enterprise telemetry in Google Security Operations with GoogleSQL pipe syntax, providing security analysts with a clean, sequential pipeline structure to aggregate, filter, and correlate security logs alongside YARA-L detection rules.
  10. Google named a Leader in the External Threat Intelligence Service Forrester Wave: Forrester evaluates global cybersecurity providers, placing Google in the Leaders category for external threat intelligence based on Mandiant frontline telemetry and proactive adversary tracking.
  11. Policy Intelligence: Agent Identity Access Troubleshooting: Policy Troubleshooter adds native support for autonomous agent identities, allowing security administrators to diagnose IAM allow, deny, and principal access boundary policies directly from principal identifiers or denied permission error IDs.
  12. Cloud Armor Managed Rulesets Available in Preview: Google Cloud Armor introduces managed rulesets in Preview, delivering automated signature updates to shield backend services, public APIs, and serverless applications against emerging web exploits without manual rule tuning.
  13. Identity and Access Management (IAM) Remote MCP Server: The official IAM remote Model Context Protocol (MCP) server reaches General Availability, enabling external AI applications to inspect IAM roles, audit deny policies, and query permissions safely across cloud resources.
  14. Managed Workload Identity for Backend mTLS on Application Load Balancers: General Availability of managed workload identity for backend mutual TLS across Google Cloud Application Load Balancers, automating certificate issuance, rotation, and trust domains via Certificate Authority Service.

📊 High-Performance Databases, Lakehouse & Big Data Analytics

Data architectures are converging around transactional queues, native hybrid BM25 search, cross-cloud lakehouse caching, and automated in-database causal modelling.

  1. Spanner: Removing cumulative mutation limits for DML transactions: Cloud Spanner eliminates the cumulative mutation ceiling for DML transactions, simplifying bulk batch mutations and large data modifications without requiring manual transaction splitting.
  2. Spanner Queues Reach General Availability: Cloud Spanner queues are Generally Available, bringing transactional messaging directly into Spanner to power horizontally scalable, event-driven workflows with strict ACID guarantees.
  3. Announcing Native BM25 Ranking in AlloyDB and Cloud SQL: AlloyDB and Cloud SQL for PostgreSQL introduce native BM25 full-text indexing via the pg_textsearch extension, eliminating external search dependencies for keyword relevance and hybrid vector search pipelines.
  4. Enterprise-grade PostgreSQL with AlloyDB Omni RPM Orchestrator is generally available: AlloyDB Omni RPM Orchestrator reaches General Availability, providing turnkey deployment automation, high availability, and dynamic read-pool scaling for on-premises and sovereign PostgreSQL workloads.
  5. Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching: BigQuery introduces cross-cloud caching in Preview, dramatically reducing cross-cloud data transfer fees and boosting query performance when querying federated lakehouse tables across AWS S3 and Azure Data Lake Storage.
  6. Agent-ready analytics: Unlocking insights with BigQuery augmented analytics: BigQuery unveils augmented analytics Table-Valued Functions (TVFs) that automatically discover statistical outliers, explain data patterns, and surface key drivers using in-database machine learning and statistical methods.
  7. How considering an Iceberg migration made me rethink BigQuery costs: Evaluating an enterprise Apache Iceberg migration on Google Cloud, examining how faster query runtimes can paradoxically increase compute slot consumption, and reassessing storage versus compute billing models.
  8. Concurrent Microbatch Support Exists for BigQuery dbt: A deep dive into unannounced support for dbt concurrent microbatch processing in BigQuery, demonstrating how data engineers can slash backfill windows by running parallel incremental partitions.
  9. Announcing Pause/Resume and NVIDIA RTX PRO 6000 Blackwell GPU support in Dataflow: Cloud Dataflow introduces pipeline Pause and Resume capabilities to optimise operational costs during maintenance windows, alongside NVIDIA RTX PRO 6000 Blackwell GPU acceleration for high-throughput streaming inference.
  10. How We Lit Up Dataflow’s Black Box with OpenTelemetry tracing: Tracing distributed streaming pipelines on Cloud Dataflow using end-to-end OpenTelemetry spans, uncovering hidden serialisation delays and pipeline stage bottlenecks that standard monitoring metrics conceal.
  11. The $5,000 GA4 BigQuery Unnesting Mistake: Scalar Subqueries vs. CROSS JOIN: Illustrating how unpacking nested GA4 event parameters via repeated scalar subqueries drives catastrophic slot consumption, and how refactoring to canonical CROSS JOIN UNNEST queries slashes Google Cloud bills.
  12. Arrays in BigQuery: what they are and how to get data out of them: A practical guide to manipulating nested structures in BigQuery SQL, detailing efficient usage of UNNEST, safe offset indexing, and shorthand aggregation functions like ARRAY_FIRST and ARRAY_AGG.
  13. Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow: Google open-sources Distributed GraphFlow, enabling telecom operators to train and deploy Graph Neural Networks (GNNs) across petabyte-scale network graphs stored in Cloud Spanner and BigQuery.
  14. How a solo founder runs a five-continent tender platform on AlloyDB and MCP: How Lucius AI operates a global procurement platform spanning 210,000 tenders using AlloyDB for PostgreSQL, achieving 47x faster ScaNN vector searches and orchestrating complex database tasks over MCP.
  15. Embedding versions management, TOAST and bloating in PostgreSQL: Investigating database bloat, index expansion, and TOAST table overhead in PostgreSQL and AlloyDB when refreshing dense vector embeddings, detailing why standard vacuuming leaves physical disk allocations bloated.
  16. Beyond DMS: Accelerating Migrations SQL Server Logins and Users to Cloud SQL: Replicating SQL Server logins, permissions, and hashed credentials into Cloud SQL databases during database migrations while maintaining strict security boundaries and pruning orphaned access.
  17. The future of orchestration: Pine59’s journey to Airflow 3 on Google Cloud: Supply chain company Pine59 details modernising its monorepo data pipelines to Apache Airflow 3 on Google Cloud, cutting pipeline execution latency and streamlining complex MLOps workflows.
  18. Unit Testing for Data Pipelines and WHERE Filters in Aggregates in BigQuery: BigQuery introduces General Availability for pipeline unit testing to validate SQL transformations against mock datasets, alongside Preview support for embedding WHERE filter clauses directly inside aggregate function invocations.

⚡ Cloud-Native Infrastructure, GKE & Serverless Workloads

Platform engineering teams are driving container density to extreme scales, orchestrating dynamic GPU sharing, and hardening multi-cloud networking boundaries.

  1. Agent Substrate brings high-density, scalable, trusted infrastructure to GKE: Google announces Agent Substrate, an open-source, high-density runtime on GKE capable of orchestrating millions of secure, sandboxed code-execution environments across a single Kubernetes cluster.
  2. For SeaVerse, GKE Agent Sandbox reduces infrastructure costs by 60%: AI-based gaming platform SeaVerse achieves 60% compute infrastructure savings by deploying GKE Agent Sandbox, delivering sub-second sandbox initialisation and hardware-level isolation for dynamic user scripts.
  3. Introducing Filestore agent volumes: fully managed storage for agent workspaces: Google Cloud launches Filestore agent volumes, offering fully managed, elastic POSIX storage designed specifically for high-concurrency, multi-agent workspaces and shared scratch filesystems.
  4. The Idle Half of a GPU Pod: An empirical analysis of high-end Google Cloud GPU nodes, proving that host CPUs and RAM sit largely idle during model inference and can be safely backfilled with compute-intensive workloads without harming inference latencies.
  5. Sharing One Physical GPU Across Multiple Pods on GKE with Dynamic Resource Allocation: Leveraging Kubernetes Dynamic Resource Allocation (DRA) on Google Kubernetes Engine to dynamically carve and allocate slices of a physical NVIDIA GPU across multiple Pods, maximising accelerator utilisation.
  6. How We Built GPU Failover on GKE Without Giving Up Infrastructure Control: Implementing declarative, priority-based GPU failovers on GKE using Config Connector and custom ComputeClasses to gracefully migrate across available node pools during capacity shortages without compromising infrastructure governance.
  7. VPC Peering, Private Services Access and Private Service Connect: What Happens at the Boundary: A definitive networking guide comparing VPC Peering, Private Services Access (PSA), and Private Service Connect (PSC), analysing route exchange, IP overlap constraints, and transitive communication boundaries in GCP.
  8. What I Learned About Cloud Run Instances: Dissecting Google Cloud Run instances to explore how long-lived container lifecycles, memory caching, and persistent connections unlock cost-effective hosting for stateful background workers and interactive agents.
  9. Taking Advantage of Cloud Run Sandboxes with Google Apps Script for Google Workspace: Bridging Google Apps Script with Cloud Run sandboxes to run deterministic, sub-second Python and Bash routines under zero-trust gVisor isolation with zero idle infrastructure cost.
  10. GitOps for AI Agents on Google Cloud: ArgoCD, Config Sync, and GKE: Implementing declarative GitOps workflows for autonomous agent deployments on GKE using ArgoCD and Config Sync, ensuring full version control, automated rollbacks, and auditability for agent infrastructure.
  11. Exit Through the Wrong Pipe: Diagnosing a classic Google Kubernetes Engine quirk where standard container stderr logs are ingested as critical platform errors, and providing structural workarounds to silence false alerts in Cloud Logging.
  12. Best practices for handling cloud reliability incidents: A field guide from Google Cloud SRE and Consulting detailing operational protocols, cross-functional incident coordination, and post-incident reviews to minimise business disruption during unexpected outages.
  13. M4N VM Family Reaches General Availability for I/O and Memory-Intensive Workloads: Google Cloud announces the General Availability of M4N instances on Compute Engine, featuring 5th Gen Intel Xeon processors, 400 Gbps network throughput, and up to 20% TCO savings for demanding relational database engines.
  14. Regional Endpoints Generally Available for Cloud SQL for PostgreSQL Admin API: Regional endpoints reach General Availability for Cloud SQL for PostgreSQL, keeping administrative API traffic strictly within the instance region to meet stringent sovereign data residency standards.

🛠️ AI Agents, Runtimes & Platform Modernisation

The developer landscape is shifting towards installable agent plugins, standardised harnesses, cross-platform mobile ADK runtimes, and open-source SDK generation tooling.

  1. Introducing the Google Cloud Developer Plugin for AI Coding Agents: Google Cloud launches developer plugins for AI coding assistants, bundling domain-specific Agent Skills and Model Context Protocol servers into modular, installable packages to streamline cloud engineering.
  2. Power agent hubs or custom harnesses with the Antigravity SDK in one toolkit: A blueprint for building enterprise agent control planes from scratch using the Antigravity SDK, providing deterministic execution sandboxes, structured event auditing, and multi-agent lifecycle management.
  3. Announcing ADK for Kotlin 1.0: Building Production-Ready AI Agents in Kotlin, Android, and Beyond: Google officially releases version 1.0 of the Agent Development Kit (ADK) for Kotlin on Kotlin Multiplatform, introducing KSP zero-reflection tool calling, Room session storage, AppSearch semantic memory, and local LiteRT-LM support.
  4. Enterprise AgentOps: Decoupling Terraform Infrastructure from Google ADK Agent Deployments on Vertex AI: Implementing a robust two-stage deployment architecture on Vertex AI that uses Terraform for Day One Reasoning Engine infrastructure while allowing developers to continuously update ADK application code via CI/CD.
  5. The Anatomy of Harness Engineering: How to Evaluate, Iterate, and Guard AI Coding Agents: Outlining a behavioural evaluation methodology for coding agents that replaces opaque end-to-end scoring with granular trajectory checks, intermediate tool audits, and regression guards.
  6. Inside Google’s Gemini 3.8 Flash and FlashAttention-3: The Hardware Physics of 2M Tokens: An engineering exploration into the hardware dynamics of Gemini 3.8 Flash, explaining how FlashAttention-3, Grouped-Query Attention, and Ring Attention conquer quadratic memory bottlenecks to sustain 2-million-token contexts.
  7. Autonomous LLM post-training with Tunix on TPUs: Introducing autofinetune, an automated multi-agent framework that runs supervised fine-tuning and reinforcement learning loops autonomously overnight on Google Cloud TPUs using Tunix and Gemma.
  8. How Google Cloud Plugins are the Ultimate Cloud Toolkit for Claude Code: Integrating Google Cloud Agent Skills and Model Context Protocol servers directly into Claude Code via the plugin interface, enabling developers to discover, configure, and invoke GCP services from the terminal.
  9. Introducing Firebase spend caps: Firebase rolls out automated spend caps for Cloud Functions and the Gemini API, giving teams automated service pauses and multi-stage alerting to prevent surprise cloud bills.
  10. Building enterprise AI agents with ADK and the Gmail MCP Server: A production guide to connecting Google ADK agents to the Gmail MCP server on Cloud Run, comparing user-consent OAuth against domain-wide delegation while outlining token caching and permission boundaries.
  11. From manual to autonomous: A blueprint for agentic process automation on GCP: Architecting a collaborative multi-agent engine on Google Cloud that transforms enterprise customer engagement data in BigQuery into autonomous, real-time retention campaigns.
  12. Why client SDK generation belongs in the open: Google partners with Speakeasy to open-source its OpenAPI code generator under the AGPLv3 licence, providing developers with deterministic, strictly-typed multi-language SDK generators and MCP documentation endpoints.
  13. Three ways to save a billion tokens with Firebase AI Logic and on-device AI for Chrome: Slashing cloud token consumption and API expenses by routing local summarisation and classification workloads to on-device Gemini Nano in Chrome while reserving Firebase AI Logic for heavy reasoning.
  14. 5 ways to use Gemini text-to-speech in your apps with Firebase AI Logic: Integrating Gemini multimodal text-to-speech models into web and mobile clients using Firebase AI Logic, delivering natural audio generation with low streaming latency.
  15. Launching MCP Apps in Toolbox: The MCP Toolbox for Databases introduces support for the MCP Apps extension, allowing database administrative tools to serve interactive graphical interfaces and charts inside AI chat clients.
  16. Zero-infrastructure managed XProf: Profiling ML workloads on Cloud TPU with ML Diagnostics: Instrumenting JAX training pipelines on Cloud TPU VMs with the Google Cloud ML Diagnostics SDK to capture XProf performance traces and streaming hardware telemetry directly within Cloud Console.
  17. Slot-Filling Framework: Guide to Building Deterministic Agents on Google Cloud CXAS: Implementing the slot-filling framework within Google Cloud Customer Experience Agent Services (CXAS) to unite natural conversational flow with guaranteed deterministic business logic execution.
  18. GCP Bytes Podcast Episode #49: Discussing always-on Cloud Run architectures, Cloud Fault Injection Testing, Gemini Flash 3.8 hardware advancements, enterprise AI agent adoption in financial services, and vector memory management with mem0.

📋 Essential Release Notes

A curated review of runtime additions, enterprise security controls, and managed infrastructure milestones rolling out across Google Cloud.

  1. API Gateway: API Gateway introduces Public Preview support for configuring gateways as remote Model Context Protocol (MCP) servers, enabling organisations to expose existing OpenAPI 3.x REST services directly to AI agents as callable tools without backend code modifications.
  2. Cloud Spanner: Spanner queues reach General Availability (GA), delivering a fully managed, transactional messaging engine that pairs asynchronous work distribution with Spanner's horizontal scalability and strong consistency.
  3. Cloud SQL for PostgreSQL: The pg_textsearch extension is available for PostgreSQL 17+, bringing native BM25 full-text ranking algorithms to Cloud SQL. In addition, point-in-time recovery (PITR) enablement after DR switchover or replica failover now executes asynchronously, significantly reducing recovery times.
  4. Cloud SQL for MySQL: Disaster recovery switchover and replica failovers now execute faster by decoupling PITR enablement into a background asynchronous process that no longer blocks failover completion.
  5. BigQuery: Data pipeline unit testing is Generally Available (GA) to test SQL transformation logic against mock datasets. Aggregate functions add Preview support for inline WHERE clauses, BigQuery Graph metadata automatically synchronises with Knowledge Catalogue, and migration lineage visualisation enters Preview.
  6. Bigtable: Parameterised views are Generally Available (GA) in the Cloud Console and Bigtable Studio, enabling teams to save, parameterise, and run secure SQL queries across Bigtable instances.
  7. Cloud Run: Cloud Run jobs introduce delayed execution in Preview, allowing non-urgent asynchronous batch processing to be deferred for up to 12 hours to take advantage of reduced pricing tiers.
  8. Google Kubernetes Engine (GKE): Version 1.35.7-gke.1222000 is now the default version for cluster creation, alongside newly released patch releases across 1.34, 1.35, and 1.36 release channels.
  9. Cloud Filestore: Small capacity Filestore instances for the Regional service tier are Generally Available (GA), starting at 100 GiB and scaling in 1 GiB increments to provide cost-effective shared NFS storage for agent workspaces and testing workloads.
  10. Compute Engine: The storage-optimised Z4D machine series powered by AMD EPYC Turin and Titanium processors is Generally Available (delivering up to 3 TB memory and 42 TB local NVMe storage). Network-optimised C4N instances with Titanium SSD enter Public Preview, and regional disks can now be created from custom and public OS images.
  11. Chronicle Security Operations: GoogleSQL support in SecOps Search enters Public Preview, allowing security engineers to interrogate UDM event logs, entity graphs, and detection rules using standard SQL or Piped SQL syntax.
  12. Dataform: Dataform unit testing reaches General Availability (GA), allowing analytics engineers to validate SQLX transformation logic against mock data and expected output assertions.
  13. Datastream: Datastream adopts MongoDB extended JSON canonical mode as the default replication format for MongoDB-to-BigQuery CDC pipelines, preserving complete BSON data type fidelity during streaming ingestion.
  14. Secret Manager: Parameter Manager introduces CRC32C checksum validation to verify data integrity across version writes and reads, alongside tag-based access control via Identity and Access Management policies.
  15. Virtual Private Cloud (VPC): General Availability of VPC Flow Logs metadata annotations for Private Service Connect interface endpoints, direct VPC egress logging for App Engine, and multiple dynamic NIC attachments within a single VPC network.
  16. Looker: The Looker extension for VS Code is Generally Available (GA), enabling local LookML development, Git branch synchronisation, and Model Context Protocol (MCP) integration for AI-assisted modelling.
0

From the Community

No community links this week.

Enjoyed this brief?

Don't miss the next drop.