Notes from Google Cloud Next 2026

Work sent me to Google Cloud Next this year. Almost everything was about AI. There were even banners in the hallway saying “Put an agent on it”. The engineering talks, like the GKE one, ran in small rooms with thin crowds.

The GKE talk

My favorite talk opened with Yahoo Mail moving petabytes of data onto GCP (yes, they’re still around). They split GKE into separate clusters for ingress, mail, add-ons like contacts and calendar, and analytics, so a runaway batch job can’t take down mail delivery. They picked GCP for Spanner’s multi-region consistency and to run GKE next to that data.

Google spent the rest of the talk on autoscaling and capacity, aimed at staying fast without over-provisioning:

  • Predictive HPA forecasts demand from history and scales out a few minutes ahead, for workloads with long startups and cyclical traffic.
  • Native custom-metric HPA reads a Prometheus metric straight off the pod’s /metrics through an AutoscalingMetric CRD, skipping the Cloud Monitoring adapter and its pile of IAM setup. The demo scaled on vLLM’s KV-cache usage.
  • Compute classes let you hand GKE a prioritized list of machine types (spot first, on-demand fallback) and have it find capacity, with node-pool auto-creation you can scope per class instead of cluster-wide.
  • Capacity buffers keep spare nodes warm so critical workloads can scale out without waiting for provisioning. The speaker pitched them as an alternative to balloon pods, low-priority placeholder pods parked to hold space.
  • VPA can now resize a running pod in place instead of restarting it, so rightsizing doesn’t interrupt the service.
  • CPU Boost uses that in-place resize to give a pod extra CPU during startup and take it back once the pod is ready, so a slow-starting JVM doesn’t hold that headroom for life.

The CPU Boost config they showed:

startupBoost:
  cpu:
    type: "Factor"
    factor: 3            # 3x the CPU request and limit while starting
    durationSeconds: 10  # keep the boost 10s past readiness

BigQuery

The BigQuery talk was mostly about making it a better source of context for models:

  • Fluid Scaling drops the 60-second minimum charge on autoscaled queries, so you pay per second for bursty, agent-driven traffic.
  • ObjectRef is a column type pointing at an object in storage (an image, an audio file), so you can join unstructured data with structured rows and run AI/ML on it in one query.
  • Hybrid Search runs vector and lexical search together and merges them with a reranker, for RAG.
  • Observability gained job-level cost attribution to answer “who is burning the budget”.

They also made it easier to call LLMs from SQL. An optimized mode for AI.CLASSIFY and AI.IF has Gemini label a sample of rows, then trains a small model on those labels to handle the rest. Google says that uses 230x fewer tokens than calling Gemini on every row. It still sounds like an easy way to accidentally burn a ton of money.

The AI talks

  • Gemini Enterprise was the keynote’s centerpiece and looked like Google’s Glean plus a drag-and-drop GUI for building no-code agent flows from predefined steps.
  • BigLake is now the Cross-cloud Lakehouse, built on Apache Iceberg so you can query data sitting in other clouds without copying it.
  • AlloyDB is Google’s Postgres-compatible answer to Aurora, and it got optimized modes for its AI functions too, like ai.if.
  • Google’s ADK agent framework picked up graph-based and multi-agent workflows and Agent Skills for Python, Go, and TypeScript.

Semantic layers like LookML have been around for years, but they matter more once an LLM writes the queries. A Looker talk said the descriptions in LookML help LLM accuracy a lot. It also seemed like a pitch for running more LLM calls through Looker, where the tokens bill separately.

L’Oreal’s talk was about opening agents to non-programmers on the business side. They built their agent platform on ADK, with MCP tools that can even pull reports, and the talk showed off some Gemini Enterprise too. I like the goal, since programmers shouldn’t be the only ones who get to build useful things. The speaker was vague about what business users are allowed to do beyond sticking to the guardrails and brand guidelines. Code review never came up, though an earlier slide seemed to show an approval process. I’d want some review, mainly around exfiltration. If those agents are limited to Gemini Enterprise’s predefined steps, how much one could leak depends on whether any step can make an arbitrary HTTP request. I couldn’t tell, and the controls I’d want might exist and just not have made the talk.

A separate Wayfair talk covered the Universal Commerce Protocol, a standard they’re building with Google and other retailers so LLMs can surface products and you can check out without leaving the chat. They said their LLM-referred traffic doubled in two quarters and converts better than their usual channels, though it’s still a single-digit share of total traffic.

The expo floor

Google had the accelerators behind AI Hypercomputer out on the floor. One was an Ironwood TPU rack, where a full pod is 9,216 liquid-cooled chips sharing 1.77 petabytes of HBM over an optical fabric that steers light between fibers with tiny tilting mirrors. The other was an A4X Max rack, NVIDIA’s GB300 NVL72, with 72 Blackwell Ultra GPUs pulling around 120 kW.

Boston Dynamics had Spot patrolling a mock industrial site, and Google and Mars ran an “AI Snack Factory” where Gemini invents candy flavors and attendees vote. “Golden Latte Flambe” was winning.