Google Cloud · Infrastructure · YoctoIT tech page

TPU & GPU

The AI infrastructure: Google's TPUs and NVIDIA GPUs for training and inference at scale — the power behind Gemini, for rent.

FOCUS · THE POWER OF AITPUs for transformers, GPUs for everything: AI compute sized and paid right
YoctoIT material for clients and partners · Google Cloud, BigQuery, Gemini and the other products mentioned are trademarks of Google LLC.
01 · What it is

Cloud TPU & GPU, made clear.

Google trains Gemini on its own TPUs — the same ones (Trillium, Ironwood) can be rented by the hour, in pods of thousands of interconnected chips. Alongside, NVIDIA GPUs (H100, H200, B200) for the CUDA ecosystem. For a company the issue isn't the chip: it's sizing, scheduling and paying the right amount.

TPU
Google's chips for transformers: the price/performance for serious AI
Pod
thousands of interconnected chips: distributed training without building anything
DWS
Dynamic Workload Scheduler: the GPU booked when needed, not owned
Cloud TPU & GPU
OFFICIAL GOOGLE CLOUD BRANDING · TPU & GPU
FOTO UFFICIALE GOOGLE · TPU IRONWOOD · FONTE: GOOGLE
OFFICIAL GOOGLE PHOTO · TPU IRONWOOD · SOURCE: GOOGLE
02 · How to use it well

The things that make the difference.

The AI stack

Training & inferenceyour models and fine-tuning
TPU pods
NVIDIA GPUs
Spot & DWS
scale · ecosystem · economics
GKE / Vertex AI on topthe orchestration of the compute
Google's AI data centersthe factory behind Gemini
Supercomputing on demand

The honest sizing

Fine-tuning ≠ pretraining: almost always you need GPU hours, not a TPU pod — the bill says thanks.

Smart scheduling

DWS and spot for patient jobs: the same GPU at a fraction of the price.

Optimized inference

Quantization, batching and autoscaling: cost per request is the metric that counts.

Getting the data out

The dataset close to the compute (regional Cloud Storage): the GPU mustn't wait for the disk.

03 · In depth

TPUs, GPUs and the AI Hypercomputer

For AI on GCP you choose the silicon: TPUs (Trillium/v5) accelerate training and inference of large models with the pod as the unit of scale (dedicated ICI interconnect), NVIDIA GPUs (H100/H200/B200 on A3/A4) cover the CUDA ecosystem; the AI Hypercomputer combines silicon, storage (Hyperdisk ML, Parallelstore) and scheduling (DWS: calendar and flex start) for capacity when needed; multislice scales the training beyond the single pod.

  • TPU pod — thousands of chips on a dedicated interconnect: foundation model training
  • GPU A3/A4 — H100→B200: the CUDA ecosystem at Google scale
  • DWS — Dynamic Workload Scheduler: capacity booked when needed, not always
  • Hyperdisk ML — the storage that feeds the accelerators: no starving GPUs
  • Multislice — training beyond the pod: scale without rewriting the code
  • Inference — TPUs for serving too: the cost per token that goes down
04 · Numbers and lifecycle

The numbers that matter.

~2x
Trillium's perf/watt over the previous generation
k
chips per pod: the scale of serious training
flex
DWS: the discount for those who can wait for capacity
$/token
the real metric of inference: you design for it
AI compute is bought with a plan: sizing, scheduling and cost per token — the right power, when needed, from us.
05 · Use cases

Where it really pays off.

Enterprise fine-tuning

Open models specialized on your data.

Production inference

The AI service with SLAs and cost per request.

Research and simulation

HPC and training when the project demands it.

AI compute is rented, competence isn't: we size and govern the power that's needed.