These steps guide you through the cluster creation process for TPUs using DWS Flex Start.
-
Complete the common setup steps (1-4) in the Create a cluster section.
-
In the
examples/gke-consumption-options/dws-flex-start/gke-tpu-v6e/gke-tpu-v6e-deployment.yamlfile, fill in the following settings in the terraform_backend_defaults and vars sections to match the specific values for your deployment:bucket: the name of the Cloud Storage bucket you created in the previous step.deployment_name: the name of the deployment.project_id: your Google Cloud project ID.region: the compute region for the cluster.zone: the compute zone for the node pool of TPU v6e machines.enable_flex_start: set totrueto enable DWS Flex Start.autoscaling_min_node_count: set to0(required for Flex Start).autoscaling_max_node_count: set to the required node count for your topology (e.g.,4for a4x4topology).authorized_cidr: The IP address range that you want to allow to connect with the cluster.system_node_pool_disk_size_gb: the size of disk for each node of the system node pool.v6e_node_pool_disk_size_gb: the size of disk for each node of the TPU node pool. To modify advanced settings, editexamples/gke-consumption-options/dws-flex-start/gke-tpu-v6e/gke-tpu-v6e.yaml. -
Generate Application Default Credentials (ADC) to provide access to Terraform.
gcloud auth application-default login
-
Deploy the blueprint to provision the GKE infrastructure using TPU v6e machine types:
cd ~\/cluster-toolkit ./gcluster deploy -d \ examples/gke-consumption-options/dws-flex-start/gke-tpu-v6e/gke-tpu-v6e-deployment.yaml \ examples/gke-consumption-options/dws-flex-start/gke-tpu-v6e/gke-tpu-v6e.yaml
-
When prompted, select (A)pply to deploy the blueprint.
- DWS Flex Start does not work with static nodes. So, static_node_count cannot be set.
- To use DWS Flex Start,
auto_repairshould be set tofalse.
The Cluster Toolkit automatically generates a pre-configured, DWS-compliant job file located in your deployment folder (e.g., gke-tpu-7x-flex/primary/my-job-xxxx.yaml).
If you wish to create your own custom job, ensure it includes the following critical settings:
The job must target the TPU v6e nodes and tolerate the standard TPU taint:
nodeSelector:
cloud.google.com/gke-tpu-accelerator: "tpu-v6e-slice"
cloud.google.com/gke-tpu-topology: "4x4"
tolerations:
- key: google.com/tpu
operator: Equal
value: "present"
effect: NoScheduleEach pod in the job should request the full amount of TPU chips available on the node (typically 4 for ct6e-standard-4t):
resources:
limits:
google.com/tpu: 4
requests:
google.com/tpu: 4The parallelism and completions count in your Kubernetes Job (or ReplicatedJob replicas in a JobSet) must exactly match the number of nodes in your TPU slice (e.g., 4 for a 4x4 topology).
When using the TPU Flex Start model, the cluster begins with 0 nodes in the TPU node pool. You can verify the dynamic scaling by following these steps:
-
Monitor the cluster status: Open two terminals. In the first, watch the pods:
kubectl get pods -w
In the second, watch the nodes:
kubectl get nodes -w
-
Submit the TPU job: Submit the generated job file:
kubectl apply -f <deployment_folder>/primary/my-job-xxxx.yaml
-
Observe Scale-Up:
- Initial State: Pods will show as
Pending.
NAME READY STATUS RESTARTS AGE my-job-775c-0-stfds 0/2 Pending 0 10s-
Autoscaling Triggered: Check events to see the scale-up trigger:
kubectl get events. You will seeTriggeredScaleUp. -
Nodes Joining: After a few minutes, nodes will appear and transition to
Ready.
NAME STATUS ROLES AGE VERSION gke-tpu-46f78452-4hkf Ready <none> 10s v1.33.5-gke.2392000- Running: Once nodes are ready, pods will transition to
Running.
- Initial State: Pods will show as
-
Observe Scale-Down:
- Completion: Once the workload finishes, pods will move to
Completed.
NAME READY STATUS RESTARTS AGE my-job-775c-0-stfds 0/2 Completed 0 5m- Automatic Removal: After a short idle period (typically 1-10 minutes), the Cluster Autoscaler will delete the nodes.
gke-tpu-46f78452-4hkf NotReady,SchedulingDisabled <none> 6m v1.33.5-gke.2392000- Final State: The node pool will return to 0 nodes.
- Completion: Once the workload finishes, pods will move to