Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are looking for a Senior/Lead Cloud DevOps Engineer to provision and manage the GCP/GKE infrastructure foundation for TPU workloads, including topology-shaped node pools, automation, and cost/quota guardrails.
Responsibilities
- Provision GKE topology-shaped node pools using Terraform/Helm
- Establish quota and cost guardrails for TPU accelerator usage
- Manage TPU capacity as a project-critical dependency (quota requests, reservations, regional availability)
- Set up logging/monitoring (Prometheus, Grafana, Cloud Logging)
- Support cloud networking, security, and IAM configuration for TPU workloads
Requirements
- Strong GCP expertise (GCE, GKE) and Kubernetes container orchestration
- Infrastructure-as-Code proficiency (Terraform, Helm)
- Scripting proficiency (Python, Bash)
- Experience with monitoring/observability stacks (Prometheus, Grafana, Cloud Logging)
- Understanding of cloud networking, security, and IAM
- Experience running/fine-tuning LLM models is a plus
Nice to have
- Experience managing accelerator (GPU/TPU) quota and capacity planning at scale
- Cloud financial/cost-optimization experience (FinOps)
