Lead Platform Engineer - HPC, Kubernetes
Remote in Latvia, Republic of Lithuania
Platform Engineering& 16 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are looking for a Lead Platform Engineer to support a customer that develops and manages several HPC clusters across AWS, CoreWeave, GCP and other providers, operating several thousand GPUs today and scaling 10x. This role is Kubernetes-heavy and requires strong software engineering skills to operate multi-cloud platform infrastructure where misconfigurations or failed upgrades cost thousands of GPU-hours, and at this scale, new kinds of failure come up routinely.
Responsibilities
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle, node pool management, networking policy, and stability during rapid growth
- Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI, with more providers coming
- Manage job scheduling to allocate GPU compute across training and inference workloads
- Define and maintain SLIs/SLOs
- Build monitoring and alerting systems
- Take part in incident response and write post-incident reviews
- Build tooling and automation in production-quality code
- Coordinate daily with the Networking, Storage, Security and AI/ML platform teams
Requirements
- 5+ years of experience in infrastructure engineering, cloud platforms or HPC
- Expertise in Kubernetes at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
- Proficiency in advanced Python with experience writing production-grade tools, not only scripts
- Familiarity with Go, Rust or C++ is a strong plus
- Proficiency in Terraform, writing and reviewing infrastructure as code daily
- Working knowledge of AWS, including EC2, S3, EFS and FSx for Lustre
- Background in Site Reliability Engineering
- English proficiency at B2 level or higher
Nice to have
- Familiarity with Amazon Elastic Kubernetes Service, Google Kubernetes Engine and Google Cloud Platform
- Knowledge of High-performance computing (HPC), Slurm, Lustre and Amazon FSx
- Familiarity with Go, Rust and C++
- Familiarity with CI/CD
