Principal Site Reliability Engineer (SRE)
Remote in Kazakhstan, & 2 others
Platform Engineering& 9 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are seeking an experienced Principal Site Reliability Engineer (SRE) to architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS. This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations. Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch.
Responsibilities
- Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independently
- Provision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patterns
- Design and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground up
- Lead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rules
- Architect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verification
- Partner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries
Requirements
- 7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructure
- Expertise in AWS, Amazon EKS, and Kubernetes in production environments
- Proficiency in Terraform, Karpenter, and GitOps delivery patterns
- Background in designing incident response frameworks, on-call models, and blameless post-mortem processes
- Skills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rules
- Competency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releases
- Understanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundaries
- Familiarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablement
- English proficiency at B2 level or higher
