Lead Site Reliability Engineer
Remote in Mexico, & 3 others
Site Reliability Engineering& 10 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are seeking a Lead Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate DevOps maturity for high-impact services. You will design scalable automation, improve CI/CD and release practices, and lead rapid incident response.
Responsibilities
- Design reliability strategies and SRE practices for business-critical infrastructure
- Build automation and tooling in Python to improve stability, consistency, and operational leverage
- Develop and maintain CI/CD workflows and source control practices using GitLab
- Lead incident response during on-call rotations and restore service for business-critical issues
- Improve release management processes to support enterprise-scale delivery
- Harden cloud infrastructure across networking, compute, security, and IAM controls
- Implement configuration automation to reduce manual work and prevent drift
- Operate and troubleshoot Kubernetes-based workloads and developer-facing platform usage
- Partner with engineering stakeholders to prioritize reliability work and manage change safely
- Assess systemic risks and drive corrective actions to prevent recurring incidents
Requirements
- 5+ years of site reliability engineering or DevOps experience in cloud environments
- Hands-on experience with a leading cloud provider, with practical work across Amazon Web Services and Microsoft Azure
- Leadership ability to guide technical direction and take ownership of critical infrastructure outcomes
- Project delivery experience improving DevOps tools, processes, and engineering maturity at scale
- Deep CI/CD knowledge across pipelines, source control, and release management
- Strong Kubernetes skills with practical usage as a developer
- Advanced Python programming skills for automation and tooling
- Enterprise-scale release management experience supporting complex systems
- Solid infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
- Strong analytical skills to diagnose complex issues and identify high-leverage solutions
- Effective incident response skills, including on-call ownership and rapid restoration of service
- Upper-Intermediate English proficiency (B2, Upper-Intermediate)
Nice to have
- Amazon Web Services certification or proven advanced AWS operational experience
- Microsoft Azure certification or proven advanced Azure operational experience
- AI Architecture experience for reliability-focused platform design
- AI Solution Engineering experience integrating AI-enabled capabilities into operations
- Gen AI Solutions Development experience for operational intelligence and automation use cases
