Senior Site Reliability Engineer
Remote in Mexico, & 3 others
Site Reliability Engineering& 10 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are seeking a Senior Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will drive improvements across CI/CD, cloud platforms, and operational readiness while solving high-impact production issues.
Responsibilities
- Design reliability strategies for critical infrastructure and services
- Build and maintain CI/CD pipelines and release workflows using GitLab
- Automate infrastructure operations and tooling with Python
- Operate cloud environments across AWS and Azure to meet availability goals
- Harden and standardize infrastructure practices across networking, security, IAM, and compute
- Coordinate incident response during on-call shifts and restore service quickly
- Implement Kubernetes-based deployment and operational patterns to improve stability
- Analyze reliability signals and root causes to prevent repeat incidents
- Improve DevOps processes and engineering capabilities to enable faster change
- Partner with stakeholders to prioritize resilience work over short-term fixes
Requirements
- 3+ years of site reliability engineering experience in cloud environments
- Proven leadership ability to influence reliability practices across teams
- Enterprise-scale release management experience supporting frequent deployments
- Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure
- Advanced Python programming skills for automation and tooling
- Solid Kubernetes skills using clusters as a developer
- Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms
- Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
- Strong analytical skills for complex problem solving under pressure
- Upper-Intermediate English proficiency (B2)
- Reliable on-call readiness to assess and resolve business-critical issues
Nice to have
- Amazon Web Services expertise, including design patterns for resilient systems
- Microsoft Azure expertise, including governance and operational best practices
- AI Architecture experience applied to platform reliability and automation
- AI Solution Engineering experience for production-grade AI-enabled operations
- Gen AI Solutions Development experience focused on operational use cases
