Site Reliability Engineer
Remote in Mexico
Site Reliability Engineering& 8 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are looking for a Site Reliability Engineer to keep cloud services reliable, observable, and automated across multi-tenant Kubernetes environments on Azure. You will own production health, reduce toil through automation, and partner with developers to improve resilience.
Responsibilities
- Operate Kubernetes clusters and containerized workloads running on Azure
- Troubleshoot production incidents end-to-end across network, OS, platform, and application layers
- Automate repetitive operational tasks with Python, Bash, or PowerShell to eliminate toil
- Define and track SLIs/SLOs and drive improvements to meet reliability targets
- Build and tune monitoring and alerting to detect issues before clients are impacted
- Improve platform reliability through capacity, performance, and failure-mode analysis
- Partner with development teams to harden services and improve operability standards
- Document runbooks and operational procedures to speed up diagnosis and recovery
- Perform root cause analysis and implement preventive actions after incidents
Requirements
- 2+ years of experience in Site Reliability Engineering or DevOps for production systems
- 2+ years of experience operating Kubernetes and containerized workloads
- Hands-on experience with Microsoft Azure services for running workloads
- SLA/SLO adherence experience including defining and tracking SLIs/SLOs
- Infrastructure fundamentals in networking and operating systems
- Strong Linux administration skills
- Strong scripting skills in Python, Bash, or PowerShell
- Incident response skills with calm, structured troubleshooting under pressure
- Clear communication skills for cross-team collaboration during incidents and reviews
- Collaborative mindset to work effectively with development teams
- English proficiency level B2 (Upper-Intermediate)
Nice to have
- Argo CD experience for GitOps-based deployments
- Elastic Stack experience for observability workflows
- Istio experience for service mesh traffic management
- Windows Administration experience including Windows Server operations
