Skip To Main Content
backBack to Search

Lead Site Reliability Engineer

Remote in Argentina, & 3 others
Site Reliability Engineering& 14 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.

Responsibilities
  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex distributed-system issues using logs and telemetry to identify root causes
  • Deploy patches to address issues in cloud infrastructure
  • Improve reliability posture for key backend services through tangible changes
  • Build metrics and dashboards to quickly assess overall platform health
  • Monitor SLOs across services and submit code changes to improve SLOs as errors occur
  • Create and refine runbooks for backend services to standardize operations and response
Requirements
  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
  • 5+ years of experience with Amazon Web Services in production environments
  • 3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in production
  • Leadership experience guiding on-call practices and operational improvements
  • Strong project execution skills delivering reliability and observability enhancements
  • Strong troubleshooting skills using logs and telemetry to find root causes
  • Strong communication skills for clear and concise written incident updates
  • Fast learning ability to absorb information quickly and apply it during on-call
  • Upper-Intermediate English proficiency (B2)
Nice to have
  • Kubernetes
  • Terraform
  • Apache Kafka
  • Grafana
  • New Relic