Skip To Main Content
backBack to Search

Lead Site Reliability Engineer

Remote in Mexico
Amazon Web Services& 5 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are seeking a skilled Lead Compute Platform SRE to support EPAM's Compute Managed Services project for our client.

The role focuses on KTLO (Keep the Lights On) activities, ensuring 24x7 monitoring, incident management, and operational stability across multi-cloud environments (GCP, AWS, Azure). The SRE will drive observability improvements, automate processes, and maintain compliance while collaborating with cross-functional teams to deliver high-quality compute services.

Responsibilities
  • Perform continuous 24x7 monitoring of compute platforms using tools such as ELK and PagerDuty
  • Manage incidents and problems across servers, middleware, operating systems, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolution
  • Execute repaving activities, change management processes, and disaster recovery procedures
  • Ensure security and vulnerability compliance, including user management and certificate lifecycle management
  • Handle service requests, configuration updates, and audit-related data extracts
  • Develop and maintain Standard Operating Procedures (SOPs) for infrastructure operations
  • Collaborate on cell-based automation improvements and drive continuous service enhancements
Requirements
  • A minimum of 5 years of relevant experience
  • At least one year of experience leading and managing teams
  • Experience working with cloud platforms such as GCP, AWS, and Azure
  • Proficiency in OS administration across Windows and Linux environments
  • Proficiency in automation tools such as Ansible, Terraform, Python, and Bash
  • Strong knowledge of observability tools such as ELK and Grafana
  • Solid understanding of incident management processes
  • Experience using GitHub for version control and collaborative development
  • Experience in security hardening, vulnerability management, and compliance practices
  • Excellent problem-solving, communication, and collaboration skills
  • Familiarity with disaster recovery and operational recovery processes
  • English level B2 or higher, with strong written and verbal communication skills