Skip To Main Content
backBack to Search

Lead Site Reliability Engineer

Hybrid in Brazil: São Paulo
Site Reliability Engineering& 4 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are seeking a Lead Site Reliability Engineer to become part of our team. In this position, you'll partner closely with backend developers, product managers, and fellow SREs to spot potential risks before they escalate into full-blown incidents — and when problems do surface, you'll take charge of driving the response and subsequent analysis.

Responsibilities
  • Architect and sustain monitoring, alerting, and observability systems, covering metrics, logs, and traces, to support debit card services
  • Develop and manage dashboards that display availability, latency, error rates, and other critical reliability indicators for both engineering teams and leadership
  • Establish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets tied to essential debit card processes, such as authorization, settlement, card issuance, and dispute handling
  • Continuously oversee system health and stability, catching signs of degradation early to prevent negative effects on customers
  • Take part in on-call schedules, driving or contributing to incident resolution, root cause investigation, and blameless retrospectives
  • Team up with product and engineering groups to assess architecture from the standpoint of reliability, scalability, and resilience to failure
  • Streamline repetitive operational tasks through automation and custom tooling to cut down on manual toil
  • Perform capacity forecasting and load testing to confirm systems can handle growing transaction demands
  • Enhance the safety of releases by leveraging canary deployments, automated rollbacks, and progressive rollout techniques
  • Help develop and uphold reliability standards, runbooks, and operational documentation throughout the team
Requirements
  • At least 5 years of relevant experience working as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure Engineer
  • A minimum of one year of experience guiding and overseeing teams
  • Practical experience using observability and monitoring platforms such as Grafana, Prometheus, Datadog, or comparable internal metrics systems
  • Solid grasp of SLOs, SLIs, error budgets, and broader reliability engineering concepts
  • Command of at least one programming language typically used for automation and tooling, such as Go, Python, or Java
  • Background working with distributed systems and awareness of common failure patterns in high-volume, low-latency settings
  • Working familiarity with container orchestration tools and infrastructure, including Kubernetes and Docker, plus experience with cloud or on-premises environments at scale
  • Exposure to incident management workflows, covering on-call duties, root cause diagnosis, and postmortem reviews
  • Solid scripting and automation capabilities aimed at minimizing operational toil, using languages such as Bash or Python
  • Practical understanding of CI/CD pipelines and secure deployment techniques, including canary releases, blue-green deployments, and rollback approaches
  • Strong communication abilities, capable of turning system performance data into understandable insights for technical and non-technical audiences alike
  • Excellent English communication skills (B2 level or higher)
Nice to have
  • Background working within payments, fintech, or other tightly regulated, transaction-sensitive industries
  • Understanding of PCI-DSS or similar financial industry compliance and security standards
  • Exposure to chaos engineering or fault-injection methodologies for resilience testing
  • Experience with database reliability topics, such as query optimization, replication, and failover, within transactional systems
  • Background developing or supporting internal tools and platforms designed for large-scale observability