Skip To Main Content
backBack to Search

Senior Site Reliability Engineer

Hybrid in Brazil: São Paulo
Site Reliability Engineering& 4 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are looking for a Senior Site Reliability Engineer to join our team. You'll work closely with backend engineers, product managers, and other SREs to identify risk before it becomes an incident — and when incidents do happen, you'll help lead the response and the follow-up.

Responsibilities
  • Design, build, and maintain monitoring, alerting, and observability tooling, including metrics, logs, and traces, for debit card services
  • Build and maintain dashboards that surface availability, latency, error rates, and other key reliability signals to engineering and leadership
  • Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical debit card flows, such as authorization, settlement, card issuance, and disputes
  • Monitor system availability and reliability on an ongoing basis, proactively identifying degradation trends before they cause customer-facing impact
  • Participate in on-call rotations, leading or supporting incident response, root cause analysis, and blameless postmortems
  • Partner with product and backend engineering teams to review architecture for reliability, scalability, and fault tolerance
  • Automate manual operational work through tooling and scripting to reduce repetitive tasks and operational toil
  • Conduct capacity planning and performance testing to ensure systems scale effectively with transaction volume
  • Improve deployment safety through canary releases, rollback automation, and progressive delivery practices
  • Contribute to and enforce reliability best practices, runbooks, and operational documentation across the team
Requirements
  • A minimum of 3 years of relevant experience as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure Engineer
  • Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, Datadog, or M3/Uber's internal metrics stack
  • Strong understanding of SLOs, SLIs, error budgets, and other reliability engineering principles
  • Proficiency in at least one programming language commonly used for tooling and automation, such as Go, Python, or Java
  • Experience with distributed systems and a solid understanding of failure modes in high-throughput, low-latency environments
  • Familiarity with container orchestration and infrastructure, such as Kubernetes and Docker, along with cloud or on-prem infrastructure at scale
  • Experience with incident management processes, including on-call response, root cause analysis, and postmortems
  • Strong scripting and automation skills for reducing operational toil, using tools such as Bash, Python, or similar languages
  • Working knowledge of CI/CD pipelines and safe deployment practices, including canary, blue-green, and rollback strategies
  • Excellent communication skills, with the ability to translate system health data into clear insights for both engineers and non-technical stakeholders
  • Excellent English communication skills (B2 level or higher)
Nice to have
  • Experience in payments, fintech, or other high-compliance, transaction-critical environments
  • Familiarity with PCI-DSS or other financial services compliance and security requirements
  • Experience with chaos engineering or fault-injection testing
  • Background in database reliability, including query performance, replication, and failover, for transactional systems
  • Experience building or maintaining internal tooling and platforms for observability at scale