Senior Site Reliability Engineer
Hybrid in Brazil: São Paulo
Site Reliability Engineering& 4 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are looking for a Senior Site Reliability Engineer to join our team. You'll work closely with backend engineers, product managers, and other SREs to identify risk before it becomes an incident — and when incidents do happen, you'll help lead the response and the follow-up.
Responsibilities
- Design, build, and maintain monitoring, alerting, and observability tooling, including metrics, logs, and traces, for debit card services
- Build and maintain dashboards that surface availability, latency, error rates, and other key reliability signals to engineering and leadership
- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical debit card flows, such as authorization, settlement, card issuance, and disputes
- Monitor system availability and reliability on an ongoing basis, proactively identifying degradation trends before they cause customer-facing impact
- Participate in on-call rotations, leading or supporting incident response, root cause analysis, and blameless postmortems
- Partner with product and backend engineering teams to review architecture for reliability, scalability, and fault tolerance
- Automate manual operational work through tooling and scripting to reduce repetitive tasks and operational toil
- Conduct capacity planning and performance testing to ensure systems scale effectively with transaction volume
- Improve deployment safety through canary releases, rollback automation, and progressive delivery practices
- Contribute to and enforce reliability best practices, runbooks, and operational documentation across the team
Requirements
- A minimum of 3 years of relevant experience as a Site Reliability Engineer, DevOps Engineer, or Production/Infrastructure Engineer
- Hands-on experience with monitoring and observability tools such as Grafana, Prometheus, Datadog, or M3/Uber's internal metrics stack
- Strong understanding of SLOs, SLIs, error budgets, and other reliability engineering principles
- Proficiency in at least one programming language commonly used for tooling and automation, such as Go, Python, or Java
- Experience with distributed systems and a solid understanding of failure modes in high-throughput, low-latency environments
- Familiarity with container orchestration and infrastructure, such as Kubernetes and Docker, along with cloud or on-prem infrastructure at scale
- Experience with incident management processes, including on-call response, root cause analysis, and postmortems
- Strong scripting and automation skills for reducing operational toil, using tools such as Bash, Python, or similar languages
- Working knowledge of CI/CD pipelines and safe deployment practices, including canary, blue-green, and rollback strategies
- Excellent communication skills, with the ability to translate system health data into clear insights for both engineers and non-technical stakeholders
- Excellent English communication skills (B2 level or higher)
Nice to have
- Experience in payments, fintech, or other high-compliance, transaction-critical environments
- Familiarity with PCI-DSS or other financial services compliance and security requirements
- Experience with chaos engineering or fault-injection testing
- Background in database reliability, including query performance, replication, and failover, for transactional systems
- Experience building or maintaining internal tooling and platforms for observability at scale
