Senior Site Reliability Engineer (SRE)
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe're looking for a Senior Site Reliability Engineer (SRE) to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).
This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.
- Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services
- Collaborate with cross-functional teams to embed reliability into application and infrastructure design
- Automate operational tasks to reduce manual toil and improve service performance
- Troubleshoot and resolve infrastructure and application incidents quickly and effectively
- Implement robust monitoring and observability systems to detect and prevent outages
- Plan capacity and scaling strategies to ensure high availability and resiliency
- Contribute to incident postmortems and continuous improvement initiatives
- Support the adoption of SRE best practices across all SDLC stages
- Bachelor’s degree in Computer Science, Engineering or related field
- Proven experience working in cloud environments (AWS, GCP or Azure)
- Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation)
- Proficiency in Python or other scripting language for automation tasks
- Strong understanding of monitoring tools and observability frameworks
- Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab)
- Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes
- Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
- Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies
- Background in DevOps practices and agile delivery frameworks
- Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
