Senior Observability Engineer
Office in Ecuador: Quito
Site Reliability Engineering& 10 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are seeking a Senior Observability Engineer to implement, maintain, and optimize the bank's monitoring platforms, driving log integrations, new metrics, traces, and digital experience monitoring. In this role, you will collaborate on the development of L1-L3 dashboards and the automation of signals for incident detection and remediation, working closely with Champions, SRE, and DevSecOps teams to ensure critical service coverage and enable data-driven decisions.
Responsibilities
- Configure, deploy, and support the observability infrastructure according to design specifications, including agents/collectors, ingestion pipelines, storage, and L1-L3 dashboards
- Collaborate with champions to translate their tribes' needs into custom extensions or integrations and maintain dashboard coverage/fidelity
- Coordinate and manage the execution of observability architecture designs, including service maps, tagging, security, and compliance, with IT departments and vendors
- Work with observability architects to identify gaps between needs and available tools
- Monitor champions' compliance with observability regulations, standards, and best practices
- Participate in war rooms and post-incident reviews to accelerate root cause analysis and document preventative actions
- Train champions on instrumentation and the effective creation/use of dashboards and alerts
- Optimize telemetry cardinality, retention, and costs under FinOps principles, while maintaining data quality
Requirements
- More than 5 years of experience in observability engineering
- Expertise in Dynatrace, Grafana, and Zabbix for observability instrumentation
- Skills in modeling metrics, logs, and traces for SLO
- Knowledge of designing data ingestion and transformation pipelines using Kafka and Fluent Bit
- Proficiency in scripting and automation with Python, Bash, Terraform, and Ansible for deployment and runbooks
- Background in performance management in Kubernetes/EKS, cloud services (AWS, Azure, GCP), and databases
- Understanding of infrastructure and networking
- Competency in advanced troubleshooting and root cause analysis using AIOps and distributed tracing techniques
- Ability to translate technical data into actionable insights for product teams
- Familiarity with Agile collaboration (Scrum/Kanban) and tools such as Jira and Confluence for work tracking
- Capability to communicate clearly in war rooms and post-mortems, including documentation of lessons learned
