Senior Data Reliability Engineer
*Ви можете працювати віддалено з території країни (або країн), для яких відкрита ця позиція.
EPAM's Operational Intelligence practice is expanding from classic infrastructure and application observability into Data & AI Reliability Engineering. We are looking for a Data Reliability Engineer who will apply SRE principles to data pipelines and data platforms — moving our clients from reactive firefighting to proactive management of data quality, so that data stays accurate, complete, fresh and available end-to-end.
This is an engineering role, not an L1/L2 support role: you will build detection, automation and prevention rather than sit in a 24/7 on-call rotation. You will work across client engagements and practice-level initiatives — reusable data quality frameworks, accelerators, and internal enablement.
What You’ll Get
- A practice that is actively building a new capability — your work becomes the standard, not a copy of someone else's runbook
- Engineering focus without 24/7 on-call rotation
- Internal certification and enablement tracks (Databricks, AI/LLM, Data Observability learning path in Learn)
- Cross-client exposure: enterprise-scale retail, financial services and manufacturing accounts
- Design and implement end-to-end data observability across pipelines and data platforms, covering the core data quality pillars: freshness, volume, schema, completeness and accuracy
- Embed automated data quality checks into CI/CD pipelines and orchestrators (e.g., dbt tests, Great Expectations, native platform checks)
- Configure anomaly detection — including dynamic and ML-based thresholds — for data drift, volume anomalies and pipeline failures
- Define, measure and report Data SLIs, SLOs and error budgets together with business and data product stakeholders
- Perform root cause analysis by tracing data lineage upstream to the exact origin of a failure; drive problem management so incidents do not recur
- Design architectural guardrails for data pipelines: circuit breakers, dead-letter queues, retry and rollback mechanisms, self-healing patterns
- Implement monitoring and alerting as code (Terraform / GitOps) instead of manual UI configuration
- Reduce alert noise through event correlation, tagging standards and actionable alert design
- Build reusable data quality and observability frameworks, templates and accelerators adopted by multiple data product teams
- Lead post-incident reviews and translate learnings into systemic platform improvements
- Contribute to presales activities, client-facing assessments and internal training materials for the practice
- 4+ years in SRE, DevOps, Data Engineering or Data Platform Operations, with at least 1–2 years focused on data platforms or data pipelines
- Solid SRE fundamentals: Golden Signals, SLI/SLO definition and calculation, error budgets and burn rate, incident lifecycle, ITIL basics
- Strong SQL and practical Python for automation and validation scripting
- Hands-on production experience with at least one cloud platform: Azure, AWS or GCP
- Experience with at least one modern data platform and orchestration stack — Databricks (preferred), Snowflake, Spark, Airflow, Azure Data Factory, dbt
- Working knowledge of observability tooling: New Relic, Datadog, Splunk, Elastic Stack, Grafana / OpenTelemetry (any two or more)
- Infrastructure as Code with Terraform (Ansible is a plus) and CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
- Experience with incident and change management tooling: ServiceNow, PagerDuty or equivalent
- Ability to troubleshoot complex distributed data issues under SLA pressure
- B2+ English — the role is client-facing and requires clear written and spoken technical communication
- Databricks certification (Data Engineer Associate / Professional) or equivalent cloud data certification
- Hands-on experience with dedicated data observability platforms (Monte Carlo, Soda, Anomalo, Great Expectations)
- Data catalog and lineage tooling (Unity Catalog, OpenLineage, Collibra)
- FinOps: cloud, platform and telemetry cost optimization
- Exposure to AI/ML pipeline monitoring or AI Reliability Engineering
- Power BI or another BI layer, from a monitoring and reliability perspective
- Event correlation / AIOps experience (New Relic Decisions, IBM NOI, ServiceNow ITOM)
- Mentoring or team lead experience