backНазад до пошуку

Senior Data Reliability Engineer

Віддалений* формат співпраці з території - Україна

*Ви можете працювати віддалено з території країни (або країн), для яких відкрита ця позиція.

Data Software Engineering
hot

EPAM's Operational Intelligence practice is expanding from classic infrastructure and application observability into Data & AI Reliability Engineering. We are looking for a Data Reliability Engineer who will apply SRE principles to data pipelines and data platforms — moving our clients from reactive firefighting to proactive management of data quality, so that data stays accurate, complete, fresh and available end-to-end.

This is an engineering role, not an L1/L2 support role: you will build detection, automation and prevention rather than sit in a 24/7 on-call rotation. You will work across client engagements and practice-level initiatives — reusable data quality frameworks, accelerators, and internal enablement.

What You’ll Get

  • A practice that is actively building a new capability — your work becomes the standard, not a copy of someone else's runbook
  • Engineering focus without 24/7 on-call rotation
  • Internal certification and enablement tracks (Databricks, AI/LLM, Data Observability learning path in Learn)
  • Cross-client exposure: enterprise-scale retail, financial services and manufacturing accounts
Чим ви будете займатися у цій ролі
  • Design and implement end-to-end data observability across pipelines and data platforms, covering the core data quality pillars: freshness, volume, schema, completeness and accuracy
  • Embed automated data quality checks into CI/CD pipelines and orchestrators (e.g., dbt tests, Great Expectations, native platform checks)
  • Configure anomaly detection — including dynamic and ML-based thresholds — for data drift, volume anomalies and pipeline failures
  • Define, measure and report Data SLIs, SLOs and error budgets together with business and data product stakeholders
  • Perform root cause analysis by tracing data lineage upstream to the exact origin of a failure; drive problem management so incidents do not recur
  • Design architectural guardrails for data pipelines: circuit breakers, dead-letter queues, retry and rollback mechanisms, self-healing patterns
  • Implement monitoring and alerting as code (Terraform / GitOps) instead of manual UI configuration
  • Reduce alert noise through event correlation, tagging standards and actionable alert design
  • Build reusable data quality and observability frameworks, templates and accelerators adopted by multiple data product teams
  • Lead post-incident reviews and translate learnings into systemic platform improvements
  • Contribute to presales activities, client-facing assessments and internal training materials for the practice
Навички
  • 4+ years in SRE, DevOps, Data Engineering or Data Platform Operations, with at least 1–2 years focused on data platforms or data pipelines
  • Solid SRE fundamentals: Golden Signals, SLI/SLO definition and calculation, error budgets and burn rate, incident lifecycle, ITIL basics
  • Strong SQL and practical Python for automation and validation scripting
  • Hands-on production experience with at least one cloud platform: Azure, AWS or GCP
  • Experience with at least one modern data platform and orchestration stack — Databricks (preferred), Snowflake, Spark, Airflow, Azure Data Factory, dbt
  • Working knowledge of observability tooling: New Relic, Datadog, Splunk, Elastic Stack, Grafana / OpenTelemetry (any two or more)
  • Infrastructure as Code with Terraform (Ansible is a plus) and CI/CD experience (Azure DevOps, GitLab CI, GitHub Actions)
  • Experience with incident and change management tooling: ServiceNow, PagerDuty or equivalent
  • Ability to troubleshoot complex distributed data issues under SLA pressure
  • B2+ English — the role is client-facing and requires clear written and spoken technical communication
Буде перевагою
  • Databricks certification (Data Engineer Associate / Professional) or equivalent cloud data certification
  • Hands-on experience with dedicated data observability platforms (Monte Carlo, Soda, Anomalo, Great Expectations)
  • Data catalog and lineage tooling (Unity Catalog, OpenLineage, Collibra)
  • FinOps: cloud, platform and telemetry cost optimization
  • Exposure to AI/ML pipeline monitoring or AI Reliability Engineering
  • Power BI or another BI layer, from a monitoring and reliability perspective
  • Event correlation / AIOps experience (New Relic Decisions, IBM NOI, ServiceNow ITOM)
  • Mentoring or team lead experience