backНазад до пошуку

Senior AI Reliability Engineer

Віддалений* формат співпраці з території - Україна

*Ви можете працювати віддалено з території країни (або країн), для яких відкрита ця позиція.

Operational Intelligence

EPAM's Operational Intelligence practice is developing a new capability called AI Reliability Engineering (AIRE), which applies SRE principles and cloud-native practices throughout the lifecycle of production AI/ML and LLM systems. As clients transition their GenAI applications and agentic systems from pilot to production, they find that traditional APM provides no visibility into token latency, cost per request, semantic drift, or hallucinations. This role exists to close that gap.

Your work will involve instrumenting, monitoring, and hardening production AI systems, establishing AI-native service level objectives, and creating accelerators and reference implementations that the practice can reuse across accounts. Note that this is an engineering position rather than an L1/L2 support role, and it does not require a 24/7 on-call rotation.

What You'll Get

  • A brand-new discipline within EPAM where you shape the approach rather than follow an existing playbook
  • Focus on engineering work without 24/7 on-call responsibilities
  • Sponsored certification and training programs (Anthropic/Claude, Databricks, AI & Data Observability learning paths)
  • Exposure to multiple clients and a direct route into presales and solution engineering
Чим ви будете займатися у цій ролі
  • Add AI telemetry to production LLM, RAG, and agentic applications using OpenTelemetry and APM-native AI monitoring tools
  • Deploy distributed tracing across multi-model chains, agent workflows, and retrieval-augmented generation pipelines to identify systemic latency and failure points
  • Establish and track AI-native SLIs and SLOs, including time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, and contextual accuracy
  • Implement structured semantic logging and prompt/response monitoring to support quality analysis
  • Create and maintain evaluation loops for output quality and safety using golden sets, LLM-as-a-judge methods, and Ragas/DeepEval-style frameworks, integrating them into CI/CD and runtime environments
  • Monitor token-based cloud spend, model API rate limits, and quota usage while driving AI cost optimization efforts
  • Set up AI gateways to manage API load balancing, failover, and fallback models across multiple LLM providers
  • Build guardrails covering prompt injection and jailbreak filtering, output compliance, and bias and safety constraints
  • Develop detection, triage, restoration, and problem management workflows for AI incidents, incorporating autonomous AI agents into root cause analysis to parse logs, generate hypotheses, and correlate state changes
  • Enable rollback, canary, and fail-safe patterns for model, prompt, and configuration releases, while maintaining reproducibility through versioning of data, code, prompts, and models
  • Develop practice accelerators, reference architectures, and internal training materials, and contribute to presales activities and client assessments
Навички
  • 4+ years of experience in SRE, DevOps, platform, or observability engineering, with hands-on exposure to production AI/ML or LLM workloads
  • Strong grasp of SRE fundamentals, including Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, and ITIL basics
  • Strong Python skills for building instrumentation, automation, and evaluation tools
  • Hands-on experience with OpenTelemetry and at least one APM/observability platform such as New Relic, Datadog, Grafana LGTM stack, Splunk, or Elastic
  • Production experience with at least one cloud platform (Azure preferred, AWS or GCP also acceptable) and Kubernetes
  • Solid understanding of LLM application architecture, including prompts, embeddings and vector stores, RAG, and agent orchestration tools like LangChain/LangGraph or similar
  • Awareness of MLOps concepts, including model lifecycle (training vs. inference), model endpoints, containerization, and deployment/rollback patterns
  • Experience with Infrastructure as Code using Terraform, along with CI/CD tools such as Azure DevOps, GitLab CI, or GitHub Actions
  • B2+ English proficiency, since the role involves direct client interaction and requires clear technical communication in both writing and speech
Буде перевагою
  • Familiarity with AI-specific observability and evaluation tools such as Traceloop/OpenLLMetry, Langfuse, Arize Phoenix, Ragas, DeepEval, or MLflow
  • Experience with distributed inference serving at scale, including vLLM, KServe, Ray Serve, Kubernetes-native LLM orchestration, and GPU capacity planning
  • Knowledge of AI security practices, including the OWASP LLM Top 10, prompt injection defense, and guardrail frameworks like NeMo Guardrails or Llama Guard
  • Experience with Databricks (including Mosaic AI/MLflow) or Azure AI Foundry
  • Relevant certifications such as Anthropic Claude, Azure AI Engineer, AWS ML Specialty, or Databricks GenAI
  • Background in FinOps for AI workloads, particularly token and GPU cost modeling
  • Experience in Data Reliability Engineering, covering data quality and pipeline SLOs, since AI reliability depends on data reliability
  • Prior mentoring or team lead experience