Skip To Main Content
backBack to Search

Senior Site Reliability Engineer

Hybrid in Poland: Warsaw
Site Reliability Engineering& 5 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.

Responsibilities
  • Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
  • Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
  • Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
  • Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
  • Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
  • Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
  • Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
  • Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect
Requirements
  • 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
  • Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
  • Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
  • Knowledge of FinOps basics for AI workloads
  • Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code