Senior Site Reliability Engineer
Hybrid in Poland: Warsaw
Site Reliability Engineering& 5 others
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are seeking a Senior Site Reliability Engineer to own the reliability, observability, and operational health of production AI systems, bridging the gap between deployment and long-term operability while embedding cost, security, and quality discipline into every solution's lifecycle.
Responsibilities
- Own deployment end-to-end, including infrastructure as code, CI/CD pipelines, and environment management on Azure, ensuring every environment is rebuildable from source
- Build LLM-aware observability with traces on every model call, production quality signals such as eval sampling, drift detection, and guardrail-trigger rates, plus cost and latency dashboards
- Define and defend SLOs covering availability, latency, and quality objectives per solution, balancing delivery speed against stability with data-driven error budgets
- Run incident management, including on-call models, runbooks written before incidents occur, and blameless postmortems afterward
- Manage the cost of intelligence by monitoring token economics per solution, wiring in budgets and alerts, and conducting capacity planning proactively
- Keep the security posture current through patching, secret rotation, access reviews, and audit readiness across the solution's entire lifecycle
- Shape operability requirements before handover, ensuring they land in the pod's definition of done, and run hypercare jointly with sign-off on what will be operated
- Feed operational patterns, failure modes, and cost learnings back to the pods and the Architect
Requirements
- 5+ years of experience operating cloud production systems, with a track record in scaling, defining SLOs, managing on-call rotations, and automating manual work away
- Expertise in Azure IaaS/PaaS operations, infrastructure as code (Terraform/Bicep), and CI/CD tooling
- Proficiency in observability stacks with LLM tracing, container orchestration, and Python/Bash automation
- Knowledge of FinOps basics for AI workloads
- Familiarity with daily AI use in operations work, including incident triage, runbook drafting, log analysis, and automation code
