Senior ML/Evaluation Engineer
Remote in Czech Republic, Slovakia
AI Native Engineering
Looking for something else?
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobChoose an option
We are building an Enterprise Agent Development Platform, a production-grade, cloud-native ecosystem that enables engineering teams to define, orchestrate, deploy, and observe AI agents at scale. The platform standardizes agent development across the organization. The initiative spans agent framework design, runtime architecture, marketplace integration, CI/CD automation, and enterprise-grade observability, reducing agent development from months to days while enforcing consistent security, quality, and governance standards.
Responsibilities
- AWS AgentCore Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
- LLM-as-judge evaluator design (built-in AgentCore evaluators — helpfulness, correctness)
- Custom code-based Lambda evaluators (Python — deterministic checks)
- Evaluation levels (TRACE for per-response, TOOL_CALL for per-invocation, SESSION for workflow)
- OTel spans from AWS AgentCore Observability as evaluation input
- Enterprise evaluation standard authoring (mandatory dimensions, pass/fail criteria)
Requirements
- 5+ years ML engineering or AI platform engineering
- LLM evaluation framework design and implementation
- Custom evaluator implementation for deterministic quality checks
- CI/CD deployment gate design for ML model or agent quality
Nice to have
- AWS AgentCore Evaluation API hands-on (CreateEvaluation, GetEvaluationResult)
- AWS Bedrock Guardrails for PII detection evaluator integration
- CloudWatch metrics output from AgentCore Evaluation for online mode
