Agent Evaluation Engineer — Build-Time Framework & Deployment Gates
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.
- Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworks
- Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
- Develop test harnesses for multi-turn conversational simulations and context-retention scoring
- Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
- Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
- Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
- Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback
- Convert production incidents into reusable regression cases for continuous quality improvement
- Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows
- 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
- Proven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
- Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
- Practical knowledge of LangGraph or similar agent orchestration frameworks
- Strong background in designing simulation-based evaluation strategies and conversation-level tests
- Experience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)
- Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
- Background in transforming production failures into build-time regression tests for agent workflows
