Skip To Main Content
backBack to Search

Senior ML / Evaluation Engineer

Remote in Portugal
Python.AI& 5 others
Looking for something else?

Find a vacancy that works for you. Send us your CV to receive a personalized offer.

Find me a job

We're looking for a Senior ML / Evaluation Engineer to join our team in Portugal in a fully remote working mode. In this role, you will own the design and implementation of advanced evaluation frameworks for an Enterprise Agent Development Platform—a production-grade, cloud-native ecosystem enabling scalable, secure AI agent deployment. You will create evaluation strategies that combine LLM-as-judge grading with deterministic checks, define enterprise evaluation standards, and implement CI/CD deployment gates to enforce quality metrics prior to release. This position requires strong expertise in ML system testing, evaluation design, and integration into automated pipelines for agentic environments.

Responsibilities
  • Design and implement multi-layer evaluation frameworks for agentic workflows and AI-driven applications
  • Build LLM-as-judge evaluators leveraging AWS AgentCore built-in modules and custom logic for correctness and helpfulness checks
  • Develop deterministic evaluators as AWS Lambda functions for rule-based validation
  • Define enterprise evaluation standards, including mandatory dimensions, scoring criteria, and pass/fail thresholds
  • Implement CI/CD deployment gates using on-demand evaluation modes to enforce quality in automated pipelines
  • Enable online evaluation in production by integrating sampling-based evaluation strategies and PII detection guardrails
  • Incorporate observability signals (OpenTelemetry spans) from AWS AgentCore into grading frameworks for trace-level assessment
  • Generate metrics, logs, and dashboards from evaluation outcomes via CloudWatch or equivalent monitoring platforms
  • Collaborate with platform, orchestration, and DevOps teams to maintain evaluation reliability and scalability
Requirements
  • 5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform development
  • Hands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)
  • Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assurance
  • Proficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)
  • Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoring
Nice to have
  • Familiarity with AWS AgentCore Evaluations API (CreateEvaluation, GetEvaluationResult)
  • Exposure to AWS Bedrock Guardrails for compliance and sensitive data validation
  • Experience integrating evaluation metrics into AWS CloudWatch for monitoring and alerting
  • Knowledge of OTel instrumentation and trace ingestion for quality scoring inputs