Senior QA / ML Tester
Find a vacancy that works for you. Send us your CV to receive a personalized offer.
Find me a jobWe are looking for a Senior QA / ML Tester to join the AI Platform team and take ownership of quality assurance for the Agent Evaluation Framework built on AWS AgentCore. This role involves designing, implementing, and maintaining a functional test suite that validates the correctness of AI agent evaluation pipelines — covering on-demand integration testing, online sampling accuracy, and multi-evaluator execution — and delivering a feasibility assessment for non-AgentCore runtime evaluation scenarios. This is a hands-on, production-focused role at the intersection of software quality engineering and AI/ML system testing, operating within an Agile delivery team and contributing to the reliability of enterprise-grade agentic AI infrastructure.
- Design and implement a functional test suite for the AWS AgentCore Evaluation API using pytest, covering known-good / known-bad session pair validation, multi-evaluator execution correctness, and edge case handling
- Develop and maintain integration tests for on-demand evaluation mode, integrated into the CI/CD pipeline with automated execution on each build
- Validate online mode sampling accuracy, design test scenarios, define acceptance criteria, and report deviations with reproducible evidence
- Conduct and document a feasibility assessment for non-AgentCore runtime evaluation: analyze alternative runtimes, define evaluation methodology, and deliver a structured findings report
- Test OpenTelemetry trace-based evaluation inputs and validate ADOT trace ingestion, trace structure correctness, and evaluator input integrity
- Collaborate with platform engineers to clarify evaluation contracts, reproduce defects, and align on quality gates
- Maintain test documentation, including test plans, test reports, defect logs, and evaluation feasibility artifacts in Confluence/Jira
- Participate in Agile ceremonies, including sprint planning, daily standups, demos, and retrospectives
- Contribute to EngX practices such as code review of test scripts, CI/CD pipeline integration, and test coverage reporting
- 5+ years of production experience in QA automation or ML/AI system testing
- Proven experience testing AI/LLM systems, including evaluation pipelines, model outputs, or agent behavior validation
- Proficiency in Python test automation, including pytest, fixtures, parametrize, and mocking (unittest.mock, moto)
- Knowledge of AWS AgentCore Evaluation API, covering on-demand and online evaluation modes
- Familiarity with OpenTelemetry / ADOT for trace-based evaluation input testing and trace structure validation
- Skills in REST API testing, including request/response validation and authentication (SigV4, bearer tokens)
- Experience with CI/CD integration using GitHub Actions, Jenkins, or equivalent, including test pipeline configuration
- Background in test data management, including known-good / known-bad session pair design and synthetic trace generation
- Expertise in functional and integration test design for AI/ML evaluation pipelines
- Competency in defect lifecycle management, including Jira, reproducible bug reports, and root cause analysis
- Understanding of LLM/agent evaluation concepts, such as correctness scoring, sampling strategies, and evaluator chaining
- Ability to work independently after onboarding, manage own tasks, report status, and escalate blockers proactively
- Strong analytical skills to define test scenarios from ambiguous or evolving specifications
- English B2+ level, written and verbal, for daily collaboration with distributed teams
- Hands-on experience with AWS AgentCore Evaluation API or AWS Bedrock testing
- Experience testing OpenTelemetry / distributed tracing pipelines
- Familiarity with multi-evaluator execution patterns and correctness validation strategies
- Experience writing feasibility assessments or technical reports for stakeholders
- Knowledge of agentic AI frameworks (LangGraph, Strands Agents) sufficient to understand evaluation contracts
- ISTQB CT-AI certification or equivalent AI testing qualification
- Experience with AI Ready / AI Practitioner practices at EPAM (prompt engineering, AI-assisted test design)
