Overview

We are looking for a Senior ML Evaluation Engineer to help design, implement, and operationalize evaluation frameworks for enterprise AI systems. In this role, you will define quality standards for AI agents and machine learning solutions, build scalable evaluation pipelines, and integrate automated quality gates into CI/CD processes. You will work closely with AI engineers, platform teams, and governance stakeholders to ensure reliable, measurable, and production ready AI outcomes.
Project Overview:
The project focuses on establishing enterprise grade evaluation standards for AI agents and machine learning systems. The platform enables automated quality assessment, production monitoring, and governance through evaluation frameworks, observability data, and continuous validation processes.

Responsibilities:
  • Develop and maintain evaluation frameworks for AI agents and machine learning solutions
  • Design LLM as judge evaluation methodologies using built in helpfulness and correctness evaluators
  • Build custom Python based evaluators to perform deterministic quality and compliance checks
  • Define enterprise evaluation standards, including mandatory assessment dimensions and pass or fail criteria
  • Implement evaluation workflows across response, tool invocation, and end to end session levels
  • Integrate observability telemetry and OpenTelemetry spans into evaluation pipelines
  • Design and maintain CI/CD quality gates for machine learning models and AI agents
  • Collaborate with AI platform teams to improve evaluation coverage, automation, and reporting
  • Analyze evaluation results and provide recommendations to improve agent reliability and performance
  • Support production monitoring strategies and continuous quality verification processes
  • Contribute to AI governance initiatives and best practices for model and agent evaluation
Required Qualifications:
  • 5+ years of experience in machine learning engineering or AI platform engineering
  • Hands on experience designing and implementing LLM evaluation frameworks
  • Experience creating custom evaluators for deterministic quality validation and policy enforcement
  • Experience building CI/CD deployment gates for machine learning models, AI applications, or agent based systems
  • Strong Python development skills
  • Experience working with AI quality metrics, automated testing methodologies, and evaluation pipelines
  • Understanding of agent based architectures and modern AI application development practices
  • Experience collaborating with engineering teams on quality assurance and governance initiatives
  • Strong analytical and problem solving skills
  • Effective written and verbal communication skills
Nice To Have:
  • Hands on experience with AWS Agent Evaluation APIs, including evaluation execution and results analysis
  • Experience integrating AWS Bedrock Guardrails for PII detection and evaluation workflows
  • Experience using CloudWatch metrics for online evaluation monitoring and reporting
  • Knowledge of observability frameworks and OpenTelemetry based monitoring
  • Experience with enterprise AI governance and compliance programs
  • Experience evaluating production AI agents and large scale machine learning systems
Note:

✨ Our intelligent job search engine discovered this job and republished it for your convenience.
Please be aware that the job information may be incorrect or incomplete. The job announcement remains the property of its original publisher. To view the original job and its full details, please visit the job's URL on the owner’s page.

Please clearly mention that you have heard of this job opportunity on https://ijob.am.