Overview
We are looking for a Senior AI Evaluation Engineer to design and implement enterprise grade evaluation frameworks for agent based AI systems. In this role, you will build automated quality assessment capabilities, define deployment gate strategies, and create scalable evaluation methodologies that improve the reliability, accuracy, and safety of AI driven solutions throughout the development lifecycle.
Project Overview:
The project focuses on establishing a comprehensive evaluation and quality assurance platform for agent based AI systems. The solution provides automated testing, deployment validation, production feedback integration, and continuous quality monitoring to ensure high standards for agent performance and user experience.
- Design and develop build time evaluation frameworks for LangGraph based agent systems
- Create automated test harnesses for graph level and node level validation
- Design evaluation strategies that combine deterministic grading and LLM as judge methodologies
- Implement multi layer evaluation frameworks covering tool selection accuracy, execution trajectory quality, reasoning effectiveness, and output quality
- Develop multi turn conversation simulations and context retention scoring mechanisms
- Implement multi trial reliability testing methodologies, including pass at k and pass power k approaches
- Design and maintain CI/CD deployment gates that validate quality metrics before production releases
- Build staging validation, shadow mode comparison, and controlled rollout evaluation workflows
- Integrate production evaluation feedback into build time testing frameworks to improve quality coverage
- Collaborate with platform and engineering teams to establish evaluation standards, thresholds, and governance practices
- Analyze evaluation results and provide recommendations for improving agent reliability and performance
- Contribute to technical documentation, testing standards, and quality engineering best practices
- 4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for machine learning, large language model, or agent based systems
- Hands on experience designing multi layer evaluation frameworks that combine deterministic and LLM as judge grading approaches
- Experience implementing automated quality gates that can block deployments based on predefined metric thresholds
- Experience working with LangGraph or a comparable agent orchestration framework
- Strong understanding of agent behavior evaluation, workflow validation, and AI quality measurement techniques
- Experience designing scalable testing and validation processes for production AI systems
- Strong Python development experience
- Knowledge of CI/CD practices, deployment automation, and release governance
- Experience analyzing evaluation data and translating findings into platform improvements
- Strong communication and collaboration skills
- Hands on experience with AWS AgentCore Evaluations, including evaluation execution and custom evaluator development
- Experience designing and operating shadow mode or canary deployment strategies for machine learning or AI systems
- Experience creating automated feedback loops that convert production incidents into regression test scenarios
- Knowledge of production observability, monitoring, and evaluation pipelines
- Experience with enterprise AI governance and quality assurance programs
- Understanding of agent observability and telemetry driven quality improvement processes
✨ Our intelligent job search engine discovered this job and republished it for your convenience.
Please be aware that the job information may be incorrect or incomplete. The job announcement remains the property of its original publisher. To view the original job and its full details, please visit the job's URL on the owner’s page.
Please clearly mention that you have heard of this job opportunity on https://ijob.am.
