Apply with hirly
Senior/Staff AI Evaluation Engineer
Haparz Pvt Ltd · India
Upload your resume to see how well you match this job — free, in seconds, no account needed.
Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.
Job Description Role: Senior/Staff AI Evaluation Engineer Experience: 6+ Years Location: Remote (Pan India) Mode: Hyparz Payroll About the Role We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms . This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML . You will build automated evaluation frameworks, benchmark AI behavior, validate LLM/RAG/agentic workflows, and establish measurable quality standards for production AI systems. What You'll Do
- Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications .
- Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
- Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
- Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
- Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
- Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
- Work extensively with LangChain and LangGraph , including sub-graphs and conditional routing.
- Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
- Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
- Validate AI systems against safety, security, privacy, and Responsible AI expectations.
- Integrate evaluation and regression testing into GitHub-based CI/CD workflows .
- Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
- Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
- Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
- Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches. What We're Looking For
- 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
- Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications .
- Strong Python development and automation experience.
- Deep hands-on knowledge of LangChain, LangGraph, and LangSmith .
- Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows .
- Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
- Experience building golden datasets, benchmark datasets, or structured AI test datasets.
- Experience with adversarial testing and AI safety evaluation.
- Practical understanding of security, privacy, and Responsible AI evaluation .
- Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
- Experience with Claude Code or comparable AI-assisted development tools .
- Experience with AWS and Amazon Bedrock .
- Strong REST API, JSON, and SQL knowledge.
- Strong experience in automated, regression, integration, and end-to-end testing.
- Excellent analytical, troubleshooting, communication, and problem-solving abilities Skills:- Large Language Models (LLM), Retrieval Augmented Generation (RAG), AI Evaluation, Agentic AI and AWS Bedrock