hirly

Apply with hirly

Senior/Staff AI Evaluation Engineer

Haparz Pvt Ltd · India

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

Job Description Role: Senior/Staff AI Evaluation Engineer Experience: 6+ Years Location: Remote (Pan India) Mode: Hyparz Payroll About the Role We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms . This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML . You will build automated evaluation frameworks, benchmark AI behavior, validate LLM/RAG/agentic workflows, and establish measurable quality standards for production AI systems. What You'll Do

  • Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications .
  • Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
  • Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
  • Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
  • Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
  • Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
  • Work extensively with LangChain and LangGraph , including sub-graphs and conditional routing.
  • Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
  • Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
  • Validate AI systems against safety, security, privacy, and Responsible AI expectations.
  • Integrate evaluation and regression testing into GitHub-based CI/CD workflows .
  • Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
  • Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
  • Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
  • Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches. What We're Looking For
  • 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
  • Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications .
  • Strong Python development and automation experience.
  • Deep hands-on knowledge of LangChain, LangGraph, and LangSmith .
  • Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows .
  • Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
  • Experience building golden datasets, benchmark datasets, or structured AI test datasets.
  • Experience with adversarial testing and AI safety evaluation.
  • Practical understanding of security, privacy, and Responsible AI evaluation .
  • Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
  • Experience with Claude Code or comparable AI-assisted development tools .
  • Experience with AWS and Amazon Bedrock .
  • Strong REST API, JSON, and SQL knowledge.
  • Strong experience in automated, regression, integration, and end-to-end testing.
  • Excellent analytical, troubleshooting, communication, and problem-solving abilities Skills:- Large Language Models (LLM), Retrieval Augmented Generation (RAG), AI Evaluation, Agentic AI and AWS Bedrock
Apply: Senior/Staff AI Evaluation Engineer at Haparz Pvt Ltd