Back to roles

Socium - Teams Done Differently

AI Evaluation Engineer

On-siteAbu Dhabi (On-site)Posted last wk.
Fast Track available

About the role

Location: Abu Dhabi (Onsite)

Contract Duration: Initial 12 months extendable contract

Key Responsibilities

  • Design and implement AI/LLM evaluation frameworks, benchmarks, and test datasets.
  • Develop metrics and automated testing for accuracy, relevance, groundedness, hallucination, safety, and reliability.
  • Evaluate LLMs, RAG pipelines, agents, and multimodal AI applications.
  • Build regression testing and quality gates for AI/ML releases.
  • Analyze evaluation results and identify opportunities for model and system improvements.
  • Develop scalable evaluation pipelines and services using cloud-native and MLOps practices.
  • Collaborate with ML engineers, data scientists, software engineers, and product teams.

Requirements

  • 4–6 years of experience in AI/ML Engineering, Data Science, NLP, or a related field.
  • Strong Python and software engineering skills.
  • Hands-on experience with LLMs, Generative AI, RAG, embeddings, and vector databases.
  • Experience with AI/LLM evaluation, benchmarking, model testing, or quality measurement.
  • Familiarity with LLM-as-a-Judge, RAG evaluation, prompt testing, and human-in-the-loop evaluation.
  • Experience with AWS and/or Azure, Docker, Kubernetes, MLflow, and CI/CD.
  • Strong SQL skills and experience with APIs and modern AI/ML architectures.

Already applied?

Track every application and share your Digital Profile.

Open your space