Socium - Teams Done Differently
AI Evaluation Engineer
On-siteAbu Dhabi (On-site)Posted last wk.
About the role
Location: Abu Dhabi (Onsite)
Contract Duration: Initial 12 months extendable contract
Key Responsibilities
- Design and implement AI/LLM evaluation frameworks, benchmarks, and test datasets.
- Develop metrics and automated testing for accuracy, relevance, groundedness, hallucination, safety, and reliability.
- Evaluate LLMs, RAG pipelines, agents, and multimodal AI applications.
- Build regression testing and quality gates for AI/ML releases.
- Analyze evaluation results and identify opportunities for model and system improvements.
- Develop scalable evaluation pipelines and services using cloud-native and MLOps practices.
- Collaborate with ML engineers, data scientists, software engineers, and product teams.
Requirements
- 4–6 years of experience in AI/ML Engineering, Data Science, NLP, or a related field.
- Strong Python and software engineering skills.
- Hands-on experience with LLMs, Generative AI, RAG, embeddings, and vector databases.
- Experience with AI/LLM evaluation, benchmarking, model testing, or quality measurement.
- Familiarity with LLM-as-a-Judge, RAG evaluation, prompt testing, and human-in-the-loop evaluation.
- Experience with AWS and/or Azure, Docker, Kubernetes, MLflow, and CI/CD.
- Strong SQL skills and experience with APIs and modern AI/ML architectures.