Socium - Teams Done Differently
AI Ops Engineer
On-siteAbu Dhabi (On-site)Posted last wk.
About the role
Location: Abu Dhabi (Onsite)
Contract Duration: Initial 12 months extendable contract
Key Responsibilities
- Build and maintain end-to-end MLOps/LLMOps pipelines using DevOps best practices.
- Deploy, monitor, version, and roll back AI/ML models in production.
- Build scalable and GPU-optimized workflows for AI training and inference.
- Operate AI models and applications as reliable production services.
- Support RAG and agentic AI workflows in production environments.
- Monitor AI inference usage, including TPM/quota, performance, and consumption/cost.
- Implement data preprocessing, feature engineering, and model evaluation workflows as required.
- Support fine-tuning and deployment of pre-trained LLM, NLP, and computer-vision models.
- Ensure AI solutions are reproducible, governed, documented, and compliant.
- Contribute to AI infrastructure strategy and responsible/ethical AI practices.
Requirements
- Experience in DevOps, Cloud, SRE, Platform Engineering, or Infrastructure.
- Experience transitioning into or working with AI Ops, MLOps, or LLMOps.
- Hands-on experience with Kubernetes, CI/CD, containers, and cloud environments.
- Experience deploying and operating AI/ML models in production.
- Familiarity with GPU-based inference and AI/ML workloads.
- Understanding of model monitoring, versioning, deployment, and rollback.
- Familiarity with LLMs, RAG, agentic workflows, or other production AI applications.
- Strong automation and troubleshooting skills.