Hire Rightt - Executive Search & HR Advisory
Senior LLMOps / AI Platform Engineer
On-siteDubai, United Arab Emirates (On-site)Posted 5 days ago
About the role
Position: Senior LLMOps / AI Platform Engineer
Location: Dubai
Industry: Financial Services
Salary: AED 25 - 30K/ month
Role Overview:
The job holder will be responsible for building, deploying, optimizing, and operating production-grade LLM and Generative AI infrastructure. The role combines LLM inference, GPU optimization, Kubernetes, cloud infrastructure, observability, RAG, and AI platform engineering.
Responsibilities:
- Deploy and operate self-hosted LLMs using vLLM, SGLang, Ollama.
- Optimize LLM inference for latency, throughput, concurrency, GPU memory, KV cache, and cost.
- Manage GPU workloads across multiple NVIDIA GPUs.
- Deploy and maintain AI services on Kubernetes / AWS EKS using Docker and Helm.
- Implement LLM reliability mechanisms including health checks, monitoring, automated recovery, and model restart/refresh strategies.
- Implement observability using Langfuse/LangSmith, OpenTelemetry, Prometheus, and Grafana.
- Deploy and optimize RAG systems, embedding models, and vector databases such as Qdrant, Milvus.
- Support AI agents and workflows built with LangChain and LangGraph.
- Build and maintain CI/CD pipelines for AI services and infrastructure.
- Troubleshoot production issues across LLMs, GPUs, Kubernetes, networking, and AI applications.
Requirements:
- Strong Python and FastAPI experience
- vLLM, Hugging Face and self-hosted LLM deployment
- Kubernetes, Docker, Helm and AWS
- NVIDIA GPU inference and performance optimization
- LangChain / LangGraph/LangSmith
- RAG, embeddings and vector databases (Qdrant)
- LLM observability and monitoring
- PostgreSQL / Redis
- GitHub Actions / CI/CD
- Strong production troubleshooting skills
Email CVs to: mahin@hirerightt.com.