Halian | Managed Services, Recruitment Agency & Contract Staffing
Senior DevOps Engineer (m/f/d)
About the role
Key Responsibilities
-
Lead the deployment, administration, and optimization of enterprise Kubernetes environments running on Nutanix infrastructure, ensuring resilience, scalability, and operational excellence.
-
Design and build infrastructure solutions that support advanced AI and machine learning workloads, including GPU-enabled environments, model deployment frameworks, and large-scale data processing pipelines.
-
Develop and maintain automated infrastructure provisioning and configuration management using tools such as Terraform, Ansible, Helm, and GitOps frameworks.
-
Establish and enhance CI/CD pipelines to streamline the deployment of AI services, containerized applications, and machine learning workflows.
-
Drive operational maturity through the implementation of monitoring, logging, tracing, and alerting solutions utilizing platforms such as Prometheus, Grafana, ELK, and OpenTelemetry.
-
Implement and maintain security controls including RBAC, container image governance, secrets management, network segmentation, and workload protection aligned with enterprise compliance standards.
-
Continuously optimize cluster performance, storage utilization, networking efficiency, and compute resource allocation across CPU and GPU workloads.
-
Work closely with AI engineers, platform architects, and development teams to onboard applications, improve user experience, and establish operational best practices.
-
Define, implement, and validate backup, recovery, and disaster recovery procedures for Kubernetes platforms and associated data services.
Required Qualifications
-
Minimum 5 years of experience in DevOps, Platform Engineering, Cloud Operations, or Site Reliability Engineering, including at least 3 years of hands-on Kubernetes administration in production.
-
Strong experience working with Nutanix technologies including AHV, Prism, Karbon, Files, and Objects within enterprise environments.
-
Demonstrated experience supporting AI/ML platforms and tools such as Kubeflow, MLflow, KServe, Ray, or similar ecosystems.
-
Advanced knowledge of Infrastructure as Code and platform automation using Terraform, Ansible, Helm, and GitOps methodologies.
-
Proficiency in scripting and automation using Python, Bash, Go, or equivalent languages.
-
Experience managing GPU-enabled Kubernetes environments and NVIDIA accelerator technologies.
-
Sound understanding of Kubernetes networking, storage, security, and policy enforcement frameworks.
-
Hands-on experience with CI/CD tools such as GitLab CI, GitHub Actions, Jenkins, or similar platforms.
-
Experience implementing and managing observability and monitoring solutions in cloud-native environments.
-
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience.
Preferred Qualifications
-
Nutanix certifications such as NCP-MCI, NCP-DS, or related credentials.
-
CNCF certifications including CKA, CKAD, or CKS.
-
Experience managing multiple Kubernetes clusters across hybrid or enterprise environments using solutions such as Rancher, OpenShift, or Anthos.
-
Strong understanding of MLOps principles, model lifecycle management, and machine learning deployment frameworks.
-
Exposure to regulated industries and compliance-driven environments, including financial services, healthcare, or government sectors.
-
Familiarity with modern governance and security standards such as SOC 2, HIPAA, GDPR, or ISO 27001.
Senior DevOps Engineer in Abu Dhabi, United Arab Emirates … more