Senior MLOps Engineer

🏢 Global Corporation
📍 Abu Dhabi, United Arab EmiratesFull-timeOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
✨ AI Summary
MBZUAI is seeking a Senior MLOps Engineer to design, build, and maintain robust ML infrastructure for training, inference, and deployment pipelines, focusing on LLMs and speech models within Kubernetes environments. This role requires hands-on experience with Kubernetes (EKS), Helm, AWS, and MLOps toolchains. Responsibilities include managing the ML model lifecycle, automating MLOps pipelines, implementing Infrastructure as Code, and supporting GPU-accelerated environments. The ideal candidate will have a Bachelor's degree and at least 4 years of relevant experience, with strong collaboration skills and a passion for innovation.
Required Skills
Information Technology
KubernetesEKSEC2AWSS3IAMArgo CDTerraformPrometheusGrafanaCI/CDJenkinsAWS CodePipelinePyTorchCloudWatchMonitoringPythonDockerGit
Soft Skills & Professional Competencies
Collaboration
Nice to have:
Business, Sales & Management
Scheduling
Information Technology
AI IntegrationKubernetesCybersecurityModel DeploymentFine-TuningData Governance
Soft Skills & Professional Competencies
Research
Requirements
Requires a Bachelor's degree in Computer Science, AI Systems Engineering, or related field. Must have at least 4 years of experience in MLOps, DevOps, or cloud infrastructure engineering for ML systems, with strong proficiency in Kubernetes, Helm, and AWS services. Experience with ML model deployment tools like vLLM and CI/CD pipelines is essential.
Description
Job description / Role Job Type Full Time Job Location Abu Dhabi, UAE Nationality Any Nationality Salary Not Specified Gender Not Specified Arabic Fluency Not Specified Job Function IT - Software & Web Development Company Industry Software & Internet Services Application open: Full-time MBZUAI is looking to recruit a Senior MLOps Engineer for our Institute of Foundation Models department. The Institute of Foundation Models (IFM) at MBZUAI is dedicated to pioneering academic research at the forefront of global AI innovation, driven by real-world societal needs. IFM builds some of the world’s most powerful foundation models – open, fast, and focused on solving real-world problems. With deep scientific roots and world-class talent in Abu Dhabi, Paris, and Silicon Valley, IFM is shaping the future of AI. The MLOps Engineer will design, build, and maintain robust ML (Machine Learning) infrastructure across training, inference, and deployment pipelines. The role will take ownership of the model lifecycle from data ingestion to real-time serving and ensure large language models (LLM) and speech models are deployed efficiently, securely, and reproducibly in Kubernetes-based environments. This position requires hands-on experience with Kubernetes (EKS), Helm, AWS cloud infrastructure, and modern MLOps toolchains (e.g., vLLM, SGLang, OpenWebUI, Weights & Biases, MLflow). Familiarity with speech/voice AI frameworks like ElevenLabs, Whisper, and RVC is also valuable. Key experience required Infrastructure design and cloud management Design, build, and maintain scalable ML infrastructure on AWS (EKS, EC2, RDS, S3, IAM), Azure, or GCP to support AI and data-intensive workloads. Deploy and manage Kubernetes clusters using Helm, ArgoCD, and Terraform for reproducible and secure environments. Ensure observability, cost optimization, and reliability of multi-environment cloud resources with integrated monitoring (Prometheus, Grafana). MLOps and pipeline automation Develop and maintain automated MLOps pipelines for data versioning, model validation, and deployment using GitHub Actions, Jenkins, or AWS CodePipeline. Implement and optimize high-throughput model serving pipelines using vLLM, TensorRT, SGLang, or similar frameworks. Manage CI/CD workflows for model and application releases, integrating continuous testing and rollback strategies. Support real-time multimodal inference workloads (voice, text, vision) across distributed clusters. Infrastructure as code and system automation Implement Infrastructure as Code (IaC) using Terraform, Helm, and Ansible for automated configuration, provisioning, and governance. Create and manage ISO images, operating systems, and environment rebuilds for consistency across environments. Automate workstation, server, and network configurations (DHCP, DNS, TLS) across on-premises and cloud systems. GPU and ML environment support Set up and maintain GPU-accelerated environments with CUDA, cuDNN, PyTorch, NCCL, and relevant AI/ML libraries. Support containerized GPU workloads using Kubernetes GPU operators and optimize performance for LLM and TTS inference. Application deployment and monitoring Deploy and manage production-ready AI/ML applications with OpenWebUI, Gradio, or similar front-end interfaces for internal and external demos. Monitor and troubleshoot performance, resource utilization, and reliability; ensure proactive alerting and fault resolution. Security, compliance, and reliability Implement and enforce security best practices across infrastructure, data pipelines, and applications. Design and maintain disaster recovery, backup, and data protection strategies for critical systems. Ensure compliance with institutional and regulatory standards for data integrity and system resilience. Collaboration and integration Collaborate closely with ML researchers, AI engineers, and data scientists to productize and scale AI models (LLMs, ASR, TTS). Coordinate with cross-functional teams for project deployment, performance benchmarking, and workflow optimization. Innovation and continuous improvement Evaluate and integrate emerging DevOps, MLOps, and cloud-native technologies to enhance automation and scalability. Optimize cloud and hardware resource utilization to achieve operational efficiency and cost reduction. Documentation and knowledge transfer Maintain comprehensive documentation of infrastructure architectures, deployment processes, and operational workflows. Mentor junior engineers and promote best practices in DevOps, MLOps, and secure infrastructure management. Academic qualification Bachelor’s degree in Computer Science, AI Systems Engineering, or a related field. Professional experience required Essential: Minimum of 4 years of experience in MLOps, DevOps, or cloud infrastructure engineering for ML systems. Strong proficiency in Kubernetes, Helm, and container orchestration. Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve. Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM). Solid experience with Python, Docker, Git, and CI/CD pipelines. Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki). Excellent collaboration skills with ML researchers and software engineers. Preferred experience required Extensive experience with vLLM, Kubernetes, ElevenLabs, Whisper, Gradio/OpenWebUI, or custom TTS/ASR model hosting. Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration. Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and Cloudflare systems. Prior work in LLM deployment, fine-tuning pipelines, or foundation model research. Exposure to data governance and responsible AI operations in research or enterprise settings. Apply Now
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00