Senior MLOps Engineer

🏢 GCS
📍 United Arab EmiratesFull-timeRemote
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
The Senior MLOps Engineer will be responsible for owning the infrastructure that trains, serves, and scales large language models in production for a UAE-based Sovereign AI company. Key responsibilities include deploying and scaling self-hosted open-weight models, operating a multi-provider gateway, building and maintaining ML pipelines, and ensuring production reliability through monitoring, logging, and incident response. The role requires hands-on experience with MLOps, ML infrastructure, serving large models at scale, Kubernetes, Docker, and IaC (Terraform), along with experience handling real production incidents.
Required Skills
Other
DeepSpeedtritonAcceleratevLLMTGIFSDP
Information Technology
MLflowSageMakerKubeFlow
Requirements
Requires 5+ years in MLOps, ML infrastructure, or ML engineering, with hands-on experience serving large models at scale using vLLM/Triton/TGI, quantization, and GPU optimization. Strong Kubernetes, Docker, and IaC (Terraform) skills are essential. Candidates must have experience being on-call for ML services and handling real incidents.
Description
Title: Senior MLOps Engineer (Remote) | Sovereign AILocation: RemoteAbout the role:We are hiring a Senior MLOps Engineer for one of the leading players in Sovereign AI, a UAE-based company with frontier-model access and agentic AI products live across critical sectors at unprecedented scale.You'll own the infrastructure that trains, serves, and scales large language models in production, the systems that keep high-stakes AI running reliably and cost-effectively.What you'll do:Deploy and scale self-hosted open-weight models (from 7B up to 370B+ parameters) using engines like vLLM, Triton, or TGI, choosing serving strategies (continuous batching, tensor/pipeline parallelism, quantization) to hit latency and cost targetsOperate a multi-provider gateway spanning self-hosted models and external APIs, with token accounting, budget controls, and cost-aware routingBuild and maintain pipelines for fine-tuning, evaluation, versioning, and continuous delivery (MLflow, SageMaker, or Kubeflow), including distributed training (DeepSpeed, FSDP, Accelerate)Own production reliability: monitoring, logging, alerting, incident response, and safe rollbackWhat we're looking for:5+ years in MLOps, ML infrastructure, or ML engineering, owning end-to-end model lifecycles in productionHands-on experience serving large models at scale (vLLM/Triton/TGI, quantization, GPU optimization)Strong Kubernetes, Docker, and IaC (Terraform)You've been on-call for ML services and handled real incidents, not just built happy-path pipelinesThe setup:Fully remote, contract-basedOnboarding ASAPWorld-class environment: elite talent from FAANG and frontier AI labsInterested Apply here or message me directly with your CV.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00