Senior Engineer - HPC Operations

🏢 core42
📍 Abu Dhabi, United Arab EmiratesFull-timeOn-site
📅 Posted: 6d ago🔄 Updated: 6d ago
CV%
✨ AI Summary
Core42 is seeking a Senior Engineer - HPC Operations to manage and support high-performance computing clusters for AI and ML workloads. This role requires leading daily operations, maximizing system efficiency, acting as a technical escalation point, and mentoring junior engineers. The position involves managing Slurm and Kubernetes environments, GPU resource management, and utilizing monitoring tools like Prometheus and Grafana. The ideal candidate will possess a Bachelor's or Master's degree in a technical field with at least 7 years of relevant experience. Key technical skills include expertise in HPC environments, Slurm, Kubernetes, GPU resource management, scripting (Python, Bash), Linux, networking (RDMA, InfiniBand), and storage technologies (NFS, Lustre, Ceph).
Required Skills
Other
RDMAinfinibandlustreDCGMSlurm
Operations, Logistics & Supply Chain
Process Optimization
Information Technology
MLflowMLOpsKubeFlow
🎁 Benefits & Perks
Competitive Salary, Yearly Bonus, Exclusive Discount Cards (Esaad and Fazaa), Premium Family Insurance (health, dental, vision, life), Learning & Development access to top-tier learning platforms.
Requirements
The ideal candidate will have a Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field, with 7+ years of experience in HPC operations, systems engineering, or DevOps roles. Required expertise includes advanced configuration and optimization of HPC environments, hands-on experience with Slurm and/or Kubernetes for AI/ML, GPU resource management, monitoring frameworks (Prometheus, Grafana, DCGM), strong scripting/automation skills (Python, Bash, Ansible, Terraform), and in-depth knowledge of Linux, networking (RDMA, InfiniBand, RoCE), and storage technologies (NFS, Lustre, Ceph).
Description
Senior Engineer - HPC OperationsAbout UsCore42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.The opportunityWe are seeking a highly skilled Senior Engineer – HPC Operations to oversee the daily operations and support of high-performance computing clusters designed to power large-scale AI and ML workloads. This role ensures stable, secure, and high-performing infrastructure leveraging technologies such as Slurm, Kubernetes, and modern MLOps platforms. The ideal candidate will bring deep technical expertise in HPC and a strong operational mindset to drive continuous improvement and automation across globally distributed environments. Responsibilities will extend to collaborating with multidisciplinary teams, leading complex projects, implementing cutting-edge technologies, and providing mentorship to operations engineers.Your key responsibilitiesLead the daily operational support of HPC infrastructure including compute, storage, networking, and scheduler components (Slurm, Kubernetes, etc.).Lead efforts to maximize the efficiency and performance of HPC systems, ensuring optimal resource utilization and minimal downtime.Act as the primary technical escalation point for L2 support teams and ensure prompt resolution of incidents and service requests.Monitor system health, performance, and utilization using advanced tools (e.g., Prometheus, Grafana, DCGM).Manage user environments for AI/ML workloads including container orchestration (e.g., Docker, Kubernetes) and workflow tools (e.g., MLflow, Kubeflow).Implement and manage job scheduling policies, priorities, and partitions within Slurm and/or Kubernetes environments to ensure fairness and efficiency.Lead root cause analysis (RCA) of operational issues and contribute to post-mortem documentation and continuous improvement efforts.Provide mentorship and guidance to junior engineers and participate in on-call rotation if required.Ensure compliance with security and operational policies; assist in audits and documentation for change and incident management processes.Qualifications:What we're looking for(a) Required skills / qualificationsBachelor's or Master's degree in Computer Science, Engineering, or related technical field.7+ years of experience in HPC operations, systems engineering, or DevOps roles.Advanced knowledge and expertise in configuring, optimizing, and maintaining complex HPC environments, including hardware, software, and storage systems.Hands-on experience managing Slurm clusters and/or Kubernetes-based environments for AI/ML workloads.Expert knowledge of GPU resource management, workload schedulers, and performance tuning for AI/ML workloads.Experience with monitoring and observability frameworks such as Prometheus, Grafana, and DCGM.Strong scripting and automation skills (Python, Bash, Ansible, Terraform).In-depth understanding of Linux (RHEL/CentOS/Ubuntu), networking concepts (RDMA, InfiniBand, RoCE), and storage technologies (NFS, Lustre, Ceph).What working at Core42 offers With a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person's contributions propel us forward, and together, we commit to achieving extraordinary results.Competitive Salary: We offer an attractive salary package based on your skills and experienceYearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonusExclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of servicesPremium Family Insurance: We provide comprehensive health coverage, including dental, vision and life insurance, ensuring the well-being of you and your familyLearning & Development: We offer access to top-tier learning platforms to help you grow in your career. Learn at your own pace with unlimited access to premium courses.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00