Requirements
Requires a Bachelor's degree in a relevant technical field, with a Master's degree preferred. Must have 5+ years of experience in HPC, Linux infrastructure, cloud infrastructure, distributed systems, or large-scale production environments, including experience with Slurm and Linux administration, and troubleshooting compute, storage, and networking systems. GPU cluster operations, specific NVIDIA technologies, InfiniBand networking, various storage platforms, cloud platforms, and Infrastructure-as-Code are preferred.
Description
Job description / Role
Job Type
Full Time
Job Location
Abu Dhabi, UAE
Nationality
Any Nationality
Salary
Not Specified
Gender
Not Specified
Arabic Fluency
Not Specified
Job Function
1
Company Industry
Software & Internet Services
Application open:
Full-time
MBZUAI’s Institute of Foundation Models is seeking a senior HPC engineer to provide technical leadership in designing, operating, and evolving large-scale GPU infrastructure supporting frontier AI research. The Institute for Foundation Models (IFM) operates one of the world’s largest AI-focused supercomputing environments and is looking for an experienced HPC engineer to contribute to groundbreaking research and development.
Key responsibilities
Lead operation and optimization of large-scale GPU clusters.
Drive reliability, scalability, and performance improvements.
Lead troubleshooting and root cause analysis of complex issues.
Design and validate new cluster deployments and upgrades.
Collaborate with researchers to optimize distributed AI training.
Lead vendor engagement and technical reviews.
Mentor junior engineers.
Define monitoring, operational standards, and capacity planning processes.
Participate in major incident management and escalations.
Academic qualification
Bachelor’s degree in computer science, computer engineering, electrical engineering, software engineering, information technology, applied mathematics, physics, or related disciplines. Master’s degree preferred.
Professional experience required
Essential:
5+ years in HPC, Linux infrastructure, cloud infrastructure, distributed systems, or large-scale production environments.
Experience with Slurm and Linux administration.
Experience troubleshooting compute, storage, and networking systems.
Preferred:
GPU cluster operations.
NVIDIA technologies including CUDA, NCCL, NVLink, and GPUDirect.
InfiniBand networking.
Weka, Lustre, BeeGFS, or similar storage platforms.
Azure, AWS, or GCP.
Terraform, Ansible, or Infrastructure-as-Code.
PyTorch Distributed, Megatron-LM, DeepSpeed, FSDP, or large-scale AI training environments.
Apply Now