Site Reliability Engineer

🏢 open innovation ai
📍 Abu Dhabi, United Arab EmiratesFull-timeOn-site
📅 Posted: Today🔄 Updated: Today
CV%
✨ AI Summary
Open Innovation AI is seeking a Site Reliability Engineer to support and maintain their AI products and deployments in customer environments, including secure, isolated, and air-gapped on-premises infrastructures. The role involves ensuring the availability, upgrades, and stability of end-to-end AI platforms, diagnosing and resolving incidents across hardware, Linux OS, Kubernetes, middleware, and application layers. Responsibilities include performing detailed log analysis, executing approved changes, maintaining a strong understanding of product architecture, and collaborating with L1/Service Desk teams. The engineer will also conduct on-site platform health assessments, work with Systems Engineering to resolve performance issues, and update technical documentation, adhering to ITIL processes.
Required Skills
Hospitality, Retail & Customer Service
Customer Service Management
Information Technology
Virtualization
Engineering, Construction & Trades
Troubleshooting
Other
compute storage networking
Requirements
The Site Reliability Engineer requires a Bachelor's degree in computer science, Information Technology, or Engineering, with 4-7 years of experience in SRE, DevOps, Infrastructure Operations, or Platform Engineering. Must have strong Linux system administration skills, hands-on experience with Kubernetes and container runtimes, and a solid understanding of compute, storage, networking, and virtualization. Practical experience with middleware and data-layer components, and a strong understanding of ITIL-aligned operational frameworks are essential. Excellent analytical and communication skills, with a methodical troubleshooting approach and the ability to produce clear technical documentation, are required.
Description
Company OverviewOpen Innovation AI is a global technology company that specializes in developing advanced solutions for managing AI workloads. Its flagship product, the Open Innovation Cluster Manager (OICM), orchestrates complex AI tasks efficiently across diverse infrastructures. The platform is hardware-agnostic, optimized for various GPUs and accelerators hardware, and facilitates seamless integration and scalability for enterprise AI applications. Open Innovation AI focuses on optimizing and simplifying AI workload management and making AI technologies accessible to organizations of all sizes. With its innovative solutions, companies can reduce operational costs, accelerate time to value, and maximize their return on investment, ensuring that their AI strategies contribute directly to enhanced business outcomes.Role Overview:The Site Reliability Engineer is responsible for supporting and maintaining Open Innovation AI Products and deployments across customer environments, including secure and isolated on-premises infrastructures. This role requires strong troubleshooting skills across hardware, Linux OS, Kubernetes, middleware, and application layers.The engineer is expected to diagnose and resolve technical incidents, applying deep product knowledge and strong analytical skills to restore service availability. The role requires solid understanding of operational processes such as Incident, Change, and Problem Management, along with a thorough grasp of the product architecture and how customers use it in production environments.Role Responsibilities:Experienced in Customer production environment deployments, mostly restricted connectivity and air gappedInvolves ensuring availability, upgrades, stability of end-to-end AI platforms and solutionsDiagnose and resolve incidents across hardware, Linux OS, Kubernetes clusters, containerized services, middleware, and platform components.Perform detailed analysis of logs, system behavior, and application output to identify root causes and restore service functionality.Review, validate, and execute approved changes following Change Management procedures, including system updates, configuration adjustments, and component upgrades.Maintain a strong understanding of the OICM and other OI product's architecture, its services, dependencies, and typical customer usage patterns.Collaborate with L1 and Service Desk teams by providing technical guidance, clarifying issue details, and ensuring accurate ticket triage.Escalate complex, code-level or product-defect issues to L3 with complete diagnostic in-formation and structured analysis.Conduct on-site platform health assessments, validating Kubernetes cluster status, ser-vice integrity, system resources, and overall environment readiness.Work closely with the Systems Engineering team to analyze and resolve performance is-sues across compute, storage, networking, and Kubernetes layers, and ensure that identified optimizations are reflected in the product and operational practices.Update and maintain technical documentation including SOPs, runbooks, troubleshooting steps, and known-issue guides.Participate in post-incident reviews, contributing technical insights and recommending improvements to prevent recurrence.Ensure all activities adhere to established Incident, Change, and Problem Management processes.Required experience & QualificationBachelor's degree in computer science, Information Technology, Engineering, or a related field.4–7 years of experience in SRE, DevOps, Infrastructure Operations, or Platform Engineering roles within on-prem or secure environments.Strong proficiency in Linux system administration, including troubleshooting, log analysis, service management, and performance tuning.Hands-on experience with Kubernetes, container runtimes, and distributed systems deployed in on-prem environments.Solid understanding of compute, storage, networking, and virtualization layers relevant to enterprise installations.Practical experience with middleware and data-layer components such as Kafka, Redis, PostgreSQL, or similar technologies used in distributed on-prem environments.Strong understanding of ITIL-aligned and experience operating within structured operational frameworks.Ability to diagnose complex issues across multiple layers of the stack.Experience working in secure, restricted, or isolated environments is an advantage.Excellent analytical skills, communication abilities, and a methodical approach to troubleshooting.Ability to produce clear technical documentation, including SOPs, runbooks, and investigation reports.Certifications such as RHCSA/RHCE, CKA/CKAD/CKS.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00