Site Reliability Engineer - Lead

🏢 Avrioc Technologies
📍 Abu Dhabi, United Arab EmiratesFull-timeOn-site
📅 Posted: 1w ago🔄 Updated: 1w ago
CV%
✨ AI Summary
Avrioc Technologies is seeking a Site Reliability Engineering (SRE) Lead in Abu Dhabi, UAE, to design, scale, and enhance their cloud infrastructure and observability ecosystem. The role involves architecting scalable cloud infrastructure, leading SRE best practices, optimizing CI/CD pipelines, defining SLOs/SLIs, building observability frameworks, managing Kubernetes clusters, and implementing auto-healing and proactive monitoring. The ideal candidate will have 8+ years of experience in DevOps/SRE with leadership experience, hands-on expertise in major cloud platforms (AWS, GCP, Azure), Infrastructure as Code, CI/CD, monitoring, incident response, observability tools, Kubernetes, and programming languages like Python, Bash, or Go. Experience in BCP/DR planning and capacity management is also required.
Required Skills
Other
AWS FISElastic StacklitmusChaos Mesh
Information Technology
GoArgo CD
Requirements
Requires 8+ years of experience in DevOps/SRE, including leadership in enterprise environments. Hands-on experience with AWS, GCP, or Azure, and strong expertise in Infrastructure as Code (Terraform, CloudFormation, Ansible). Proven experience in CI/CD, monitoring, incident response, observability tools, Kubernetes, Helm, and proficiency in Python, Bash, or Go. Experience in BCP/DR planning and capacity management is also needed, along with strong communication, troubleshooting, and documentation skills.
Description
HIRING: Site Reliability Engineer - Lead | Abu Dhabi, UAEWe're looking for a Site Reliability Engineering (SRE) Lead to design, scale, and elevate our cloud infrastructure and observability ecosystem.Key Responsibilities:• Architect and deploy scalable, highly available cloud infrastructure• Lead SRE best practices to ensure reliability, performance, and scalability• Optimize CI/CD pipelines (Jenkins, Argo CD or similar) for seamless deployments• Define and track SLOs & SLIs to maintain uptime and service health• Build robust observability frameworks (Elastic Stack, Prometheus, Grafana, Dynatrace, New Relic)• Manage Kubernetes clusters and Helm charts for efficient orchestration• Implement auto-healing systems and proactive monitoring• Drive chaos engineering and resilience testing (Chaos Mesh, Litmus, AWS FIS)• Collaborate with engineering and product teams to embed reliability into development• Maintain clear infrastructure and incident documentationWhat We're Looking For:• 8+ years of experience in DevOps/SRE, including leadership in enterprise environments• Hands-on experience with AWS, GCP, or Azure• Strong expertise in Infrastructure as Code (Terraform, CloudFormation, Ansible)• Proven experience in CI/CD, monitoring, and incident response• Deep knowledge of observability tools and practices• Strong Kubernetes and Helm experience at scale• Experience with databases like MySQL, Cassandra, etc.• Proficiency in Python, Bash, or Go• Experience in BCP/DR planning and capacity management• Strong communication, troubleshooting, and documentation skills
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00