Expert Site Reliability Engineer

🏢 TAWANTECH
📍 Riyadh, Saudi ArabiaFull-timeOn-site
📅 Posted: 2d ago🔄 Updated: 2d ago
CV%
✨ AI Summary
The Expert Site Reliability Engineer will be responsible for driving the reliability, availability, scalability, and operational resilience of critical technology services. Key duties include defining and implementing reliability engineering practices, establishing SLIs/SLOs, designing automation to reduce manual operations, and enhancing monitoring and alerting capabilities. The role also involves leading technical analysis for incidents, conducting root-cause analysis, and designing solutions for improved system availability and scalability. The ideal candidate will have a Bachelor's degree in a related field, 5+ years of experience in SRE/DevOps/Platform Engineering, strong experience with cloud platforms, Kubernetes, monitoring tools, automation, and incident management. Experience in Banking, FinTech, or Payment environments is preferred.
Required Skills
Information Technology
Reliability EngineeringDevOpsCloud ComputingKubernetesMonitoringObservabilityAlertingScriptingCI/CDInfrastructure as CodeTerraformIncident ManagementDebuggingHigh AvailabilityScalability
Other
Platform Engineeringproduction environmentsSLIsSLOsreliability metricsperformance engineering
Productivity & Workplace Tools
RPA
Soft Skills & Professional Competencies
Root Cause AnalysisLeadershipAnalytical SkillsProblem Solving
Business, Sales & Management
Workforce Planning
Finance, Legal & Governance
Business Continuity
Requirements
QUALIFICATIONS & REQUIREMENTS Bachelor’s degree in Computer Science, Software Engineering, IT, or a related field. 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related roles. Strong experience in cloud platforms, Kubernetes, and production environments. Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics. Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform). Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA). Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery. Experience driving reliability improvements and reducing operational toil through automation. Strong analytical, problem-solving, and technical leadership skills. Experience in Banking, FinTech, or Payment environments is preferred.
Description
Purpose:To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices. Main Duties and Responsibilities: Define and implement advanced reliability engineering practices across critical technology services. Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets. Design automation to reduce manual operational activities and improve system resilience. Develop and enhance monitoring, observability, alerting, and incident detection capabilities. Lead technical analysis and resolution of complex production incidents. Conduct root-cause analysis and drive permanent corrective and preventive actions. Design solutions to improve system availability, scalability, capacity, and disaster resilience. Identify reliability risks and recommend architectural and engineering improvements. Drive performance engineering and capacity planning for critical services. Provide advanced technical guidance and mentorship on SRE practices. Promote automation and engineering approaches that reduce operational toil and improve service reliability. 
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00