Senior Site Reliability Engineer (Performance and Scalability)

🏢 Global Corporation
📍 United Arab EmiratesFull-timeOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
✨ AI Summary
The Senior Site Reliability Engineer (Performance and Scalability) will be responsible for building DigitalZone's platform scalability to handle traffic spikes and enabling engineering teams to load and failure test their systems. Key responsibilities include developing the scalability foundation (capacity planning, autoscaling, caching, queueing, graceful degradation), establishing load and failure testing as a standard practice, owning SLOs and the observability stack, hardening Postgres and AWS infrastructure, reducing toil through automation, and leading incident response. The ideal candidate will have 5+ years of experience in SRE, platform, or backend engineering with a proven track record of scaling systems and running successful testing programs. Deep AWS and Postgres experience, fluency with observability tooling and infrastructure-as-code, and scripting skills in Go, TypeScript, or similar are required. A calm, systematic approach to incidents and strong communication skills are also essential.
Required Skills
Business, Sales & Management
Workforce PlanningCoachingSales Enablement
Other
autoscalingqueueinggraceful degradationload testingfailure testingSLOserror budgetslogstraces
Information Technology
ObservabilityAlertingPostgreSQLAWSInfrastructure as CodeIncident ResponsePostmortemsGoTypeScriptLaravel
Engineering, Construction & Trades
Piping Isometrics
Productivity & Workplace Tools
RPA
Soft Skills & Professional Competencies
CommunicationInfluencing
🎁 Benefits & Perks
Immediate, large-scale impact on a high-growth business. Top-of-the-market compensation packages. Work alongside top regional talent.
Requirements
Requires 5+ years in SRE, platform, or backend engineering with production ownership of large-scale systems. Must have a track record of scaling systems through traffic spikes and running load/failure testing programs. Experience with AWS, Postgres performance/scaling, observability tooling, infrastructure-as-code, and scripting in Go, TypeScript, or similar is essential. Strong communication and systematic incident response skills are required.
Description
Job description / Role Job Type Full Time Job Location UAE Nationality Any Nationality Salary Not Specified Gender Not Specified Arabic Fluency Not Specified Job Function IT - Software & Web Development Company Industry Software & Internet Services Your mission Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure-test their own systems. This is an enablement role at its core: you raise the reliability bar across the organization by building capability, not by owning every service yourself. What you'll do Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load. Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results. Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale. Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and infrastructure-as-code. Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on. Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems. Requirements What you'll bring 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at tens of thousands of requests per minute. A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted. Deep AWS experience and a solid grasp of Postgres performance and scaling. Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar. A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep. Benefits Immediate, large-scale impact on a high-growth business. Top-of-the-market compensation packages. Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more. Apply Now
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00