Senior Site Reliability Engineer (Performance and Scalability)

🏢 Digitalzone
📍 United Arab EmiratesFull-timeOn-site
📅 Posted: 4d ago🔄 Updated: 4d ago
CV%
✨ AI Summary
DigitalZone is seeking a Senior Site Reliability Engineer (Performance and Scalability) to enhance the platform's ability to handle traffic spikes and empower engineering teams with load and failure testing capabilities. This role focuses on building scalability foundations, establishing testing as a standard practice, and owning SLOs, error budgets, and the observability stack for TypeScript, Go, and PHP/Laravel services. The engineer will also harden Postgres and AWS infrastructure, lead incident response, and drive systemic fixes. A minimum of 5 years of experience in SRE, platform, or backend engineering with production ownership of large-scale systems is required, along with deep AWS experience, familiarity with Postgres performance, observability tooling, infrastructure-as-code, and scripting in Go, TypeScript, or similar.
Required Skills
Information Technology
Infrastructure as CodeGoObservability
🎁 Benefits & Perks
Immediate, large-scale impact on a high-growth business, Top-of-the-market compensation packages, Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
Requirements
5+ years in SRE, platform, or backend engineering with production ownership of large-scale systems. Proven experience scaling systems through traffic spikes and running load/failure testing programs. Deep AWS experience, solid grasp of Postgres performance and scaling, fluency with observability tooling and infrastructure-as-code, and scripting in Go, TypeScript, or similar. Requires a calm, systematic approach to incidents and strong communication skills for enabling other teams.
Description
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.What you'll doBuild the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state loadEstablish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the resultsOwn SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scaleHarden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaCLead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted onPartner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systemsRequirementsWhat you'll bring5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minuteA track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adoptedDeep AWS experience and a solid grasp of Postgres performance and scalingFluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similarA calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeepBenefitsImmediate, large-scale impact on a high-growth businessTop-of-the-market compensation packagesWork alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00