Senior Site Reliability Engineer

🏢 Mozn
📍 Cairo, EgyptFull-timeRemote
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
Mozn is seeking a Senior Site Reliability Engineer (SRE) to build an AI-powered agentic layer that automates SRE workloads. The role involves carrying a normal on-call rotation, responding to incidents, performing root cause analysis, and working hands-on with Kubernetes and cloud environments. A key responsibility is to design, build, and ship LLM-based agents using tools like Claude Code, OpenAI Codex, or Kimi K2/K3, integrating them with existing infrastructure and defining safe operational guardrails. Candidates must have at least 3 years of experience building production software with LLMs and hands-on experience with agentic coding tools. Strong Python skills, real SRE experience (including incident response and application-level debugging), and fluency with Kubernetes, cloud providers, and observability tools are required. The role also emphasizes building trust with stakeholders and maintaining a security- and compliance-first posture, aligning with Saudi data residency requirements. Experience in Saudi Arabia/MENA and a background in ML engineering or LLMOps are considered beneficial.
Required Skills
Other
LLM agentic workflowstool/function callingKimi K2/K3API wrapperson-call rotationDatadogpermissioningapproval gatesrollback pathsauditabilitybuild trust
Soft Skills & Professional Competencies
PlanningRoot Cause Analysis
Information Technology
RAGClaudeOpenAIPythonOrchestrationIncident ResponseDebuggingKubernetesAWSGCPOCIAzurePrometheusGrafanaElasticsearch
Engineering, Construction & Trades
ToolingATS Systems
Nice to have:
Information Technology
TerraformDockerLLMOps
Other
Ansibleeval/backtest harnesses for LLM agentsplatform engineering
Engineering, Construction & Trades
Civil Engineering
Description
About MoznMOZN is a leading Enterprise AI company enabling organizations to make informed decisions in two critical domains: Financial Crime Prevention and Enterprise Knowledge Intelligence.We’re a diverse, collaborative team of innovators united by a shared purpose: to build AI that delivers tangible business value, builds trust, and empowers people and organizations with augmented intelligence. Our culture is built on the relentless pursuit of excellence and meaningful impact.If you’re passionate about working alongside exceptional talent on world-class AI, and you want the autonomy and runway to do the best work of your career, join us in shaping the future of intelligent enterprises.About the roleWe're hiring an AI Senior SRE: someone who carries a normal SRE workload — on-call rotation, incident response, root cause analysis, hands-on Kubernetes/cloud work — and builds the agentic layer that automates that workload over time. You are not exempt from operating production. You're the person best placed to know what should be automated, because you're the one doing it. This role exists because most reliability toil (triage, root-causing, remediation, SLO tracking, onboarding checks) is repetitive and well-defined enough to hand to an LLM-based agent with the right guardrails. Your job is to do the on-call/ops work like any SRE, then turn what you learn into agents — identify the workflow, build the agent, and earn the trust to let it act with increasing autonomy.What you'll doCarry a normal on-call rotation and act as a hands-on responder: investigate, fix, and document incidents yourself, exactly like any SRE on the team — especially for workflows that don' t have an agent yet.Go deep on application-level reliability, not just infrastructure: read and debug service code, understand business logic well enough to find the real root cause, and ship fixes or PRs directly into application repos when the fix belongs in the app, not the platform.Design, build, and ship LLM-based agents (using tools like Claude Code, OpenAI Codex, or Kimi K2/K3) that plug into our existing stack: Kubernetes, cloud APIs, Prometheus/Grafana/ELK/Datadog, PagerDuty/Slack.Define the "tool" interface each agent needs — the specific APIs, scripts, and read/write actions — and build safe wrappers around them.Set clear guardrails for every agent: what it may do autonomously vs. what it must propose for human approval, with a bias toward human-in-the-loop until an agent has earned trust.Own agent evaluation — define what "correct" and "safe" look like per agent, and build test/backtest suites against real historical incidents.Continuously tune prompts, context, and tool schemas as an agent' s scope grows.Partner with the SRE/platform team to find good agent candidates: repetitive, well-scoped, auditable workflows.Report on agent impact — MTTD/MTTR/MTTX movement, false positive/negative rates, and engineer-hours of toil removed.Keep a security- and compliance-first posture: audit trails for every autonomous action, least-privilege access to production, and alignment with Saudi data residency/regulatory requirements.Requirements3+ years building production software with LLMs — agentic workflows, tool/function calling, multi-step planning, RAG — not just personal use of a chat assistant.Hands-on experience shipping real work with an agentic coding tool such as Claude Code, OpenAI Codex, or Kimi K2/K3.Strong Python (or similar) for building agent tooling, API wrappers, and orchestration.Real, hands-on SRE experience: comfortable being a primary on-call responder, running incident response, and doing root cause analysis under pressure — not just familiar with the concepts.Application-level debugging skill, not just infra: able to read a service' s codebase, trace a failure back to the actual line/logic causing it, and ship a fix yourself — SRE work here isn' t limited to restarting pods or scaling nodes.Solid hands-on Kubernetes and cloud provider experience (AWS/GCP/OCI/Azure) and fluency with observability tools (Prometheus, Grafana, Datadog, ELK) — both as an operator and as integration points for agents.Understands guardrails for autonomous systems: permissioning, approval gates, rollback paths, auditability.Can build trust with technical stakeholders — every agent starts with limited autonomy and has to earn more.Nice to haveExperience in Saudi Arabia / MENA, ideally consulting across public or private sector clients.Familiarity with Terraform/Ansible, Docker, VM/on-prem setups — useful context for the agents you' ll build, not the core job.Experience building eval/backtest harnesses for LLM agents against historical incident data.Background in ML engineering, LLMOps, or platform engineering.BenefitsYou will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting spaceYou will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is bestThe fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do bestYou will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AIWe believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00