✨ AI Summary
The Senior Manager, Core Infrastructure Engineering role involves leading teams in the development and delivery of scalable distributed systems and components with a 2-4 quarter horizon. Responsibilities include standardizing engineering practices, optimizing for high-throughput, hyper-scale workloads, and ensuring effective use of distributed state tools and data plane platforms. The role requires guiding teams to design fault-tolerant, in-service-upgradable systems, set SLO-aligned targets, and implement resiliency mechanisms. Oversight of KPIs, telemetry, moderately complex dashboards, functional/correctness requirements, fault-injection tests, and replication/synchronization strategies are also key. Proactive incident management, operational readiness, on-call coverage, encryption/access control practices, remediation plans, and compliance documentation are essential. The manager will also oversee the development and maintenance of automation/IaC and partner with teams on change-management plans for safe patching, updates, and rollbacks. This role also involves planning and execution of projects, driving cross-functional partnerships, problem-solving complex issues, and fostering continuous learning and improvement within the teams.
🎁 Benefits & Perks
flexible medical, life insurance, and retirement options, volunteer programs
Requirements
The role requires managing teams that deliver scalable distributed systems and components, with a focus on standardization of engineering practices, optimization for high-throughput workloads, and ensuring effective use of distributed state tools and data plane platforms. Key responsibilities include guiding teams in designing fault-tolerant, in-service-upgradable systems, setting SLO-aligned targets, implementing resiliency mechanisms, overseeing KPIs and telemetry, and ensuring proactive incident management, operational readiness, and security practices.
Description
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting). Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.