The Senior Manager, Core Infrastructure Engineering role is responsible for managing a team that delivers scalable distributed systems and components. Key responsibilities include standardizing engineering practices, optimizing for high-throughput and hyper-scale workloads, and ensuring effective use of distributed state tools and data plane platforms. The role also involves guiding teams to design fault-tolerant, in-service-upgradable systems, setting SLO-aligned targets, and implementing resiliency mechanisms like load-shedding and throttling. The manager will oversee KPIs, telemetry, and dashboards, direct the design of functional and correctness requirements, and ensure proactive incident management, operational readiness, and on-call coverage. Security, compliance, and automation through Infrastructure as Code (IaC) are also critical aspects, along with managing change-management plans for safe patching and rollbacks. The position requires strong project management, cross-functional collaboration, problem-solving, and continuous improvement skills. The role is based in Nashville, TN.
Required Skills
Information Technology
○Distributed Systems○Scalability○Hyper-V○Data Platforms○KPI Reporting○OpenTelemetry○Dashboards○Data Replication○Incident Management○Encryption○Infrastructure as Code
Engineering, Construction & Trades
○Site Management
Other
○fault tolerance○in-service updates○SLOs○load shedding○throttling○rate limiting○correctness requirements○fault injection○synchronization○access control
Medical, dental, and vision insurance, including expert medical opinion, Short term disability and long term disability, Life insurance and AD&D, Supplemental life insurance (Employee/Spouse/Child), Health care and dependent care Flexible Spending Accounts, Pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, Paid time off (13 days annually for the first three years of employment and 18 days annually for subsequent years of employment, prorated for part-time), 11 paid holidays, 72 hours of paid sick leave upon date of hire (refreshes each calendar year, up to a maximum cap of 112 hours), Paid parental leave, Adoption assistance, Employee Stock Purchase Plan, Financial planning and group legal, Voluntary benefits including auto, homeowner and pet insurance.
Requirements
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms.
Description
This position is based at our Nashville, TN location.
Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting). Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.