Senior Manager, Core Infrastructure Engineering

🏢 Oracle
📍 Nashville, United StatesFull-timeOn-site
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
The Senior Manager, Core Infrastructure Engineering role is responsible for managing a team that delivers scalable distributed systems and components on a 2-4 quarter horizon. This role involves standardizing engineering practices, overseeing optimization for hyper-scale workloads, and ensuring effective use of distributed state and data plane platforms. Key responsibilities include guiding teams to design fault-tolerant, in-service-upgradable systems, setting SLO-aligned targets, and implementing resiliency mechanisms. The position also requires oversight of KPIs, telemetry, dashboards, and functional/correctness requirements, as well as fault-injection tests and replication strategies. Proactive incident management, operational readiness, on-call coverage, and driving encryption/access control practices are also critical. The role oversees the development and maintenance of automation/IaC and partners with teams on change-management plans for safe patching, updates, and rollbacks.
Required Skills
Information Technology
Distributed SystemsScalabilityData PlatformsIncident ManagementInfrastructure as CodePerformance TestingOpenTelemetryDashboardsCloud SecurityEncryption
Engineering, Construction & Trades
ATS Systems
Other
fault toleranceaccess control
Soft Skills & Professional Competencies
ResilienceLeadershipCollaborationProblem Solving
Productivity & Workplace Tools
RPA
Business, Sales & Management
Change ManagementProject ManagementProcess ImprovementPerformance ManagementTalent Acquisition
Finance, Legal & Governance
Tax Compliance
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation (13-18 days annually based on tenure); 11 paid holidays; Paid sick leave (72 hours upon hire, up to 112 hours carryover); Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
Manages team delivering scalable distributed systems and components on a 2-4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high-throughput, hyper-scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault-tolerant, in-service-upgradable systems, set SLO-aligned durability/availability targets, and implement resiliency mechanisms. Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault-injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on-call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change-management plans.
Description

This position sits onsite at our Nashville, TN location. 

Manages team delivering scalable distributed systems and components on a 2–4 quarter horizon. Standardizes engineering practices and scalability requirements across teams; oversees optimization for high‑throughput, hyper‑scale workloads; and ensures effective use of distributed state tools and data plane platforms. Guides teams to design fault‑tolerant, in‑service‑upgradable systems, set SLO‑aligned durability/availability targets, and implement resiliency mechanisms (load‑shedding, throttling, rate‑limiting). Provides oversight for KPIs, telemetry, and moderately complex dashboards; directs design of functional/correctness requirements, fault‑injection tests, and replication/synchronization strategies. Ensures proactive incident management, operational readiness, and on‑call coverage; drives encryption/access control practices, remediation plans, and compliance documentation. Oversees development and maintenance of automation/IaC and partners with teams on change‑management plans enabling safe patching, updates, and rollbacks.

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00