Principal Core Infrastructure Engineer

🏢 Oracle
📍 Nashville, United StatesFull-timeOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
✨ AI Summary
We are seeking a Principal Core Infrastructure Engineer to lead the development and architecting of scalable, elastic distributed systems. This role involves optimizing code and data paths for high-throughput, hyper-scale workloads, and designing fault-tolerant systems with redundancy, replication, and failover. The engineer will establish KPIs and telemetry, build proactive dashboards and alerts, and design complex validation and synchronization mechanisms. Key responsibilities include diagnosing and resolving production issues, mentoring peers, ensuring operational readiness, implementing robust security controls, and developing Infrastructure as Code (IaC) and automation for safe patching, updates, and rollbacks. The ideal candidate will have strong technical depth in distributed systems, Networking, and AI infrastructure protocols like RDMA. Experience in building infrastructure platforms and a passion for tackling open-ended problems are essential. This is a senior individual contributor role (IC4) focused on driving major technical components within the platform, translating high-level goals into scalable designs, and delivering reliable, cloud-native solutions. The role also involves leading design reviews, being a hands-on contributor, mentoring peers, and advocating for best practices.
Required Skills
Engineering, Construction & Trades
ATS Systems
Other
redundancyfailoverload-sheddingthrottlingrate-limitingSLOsfault injectionbrownoutssynchronizationsecurity controlsupdatesrollbacksload testingnetwork partitionsconsistencypartition tolerancedurabilityaccess controls
Information Technology
Data ReplicationKPI ReportingOpenTelemetryDashboardsAlertingInfrastructure as CodeCybersecurityData PlatformsPerformance TestingHigh AvailabilityDebuggingIncident ResponseEncryptionCloud Architecture
Finance, Legal & Governance
MediationTax Compliance
Productivity & Workplace Tools
RPA
Business, Sales & Management
Change ManagementMentoringProcess Improvement
Soft Skills & Professional Competencies
Time ManagementRoot Cause AnalysisProblem SolvingCollaborationContinuous Learning
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Flexible Vacation: 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment (pro-rated for part-time); 11 paid holidays; 72 hours of paid sick leave upon date of hire, carrying over up to a maximum cap of 112 hours; Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
Leads development and architecting of scalable, elastic distributed systems. Defines and enforces scalability requirements, optimizes code and data paths, and leverages data plane platforms. Designs fault-tolerant, in-service-upgradable systems using redundancy, replication, failover, and policies for partitions. Establishes KPIs and telemetry, builds proactive dashboards and alerts. Designs complex validation, replication, and synchronization. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation for safe patching, updates, and rollbacks.
Description

Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00