Senior Software Engineer, Core Infrastructure

🏢 Oracle
📍 Nashville, United StatesFull-timeOn-site
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
Oracle Cloud Infrastructure (OCI) is seeking Senior Software Engineers to join a new team in Nashville, TN, focusing on foundational distributed-systems, developer-platform, and reliability challenges. This role involves designing, implementing, and optimizing scalable, resilient, and operable components within large-scale cloud infrastructure. Key responsibilities include building fault-tolerant systems, implementing recovery-oriented principles, and ensuring system reliability through testing, monitoring, and incident response. The position also requires expertise in automation, Infrastructure as Code (IaC), and advanced security controls, alongside adherence to change management and compliance standards.
Required Skills
Information Technology
Distributed SystemsScalabilityCloud ArchitectureOrchestrationData PlatformsPerformance TestingData ReplicationEdge ComputingOpenTelemetryDashboardsIncident ResponseDebuggingInfrastructure as CodeEncryptionTechnical Documentation
Soft Skills & Professional Competencies
ResilienceTime ManagementRoot Cause AnalysisCollaborationProblem SolvingContinuous Learning
Other
operabilityload testingfault toleranceredundancyautomatic failoverretriescircuit breakerstimeoutsrunbooksdata synchronizationaccess controls
Engineering, Construction & Trades
AutomationFire Alarm
Finance, Legal & Governance
MediationTax Compliance
Business, Sales & Management
Change ManagementProcess Improvement
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation (13 days annually for the first three years of employment, 18 days annually thereafter for full-time employees); 11 paid holidays; Paid sick leave: 72 hours upon hire, carrying over up to 112 hours annually; Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance. May be eligible for bonus and equity.
Requirements
This role requires experience in designing, implementing, and optimizing components in distributed systems with an emphasis on scalability, resiliency, and operability. Responsibilities include delivering features, load/performance tests, leveraging data plane platforms and distributed state tools, building fault-tolerant paths, applying recovery-oriented principles, implementing retries, circuit breakers, and timeouts. Proactive issue detection and mitigation via tests, alarms, dashboards, and telemetry, authoring runbooks, and participating in incident response and RCAs are also key. Experience with standard replication and synchronization, automation/IaC for troubleshooting and maintenance, and applying advanced security controls is expected, while ensuring change, compliance, and documentation standards are met.
Description

This is an opportunity to solve foundational distributed-systems, developer-platform, and reliability problems with broad impact across OCI. We are hiring Senior Software engineers who are excited by large-scale cloud infrastructure, deployment orchestration, operational automation, and building trustworthy systems. You will help shape a new team and create the platforms that enable OCI engineers to deploy and operate services confidently at unprecedented scale.

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers' implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
 

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00