Senior Software Engineer, Core Infrastructure

🏢 Oracle
📍 Nashville, United StatesFull-timeOn-site
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
Senior Software Engineers are sought to solve foundational distributed-systems, developer-platform, and reliability problems within Oracle Cloud Infrastructure (OCI). This role involves designing, implementing, and optimizing components in distributed systems with a strong emphasis on scalability, resiliency, and operability. Key responsibilities include building fault-tolerant paths, applying recovery-oriented principles, implementing retry mechanisms and circuit breakers, and proactively detecting and mitigating issues through tests, alarms, dashboards, and telemetry. The role also requires authoring runbooks, participating in incident response and Root Cause Analyses (RCAs), implementing standard replication and synchronization, developing automation/Infrastructure as Code (IaC) for troubleshooting and maintenance, and applying advanced security controls while adhering to change, compliance, and documentation standards.
Required Skills
Information Technology
Distributed SystemsScalabilityCloud ArchitectureOrchestrationEdge ComputingOpenTelemetryIncident ResponseData ReplicationInfrastructure as CodeEncryptionPerformance TestingDashboardsSystem Design
Soft Skills & Professional Competencies
ResilienceRoot Cause AnalysisCollaborationProblem SolvingContinuous Learning
Other
operabilityretry mechanismscircuit breakerstimeoutsrunbooksdata synchronizationaccess controlsload testing
Engineering, Construction & Trades
AutomationATS Systems
Finance, Legal & Governance
Mediation
Business, Sales & Management
Change Management
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation (13 days annually for first three years of employment, 18 days annually for subsequent years); 11 paid holidays; Paid sick leave (72 hours upon hire, refreshes annually up to 112 hours); Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing. Builds fault-tolerant paths, applies recovery-oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls.
Description

This is an opportunity to solve foundational distributed-systems, developer-platform, and reliability problems with broad impact across OCI. We are hiring Senior Software engineers who are excited by large-scale cloud infrastructure, deployment orchestration, operational automation, and building trustworthy systems. You will help shape a new team and create the platforms that enable OCI engineers to deploy and operate services confidently at unprecedented scale.

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers' implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00