Senior Software Engineer, Core Infrastructure

🏢 Oracle
📍 Nashville, United StatesFull-timeOn-site
📅 Posted: 3w ago🔄 Updated: 3w ago
CV%
✨ AI Summary
Oracle Cloud Infrastructure is seeking a Senior Software Engineer for its Core Infrastructure team in Nashville, TN. This role involves solving foundational distributed-systems, developer-platform, and reliability problems with broad impact across OCI. The engineer will design, implement, and optimize components in distributed systems, focusing on scalability, resiliency, and operability. Key responsibilities include delivering features, load/performance tests, leveraging data plane platforms and distributed state tools, and building fault-tolerant paths. The role also involves proactive issue detection and mitigation, authoring runbooks, participating in incident response, implementing standard replication and synchronization, developing automation/IaC, and applying advanced security controls while ensuring compliance and documentation standards. Candidates should have experience in system design and architecture with an emphasis on scalability and reliability, including building fault-tolerant components, applying recovery-oriented computing principles, and implementing retry mechanisms. Operational troubleshooting, incident management, and security are also critical aspects of the role, along with maintaining automation scripts and adhering to change management plans. The position requires strong collaboration, problem-solving, and continuous learning skills.
Required Skills
Information Technology
Distributed SystemsScalabilityPerformance TestingData PipelinesEdge ComputingOpenTelemetryIncident ResponseData ReplicationInfrastructure as CodeEncryption
Soft Skills & Professional Competencies
ResilienceRoot Cause AnalysisProblem SolvingCollaborationContinuous Learning
Other
operabilityload testingfault toleranceretriescircuit breakerstimeoutsrunbookssynchronizationaccess controls
Productivity & Workplace Tools
RPA
Finance, Legal & Governance
MediationTax Compliance
Business, Sales & Management
Change ManagementProcess Improvement
🎁 Benefits & Perks
Medical, dental, and vision insurance, expert medical opinion, short term disability and long term disability, life insurance and AD&D, supplemental life insurance (Employee/Spouse/Child), health care and dependent care Flexible Spending Accounts, pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, paid time off (13-18 days annually depending on tenure), 11 paid holidays, 72 hours of paid sick leave annually, paid parental leave, adoption assistance, Employee Stock Purchase Plan, financial planning and group legal, voluntary benefits including auto, homeowner and pet insurance.
Requirements
The role requires experience in designing, implementing, and optimizing distributed systems with an emphasis on scalability, resiliency, and operability. Candidates should be proficient in load/performance testing, leveraging data plane platforms and distributed state tools, and building fault-tolerant paths. Experience with recovery-oriented principles, retries, circuit breakers, timeouts, telemetry, runbooks, incident response, RCA, replication, synchronization, automation/IaC, and advanced security controls is essential. A strong understanding of change management, compliance, and documentation standards is also required.
Description

This is an opportunity to solve foundational distributed-systems, developer-platform, and reliability problems with broad impact across OCI. We are hiring Senior Software engineers who are excited by large-scale cloud infrastructure, deployment orchestration, operational automation, and building trustworthy systems. You will help shape a new team and create the platforms that enable OCI engineers to deploy and operate services confidently at unprecedented scale.

Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers' implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
 

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00