✨ AI Summary
Oracle Cloud Infrastructure is seeking a Principal Software Engineer to join a new team in Nashville, TN, focused on building a platform for automated, safe, verifiable, and traceable change deployments across OCI services. This role involves leading development and architecting scalable, elastic distributed systems, optimizing for high-throughput and hyper-scale workloads, and designing fault-tolerant, in-service-upgradable systems. Key responsibilities include defining scalability requirements, leveraging data plane platforms, implementing redundancy and failover, handling network unreliability, establishing KPIs and telemetry, proactively diagnosing and resolving production issues, mentoring peers, and ensuring operational readiness. The role also requires implementing robust security controls, executing remediation, maintaining compliance documentation, and developing Infrastructure as Code (IaC) and automation for safe patching, updates, and rollbacks within change-management plans. Collaboration, problem-solving, and continuous learning are essential.
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term disability and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation (13 days annually for the first three years of employment, 18 days annually for subsequent years for full-time employees); 11 paid holidays; Paid sick leave: 72 hours upon hire, carries over up to 112 hours; Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
The ideal candidate will lead development and architect scalable, elastic distributed systems. Responsibilities include defining scalability requirements, optimizing code for high-throughput, designing fault-tolerant systems, establishing KPIs and telemetry, proactively diagnosing and resolving production issues, mentoring peers, implementing robust security controls, and developing IaC and automation for safe patching and rollbacks within change-management plans.
Description
Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.