The Principal Software Engineer, Core Infrastructure will lead the development and architecting of scalable, elastic distributed systems. This role involves optimizing code and data paths for high-throughput, hyper-scale workloads, and leveraging data plane platforms for large-scale data retrieval, storage, and processing. The engineer will design fault-tolerant, in-service-upgradable systems, implement redundancy, replication, and failover mechanisms, and handle network unreliability through load-shedding and throttling while meeting SLOs. Key responsibilities include establishing KPIs and telemetry, building proactive dashboards and alerts, designing complex validation and replication for correctness and durability, and proactively diagnosing and resolving production issues. The role also requires mentoring peers, ensuring operational readiness, implementing robust security controls, executing remediation, maintaining compliance documentation, and developing Infrastructure as Code (IaC) and automation for safe patching, updates, and rollbacks within change-management plans.
○Root Cause Analysis○Problem Solving○Collaboration○Communication
🎁 Benefits & Perks
Medical, dental, and vision insurance, expert medical opinion, short term disability, long term disability, life insurance and AD&D, supplemental life insurance (Employee/Spouse/Child), Health care and dependent care Flexible Spending Accounts, Pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, Flexible Vacation, 11 paid holidays, 72 hours of paid sick leave upon date of hire (refreshes each calendar year, up to 112 hours carryover), Paid parental leave, Adoption assistance, Employee Stock Purchase Plan, Financial planning and group legal, Voluntary benefits including auto, homeowner and pet insurance. May be eligible for bonus, equity, and compensation deferral.
Requirements
Leads development and architecting of scalable, elastic distributed systems. Optimizes code and data paths for high-throughput, hyper-scale workloads. Designs fault-tolerant, in-service-upgradable systems using redundancy, replication, failover, and policies for partitions. Establishes KPIs and telemetry, builds dashboards and alerts, and designs complex validation and replication for correctness and durability. Diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation for safe patching, updates, and rollbacks.
Description
Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.