✨ AI Summary
This Principal Software Engineer role focuses on leading the development and architecture of scalable, elastic distributed systems. Key responsibilities include optimizing code for high-throughput workloads, designing fault-tolerant and in-service-upgradable systems, and ensuring system reliability through redundancy, replication, and failover mechanisms. The role involves establishing performance indicators, building monitoring and alerting systems, and designing complex validation strategies for correctness and durability. Additionally, the engineer will proactively diagnose and resolve production issues, mentor peers, implement robust security controls, and develop Infrastructure as Code (IaC) and automation for safe system updates and rollbacks.
Ideal candidates will have experience in system design and architecture, fault-tolerant system design, performance testing, data replication, security controls implementation, Infrastructure as Code, and automation development. Strong skills in distributed systems, scalability, elasticity, fault tolerance, redundancy, replication, failover, load shedding, throttling, rate limiting, SLOs, KPIs, telemetry, dashboards, alerting, fault injection, synchronization, production issue resolution, mentoring, security controls, compliance documentation, patching, rollbacks, and change management are essential. The role also requires strong problem-solving and collaboration abilities.
🎁 Benefits & Perks
Medical, dental, and vision insurance, expert medical opinion, short term disability and long term disability, life insurance and AD&D, supplemental life insurance (Employee/Spouse/Child), health care and dependent care Flexible Spending Accounts, pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, Flexible Vacation (13-18 days annually), 11 paid holidays, 72 hours paid sick leave annually, paid parental leave, adoption assistance, Employee Stock Purchase Plan, financial planning and group legal, voluntary benefits including auto, homeowner and pet insurance.
Requirements
Leads development and architects components of scalable, elastic distributed systems. Defines and enforces scalability requirements, optimizes code and data paths for high-throughput workloads. Designs fault-tolerant, in-service-upgradable systems using redundancy, replication, and failover. Establishes KPIs and telemetry, builds dashboards and alerts, and designs complex validation for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation for safe patching, updates, and rollbacks.
Description
Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.