✨ AI Summary
The Senior Core Infrastructure Engineer will design, implement, and optimize components in distributed systems, focusing on scalability, resiliency, and operability. This role involves developing features, conducting load/performance tests, and leveraging data plane platforms for high-volume data handling. Key responsibilities include building fault-tolerant paths, applying recovery-oriented principles, and proactively detecting/mitigating issues through tests, alarms, dashboards, and telemetry. The engineer will also author runbooks, participate in incident response, implement standard replication and synchronization, develop automation/IaC, and apply advanced security controls while ensuring compliance and documentation standards are met.
Qualifications include 6+ years of distributed service engineering experience, hands-on experience with highly-available web services, service-oriented architectures, RESTful web services, and strong Java development skills. Proficiency in Java, JavaScript, Terraform, OAuth, OpenID Connect, SAML, and CI/CD is required, with a preference for cloud infrastructure experience. Experience with scripting languages for automation is also necessary. Domain knowledge of Identity and Access Management, experience with public cloud platforms, and understanding of multi-AD/AZ and regional data centers are preferred.
🎁 Benefits & Perks
competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
Requirements
Requires 6+ years of distributed service engineering experience in a software development environment, with hands-on experience building and operating highly-available, high-traffic web services. Must have experience developing service-oriented architectures and RESTful web services, strong development experience in Java, and proficiency in tools like Java, JavaScript, Terraform, OAuth, OpenID Connect, SAML, and CI/CD. Experience with at least one scripting language for automation is also required.
Description
Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.