✨ AI Summary
We are seeking an experienced Principal Site Reliability Engineer to provide senior technical leadership for a portfolio of business-critical enterprise applications. Based in Mexico, this role will strengthen global 24x7 operations by leading incident response, improving regional handoffs, and serving as a trusted escalation point for complex outages, maintenance activities, security vulnerability remediation, tech-stack patching, and application upgrades. The position combines approximately equal focus across application subject matter expert ownership, roadmap delivery, and incident management/on-call support. You will lead reliability and modernization work for a suite of applications while mentoring engineers and driving technical execution across teams and regions.
🎁 Benefits & Perks
flexible medical, life insurance, and retirement options, volunteer programs
Requirements
The role requires advanced Python skills and experience in creating automation to reduce operational toil. Experience with AI-assisted engineering tools, including Codex, is also required. The position involves designing and architecting infrastructure and services for reliability and functionality, forecasting demands, and collaborating with software development teams.
Description
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.