✨ AI Summary
As a Principal Site Reliability Engineer at Oracle Cloud Infrastructure (OCI), you will lead the design, automation, and support of highly available systems, focusing on resiliency, security, scalability, and performance. This hands-on role requires deep infrastructure and software engineering expertise to own and improve end-to-end reliability metrics, architect high-availability solutions, and serve as the ultimate escalation point for complex operational issues. You will architect and build automation tools, collaborate with development teams on service design and deployments, and mentor junior engineers.
🎁 Benefits & Perks
competitive benefits that support our people with flexible medical, life insurance, and retirement options. Volunteer programs.
Requirements
Requires advanced experience with Linux systems administration, strong programming skills in Python, advanced Bash/Shell scripting, deep understanding of distributed systems, networking, service architecture, and databases. Must have solid knowledge of CI/CD pipelines, Agile methodologies, and DevOps best practices, with proven ability to lead cross-functional efforts and technical problem-solving.
Description
As a Principal member of the Site Reliability Engineering (SRE) team, you'll take ownership of highly available systems, influence service design, and work across teams to drive resiliency, automation, and operational excellence. This is a hands-on engineering role where deep infrastructure knowledge meets software engineering expertise, ideal for experienced SREs ready to take the lead.
This is not a fully remote role but a hybrid role. Does require in office at least 3 days a week in Guadalajara.