Principal Site Reliability Engineer

🏢 Oracle
📍 ZAPOPAN, MexicoFull-timeOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
✨ AI Summary
We are seeking an experienced Principal Site Reliability Engineer to provide senior technical leadership for a portfolio of business-critical enterprise applications. Based in Mexico, this role will strengthen global 24x7 operations by leading incident response, improving regional handoffs, and serving as a trusted escalation point for complex outages, maintenance activities, security vulnerability remediation, tech-stack patching, and application upgrades. The position combines approximately equal focus across application subject matter expert ownership, roadmap delivery, and incident management/on-call support. You will lead reliability and modernization work for a suite of applications while mentoring engineers and driving technical execution across teams and regions.
Required Skills
Information Technology
PythonIncident ResponseIT InfrastructureTechnical DocumentationSSIS
Productivity & Workplace Tools
RPA
Engineering, Construction & Trades
CAPA
Other
service reliabilityexperimentation
Business, Sales & Management
Performance ManagementMentoring
Soft Skills & Professional Competencies
Root Cause AnalysisCollaborationProblem SolvingContinuous Learning
🎁 Benefits & Perks
flexible medical, life insurance, and retirement options, volunteer programs
Requirements
The role requires advanced Python skills and experience in creating automation to reduce operational toil. Experience with AI-assisted engineering tools, including Codex, is also required. The position involves designing and architecting infrastructure and services for reliability and functionality, forecasting demands, and collaborating with software development teams.
Description
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00