Senior Manager, Site Reliability Engineering

🏢 Oracle
📍 Reston, United StatesOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
✨ AI Summary
The Senior Manager, Site Reliability Engineering will lead teams in designing and architecting infrastructure and services, providing guidance on reliability and functionality practices. This role involves supervising team members to ensure accurate demand forecasting, capacity management, and resource allocation. The manager will foster collaboration with software development teams to build scalable and reliable infrastructures and oversee incident and service lifecycle management, including monitoring, incident response, and root cause analysis. Additionally, the position will drive automation initiatives, implement standards for identifying automation opportunities, and provide technical communication and guidance to teams and stakeholders. The role also serves as a senior escalation point for complex issues and promotes continuous innovation and improvement in site reliability trends and technologies.
Nice to have:
Information Technology
Reliability EngineeringIncident ResponseScriptingConcurrent ProgrammingScalability
Engineering, Construction & Trades
Structural Design
Business, Sales & Management
Workforce PlanningPerformance ManagementProject Management
Soft Skills & Professional Competencies
Root Cause AnalysisCommunicationProblem SolvingCollaborationLeadership
Productivity & Workplace Tools
RPA
Finance, Legal & Governance
Budgeting
🎁 Benefits & Perks
Medical, dental, and vision insurance, expert medical opinion, short term and long term disability, life insurance and AD&D, supplemental life insurance, health care and dependent care Flexible Spending Accounts, pre-tax commuter and parking benefits, 401(k) Savings and Investment Plan with company match, flexible vacation, 11 paid holidays, 72 hours of paid sick leave, paid parental leave, adoption assistance, Employee Stock Purchase Plan, financial planning and group legal, voluntary benefits including auto, homeowner and pet insurance.
Requirements
Requires a Bachelor's degree in Computer Science or Engineering with 5 years of experience in software engineering or infrastructure management, or a Master's degree with 3 years of experience. Must have 5 years of experience in automation, programming, and scripting. Preferred qualifications include a Bachelor's degree with 7 years of experience, or a Master's with 5 years, along with 2 years of leadership/management experience and 2 years of budget experience. Automation and programming/scripting experience is preferred at 7 years.
Description

Capacity Ingestion and Management:

-       Supports team members designing and architecting infrastructure and/or service, sharing guidance on practices and terms for reliability and functionality.

-       Supervises team members and provides direction to ensure accurate forecasting of demands for infrastructure and response to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.

-       Maintains a collaborative relationship with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.

-       Implements expectations for identifying opportunities for prototyping and manages prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding) to explore novel approaches.

Incident and Service Lifecycle Management:

-       Monitors data collection, triage, technical analysis, and redirection, ensuring team members maintain and optimize operations and infrastructure reliability.

-       Provides support to team members monitoring services, ensuring they maintain up-to-date knowledge of performance and document their condition.

-       Leverages advanced knowledge to aid team members in performing incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).

-       Monitors comprehensive health and performance reporting and ensures team members take appropriate actions based on trends in data.

-       Ensures team members adhere to procedures when performing provisioning to support infrastructure, applications, and services.

-       Encourages team members to experiment with new approaches for and perform decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.

Automation:

-       Implements standards for identifying and recommending opportunities for automation and assesses potential benefits to enhance operational efficiency.

-       Takes a proactive role in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.

-       Shares strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.

Technical Communication and Guidance:

-       Reviews and provides feedback on release notes and ensures team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.

-       Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.

-       Serves as a resource to team members on what information to communicate and how to communicate.

Troubleshooting and Resolution:

-       Serves as a senior management escalation point for incidents and complex issues arising within Oracle services.

-       Monitors the resolution of technical issues spanning multiple services, ensuring effective investigation and debugging techniques are leveraged to achieve SLOs (service level objectives).

-       Shares expectations for documenting incidents performing root cause analyses, guiding team members to capture essential information for analysis and future reference.

-       Implements guidelines for post-mortem procedures to prevent incident reoccurrence.

-       Ensures team members adhere to service level agreements (SLAs) made with customers.

Innovation and Improvement:

-       Sets expectations for conducting experiments and evaluating cutting-edge tools and technologies to optimize infrastructure performance and reliability, taking proactive steps to adhere to security standards.

-       Manages and contributes to the prioritization of initiatives to improve performance bottlenecks and deployments, ensuring efficient resource usage, speed, and scalability.

-       Implements standards for developing and maintaining knowledge of site reliability trends and sharing valuable insights and information with team members, management, and beyond to promote innovative building, testing, deploying, and running services.

-       Leverages analyses and data from teams to contribute to business development decisions (e.g., design changes).

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00