✨ AI Summary
Oracle is seeking a Principal Site Reliability Engineer to design and architect infrastructure and services, ensuring reliability and functionality. This role involves forecasting demands, responding to capacity needs, and collaborating with software development teams to build scalable infrastructures. The engineer will exercise judgment in data collection for optimization, perform incident response and maintenance, and provide comprehensive health and performance reporting. Key responsibilities include identifying automation opportunities, communicating service information, and proactively anticipating the impact of changes. The position also requires providing technical support, documenting incidents, conducting advanced experiments with new tools, and maintaining up-to-date knowledge of site reliability trends.
🎁 Benefits & Perks
competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
Requirements
Designs and architects infrastructure and services for reliability and functionality, forecasts demands, and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment in data collection for optimization, performs incident response and maintenance, and provides health and performance reporting. Identifies automation opportunities, communicates service information, and anticipates the impact of changes. Provides technical support, documents incidents, conducts experiments with new tools, and maintains knowledge of site reliability trends.
Description
Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation. Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.