Senior Site Reliability Engineer

🏢 Oracle
📍 Reston, United StatesOn-site
📅 Posted: 1mo ago🔄 Updated: 1mo ago
CV%
Required Skills
Information Technology
LinuxDebuggingPythonScriptingTerraformVMwareContainerizationKubernetesMonitoringOpenTelemetryObservability
Other
Ansiblesoftware developmentUnixkernel tuningperformance profilingGlusterFSNFSSMBiSCSINVMe-oF
Soft Skills & Professional Competencies
Problem SolvingMultitaskingPrioritizationCommunicationCollaborationResilience
Productivity & Workplace Tools
RPA
Business, Sales & Management
Process Improvement
Nice to have:
Information Technology
DockerJenkinsPrometheusGrafanaPostgreSQL
Other
CephMinIO
🎁 Benefits & Perks
Medical, dental, and vision insurance, including expert medical opinion; Short term and long term disability; Life insurance and AD&D; Supplemental life insurance (Employee/Spouse/Child); Health care and dependent care Flexible Spending Accounts; Pre-tax commuter and parking benefits; 401(k) Savings and Investment Plan with company match; Paid time off: Flexible Vacation (13 days annually for the first three years, 18 days thereafter for full-time employees); 11 paid holidays; Paid sick leave (72 hours upon hire, up to 112 hours carryover); Paid parental leave; Adoption assistance; Employee Stock Purchase Plan; Financial planning and group legal; Voluntary benefits including auto, homeowner and pet insurance.
Requirements
The role requires a Bachelor's or Master's degree in Computer Science or a related engineering field, with at least 5 years of experience in software development/IT operations. Must have 5+ years in SRE, Infrastructure, or Systems Engineering roles, deep expertise with Unix/Linux systems (particularly Oracle Linux), and proficiency in Python and Bash scripting. Strong grasp of IaC tools like Ansible and Terraform, and experience with hybrid infrastructure, monitoring, and distributed storage systems like GlusterFS are essential. A US Government TS/SCI with Polygraph clearance and U.S. Citizenship are mandatory.
Description

Capacity Ingestion and Management:

-       Participates and listens in on discussions for the design and architecture of infrastructure and/or service according to terms for reliability and functionality.

-       Assists team members responding to infrastructure demands and capacity increases to support current and future workloads.

-       Supports collaborations with the software development team to contribute to the development of reliable and scalable infrastructures based on detailed deployment requirements.

-       Participates in identifying opportunities for prototyping and provides support for prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding).

Incident and Service Lifecycle Management:

-       Assists in data collection, triage, and redirection to maintain and optimize operations and infrastructure reliability.

-       Monitors services and maintains up-to-date knowledge of their performance.

-       Supports incident response and/or maintenance tasks (e.g., software installs, version upgrades, and security updates, backup and recovery) under supervision.

-       Assists in providing health and performance reporting and takes appropriate actions based on trends in data.

-       May perform provisioning according to established procedures to support infrastructure, applications, and services.

-       May perform decommissioning (e.g., shutting down servers, removing data from databases) according to established procedures to remove objects that are no longer needed.

Automation:

-       Assists in identifying opportunities for automation and assessing potential benefits.

-       Supports the development of automation or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.

-       Follows detailed instructions to conduct testing to ensure automation performs tasks correctly and produces expected results, with supervision.

Technical Communication and Guidance:

-       Communicates basic information about the scale, capacity, security, and performance attributes of services and technology within immediate team.

-       Assists in identifying and communicating basic infrastructure, feature, and tool changes within immediate team.

Troubleshooting and Resolution:

-       Provides operational support for technology, escalating routine, low-impact incidents and other issues arising within Oracle services.

-       Participates in on-call shifts to address issues.

-       Assists with resolving technical issues, performing investigations, and debugging products in order to reach SLOs (service level objectives), with supervision.

-       Follows detailed instructions to document incidents and perform root cause analyses according to standard reporting methods.

-       Participates in post-mortem procedures to prevent incident reoccurrence.

Innovation and Improvement:

-       Assists in experimenting with new tools and technologies to improve infrastructure performance and reliability and helps ensure adherence to security standards, with supervision.

-       Supports the execution of improvements for performance bottlenecks and deployments.

-       Gains basic knowledge of site reliability trends and shares relevant information with immediate team members.

-       Performs analyses as assigned and assists in providing clear data on production to support business development decisions (e.g., design changes).

Core Responsibilities:

Planning & Execution:

-       Completes assigned tasks and monitors timelines to ensure timely completion of work in accordance with project requirements, with supervision. Follows direction to prioritize work and adjust to shifts in resources or timelines.

Collaboration & Partnership:

-       Collaborates with team members to better understand expectations and contribute to shared objectives. Builds basic understanding of business, stakeholder, and/or customer needs with guidance.

Problem Solving:

-       Follows standard procedures to identify and escalate issues to senior team members. Collects and reviews basic data and/or information to troubleshoot common errors.

Continuous Learning:

-       Builds knowledge and learns new skills and/or tools aligned with industry trends and best practices as directed. Incorporates feedback and participates in training to improve skills.

Continuous Improvement:

-       Begins to identify ways to increase the efficiency and effectiveness of processes, protocols, and workflows with guidance.

 

✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00