Site Reliability Engineer III

🏢 JP Morgan
📍 Jersey City, United StatesFull-timeOn-site
📅 Posted: 6d ago🔄 Updated: 6d ago
CV%
✨ AI Summary
JPMorgan Chase is seeking a Site Reliability Engineer III for their Commercial Investment Banking Fraud Prevention team. This role involves owning the day-to-day operational health of platforms, focusing on availability, latency, and error rates, and proactively identifying reliability risks. The engineer will define and evolve SLIs/SLOs, build actionable alerting, improve end-to-end observability, and participate in incident response and problem management. Key responsibilities include Kubernetes and AWS operations, release engineering, infrastructure as code using Terraform, and database reliability partnerships. The role also requires using enterprise-authorized AI capabilities to accelerate incident triage and identify reliability risks. Required qualifications include formal training or certification in software engineering with 3+ years of experience, hands-on experience with Kubernetes and AWS production environments, and proficiency in CI/CD tools like Spinnaker/Harness, Terraform, and scripting languages (Python/Bash/Go). Strong incident response skills, Linux/networking fundamentals, and experience with AI tools for SRE workflows are also necessary. Preferred qualifications include experience with SLO programs, performance and resilience testing, operational maturity improvements, and fraud screening.
Required Skills
Information Technology
KubernetesAWSTerraformPythonGoLinuxCore NetworkingDistributed SystemsGenerative AI
Other
SpinnakerHarnessBash
Nice to have:
Other
SLO definitionfault injectionchaos engineeringrunbooksautomated health checksdeployment guardrailsfraud screeningpayment flowsdatabase reliability patterns
Information Technology
AlertingPerformance TestingSLA Management
Business, Sales & Management
Workforce Planning
Soft Skills & Professional Competencies
Resilience
Finance, Legal & Governance
Business ContinuityMediation
🎁 Benefits & Perks
Comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching.
Requirements
Requires formal training or certification in software engineering concepts with 3+ years of applied experience. Must have hands-on experience in SRE/DevOps/production engineering, operating Kubernetes workloads, and practical experience with AWS production environments. Proficiency in CI/CD tools (Spinnaker/Harness), Terraform, and scripting languages (Python/Bash/Go) is essential, along with strong incident response skills, RCA writing, and fundamental knowledge of Linux, networking, and distributed systems. Familiarity with enterprise AI capabilities for SRE workflows is also required.
Description
There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. 

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial Investment Banking team of Fraud Prevention, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform. 

 You are an integral part of a team that works to develop high-quality architecture solutions for various software applications and platform products. You drive significant business impact and help shape the target state architecture through your capabilities in multiple architecture domains. You will ensure the platform is reliable, secure, performant, and resilient in production across Kubernetes-based environments and AWS. You will apply SRE principles to drive measurable improvements in availability and latency, reduce operational toil through automation, and strengthen deployment safety and recovery capabilities in close partnership with engineering and platform teams.


Job responsibilities

  • Production ownership & reliability outcomes: Own day-to-day operational health for the platform, focusing on availability, latency, throughput, and error rates; proactively identify reliability risks and drive remediation.
  • SLO/SLI and alerting strategy: Define and evolve SLIs/SLOs and error budgets; build actionable alerting aligned to customer impact and reduce noise through tuning and standardization.
  • Observability & troubleshooting: Improve end-to-end observability (metrics, logs, traces), dashboards, and runbooks across Kubernetes and AWS; perform deep technical triage of distributed system issues.
  • Incident response & problem management: Participate in on-call and lead/assist incident triage, mitigation, and recovery; conduct RCAs and drive corrective/preventive actions to closure.
  • Kubernetes operations: Support containerized workloads, autoscaling, rollout/rollback procedures, resource tuning, and resilience patterns.
  • AWS operations (container + serverless): Operate components on AWS (EKS/ECS/Lambda) and associated data services (Dynamo DB, S3); manage operational concerns such as scaling, retries, and safe failure modes.
  • Release engineering & delivery reliability: Improve the safety and repeatability of deployments using Spinnaker and Harness.
  • Infrastructure as Code & environment consistency: Build and maintain Terraform modules and automation for reliable, repeatable environments.
  • Database reliability: Partner with engineering and database/platform teams to improve reliability patterns across multiple database technologies.
  • Security & controls in operations: Apply secure operational practices and ensure operational processes meet required controls.
  • Lead small-to-medium initiatives from proposal through production adoption.
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
 
Required qualifications, capabilities, and skills
  • Formal training or certification on software engineering concepts and 3+ years applied experience
  • Experience in SRE/DevOps/production engineering or equivalent
  • Hands-on experience operating Kubernetes workloads (deployments, scaling, debugging)
  • Practical experience with AWS (EKS, ECS, Lambda, Dynamo DB, S3) in production
  • Experience with CI/CD and release tooling such as Spinnaker and/or Harness
  • Proficiency with Terraform (IaC), plus scripting/automation (Python/Bash/Go)
  • Strong incident response skills, RCA writing, and ability to drive remediation work
  • Solid fundamentals in Linux, networking, and troubleshooting distributed systems
  • Ability to independently execute well-scoped reliability work and escalate when needed
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
 
Preferred qualifications, capabilities, and skills
  • Experience implementing SLO programs and alerting aligned to customer journeys
  • Experience with performance testing, capacity planning, and resilience testing (fault injection/chaos, DR exercises)
  • Experience improving operational maturity: standardized runbooks, automated health checks, auto-remediation, and deployment guardrails
  • Experience with fraud screening/decisioning or payment flows
  • Familiarity with database reliability patterns (capacity, backups, failover readiness)
  • Experience with secure operational practices (least privilege, secrets handling)
  • Experience partnering with engineering and platform teams to drive reliability improvements
✨ Premium Match Details
Deep-dive CV analysis, customized Cover Letters, and Interview prep!
📊 Match Analysis
Insights against your active CV
📊
Personalized Match Analysis
Upload your CV to see exact matching percentages, detailed skills mapping, and gap analysis for this role.
🎯 Overalli74%
⚡ Skillsi85%
View Breakdown
Ontology Match: 85.0
Matched:✓ Requirements Matching✓ Ontology Skills Mapping
📜 Eligibilityi49%
View Breakdown
Local: 19600%
🏗️ Career Fiti91%
View Breakdown
Seniority: 91.0
📋 Requirementsi67%
View Breakdown
Domain: 67.0
🔥 Motivationi78%
View Breakdown
Title Fit: 78.00