Verified today
Site Reliability Engineer III
About the role
As a Site Reliability Engineer III within JPMorgan Chase's Cyber & Technology Controls team in Hyderabad, Telangana, India, you will drive reliability and scalability for critical applications through code, cloud, and SRE best practices. Day-to-day work includes configuring, maintaining, monitoring, and optimizing applications and infrastructure; designing and implementing CI/CD pipelines and infrastructure as code; defining SLOs/SLIs; leading incident response and blameless postmortems; and using enterprise-authorized AI to accelerate incident triage and identify reliability risks. The role is a full-time position based in Hyderabad.
What you’ll do
- Guide and assist others in building appropriate level designs and gaining consensus from peers
- Collaborate with software and infrastructure engineers to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
- Design, develop, test, and implement solutions that improve the availability, reliability, and scalability of applications
- Implement infrastructure, configuration, and network as code for applications and platforms
- Collaborate with technical experts, key stakeholders, and team members to resolve complex problems
- Use enterprise-authorized AI capabilities to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and secu
- Apply enterprise-authorized AI capabilities to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes
- Utilize service level indicators and objectives to proactively resolve issues before they impact customers
What you’ll bring
- Formal training or certification on site reliability engineering concepts and 3+ years applied experience
- Proficient in site reliability culture and principles and familiarity with implementation within applications or platforms
- Proficient in at least one programming language such as Python or Java/Spring Boot
- Proficient knowledge of software applications and technical processes within technical disciplines (e.g., Platform Infra, Cloud, Artificial Intelligence)
- Experience in observability including monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk
- Experience with continuous integration and continuous delivery tools like Jenkins, GitLab, or Terraform
- Experience with SLO/SLI definition, chaos engineering, and disaster recovery planning
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
Skills
Education
Formal training or certification on site reliability engineering concepts