Verified today
Senior Staff Site Reliability Engineer, Linux/Network troubleshooting/Scripting
About the role
SRE Cloud Infrastructure & Operations team at Zscaler builds and operates the cloud-native Zero Trust Exchange platform. The Senior Staff Site Reliability Engineer architects, automates and scales large‑scale distributed systems, leads observability and container orchestration, and drives incident management. Hyderabad, hybrid work model (3 days onsite), with relocation assistance.
What you’ll do
- Design and implement advanced cloud automation to reduce toil
- Oversee container orchestration and ensure production performance
- Lead creation and optimization of monitoring and alerting systems
- Own cloud operations, deployments, on‑call support, and incident response
- Collaborate with cross‑functional teams on technology solutions
- Mentor team members on SRE best practices
What you’ll bring
- 7+ years designing and troubleshooting large‑scale distributed systems
- Hands‑on experience with Kubernetes (EKS/GKE) and container orchestration
- Proficiency in Terraform, Ansible, Python, Go, Java, C
- Deep knowledge of AWS cloud services and networking fundamentals
- Experience building observability systems (Grafana, SLIs/SLOs)
- Strong DevOps skills: CI/CD pipelines, SCM, builds/releases
- Understanding of web security protocols (HTTP, SSL/TLS, DNS, SQL)
- Ability to lead cross‑functional projects and mentor SRE teams
Skills
Benefits
- Comprehensive health plans
- Parental leave options
- Retirement savings options
- Education reimbursement
- In‑office perks