Verified today
Staff Site Reliability Engineer
About the role
Site Reliability Engineering at Okta builds and operates highly scalable, secure infrastructure across AWS and GCP. The Staff SRE leads multi-cloud migrations, designs Kubernetes platforms, and drives automation with Terraform and CI/CD pipelines. Based in Bengaluru, Karnataka, India, this role works on-site and participates in on-call rotations, mentoring engineers and collaborating with security and compliance teams.
What you’ll do
- Design, build, and operate scalable cloud infrastructure
- Lead container platform migrations and microservice enablement
- Implement infrastructure as code using Terraform and Ansible
- Develop and maintain CI/CD pipelines and automation
- Define SLOs/SLIs, conduct postmortems, and improve incident response
- Mentor engineers and foster reliability best practices
- Collaborate with security/compliance to ensure standards
- Participate in on‑call rotation and drive observability improvements
What you’ll bring
- 8+ years SRE/DevOps experience
- 3-5 years Kubernetes (EKS/GKE)
- 3-5 years AWS and GCP
- 3-5 years Terraform multi-cloud
- 5+ years Python or Go development
- Experience with CI/CD tools (ArgoCD, GitLab CI, Spinnaker)
- Hands‑on with observability stacks (Prometheus, Grafana, ELK)
- Strong Linux and networking fundamentals