Verified today
Associate Principal Site Reliability Engineer
About the role
Cloud Services Engineering team at Saviynt builds AI-native identity platform infrastructure. The Associate Principal Site Reliability Engineer owns uptime, reliability, and performance of AWS and Kubernetes services, designs self‑healing systems with LLM‑powered automation, and leads incident response and chaos engineering. Bengaluru, hybrid work model, 3 days onsite, relocation assistance available.
What you’ll do
- Own uptime and performance of cloud services
- Design and implement self‑healing infrastructure using AI agents
- Build LLM‑powered operational tooling for alert triage and runbook automation
- Manage and scale Kubernetes workloads and cost efficiency
- Develop observability systems with Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry
- Define SLOs/SLAs and error budgets
- Automate infrastructure via CI/CD pipelines
- Lead incident response, postmortems, and chaos engineering
What you’ll bring
- 8+ years SRE/DevOps/Platform Engineering
- Hands‑on AWS at scale
- Production‑grade Kubernetes (EKS)
- Python or Go development
- Terraform automation
- Monitoring/alerting (Prometheus, Grafana)
- Experience with LLM APIs (OpenAI)
- Familiarity with AI agents (LangChain, AutoGen)
Skills
Benefits
- High‑growth environment
- Opportunities for learning and career advancement
- Positive and inclusive workplace culture