Active listing
Staff Site Reliability Engineer, Observability
About the role
Site Reliability Engineering team focuses on building reliable, scalable cloud infrastructure. The Staff Site Reliability Engineer leads observability initiatives, designs monitoring solutions, automates incident response, and mentors engineers to improve system resilience. Remote work from India, fully remote environment with flexible hours.
What you’ll do
- Design and implement enterprise observability solutions
- Lead reliability and performance improvements across AWS and Kubernetes
- Manage vulnerability remediation and patch compliance
- Operate and optimize Kubernetes platforms and GitOps workflows
- Develop AI‑assisted automation tools for incident response
- Partner with engineering teams to define monitoring standards
- Mentor engineers and promote SRE best practices
- Participate in incident management and root cause analysis
What you’ll bring
- 8+ years cloud infrastructure experience
- 5+ years Kubernetes (EKS, AKS, GKE)
- Strong Go or Python programming
- Terraform or similar IaC expertise
- Deep knowledge of observability platforms (Prometheus, Grafana, OpenTelemetry)
- Linux systems administration
- CI/CD pipeline experience
- Security and compliance awareness
Skills
Benefits
- Remote work from India
- Performance‑based bonus and incentive opportunities
- Employee equity grants and stock purchase plan
- Comprehensive health and wellness benefits
- Retirement savings and statutory benefits