Verified today
Senior Site Reliability Engineer
About the role
Senior Site Reliability Engineer at ISS STOXX in Mumbai. Responsible for designing, expanding, and optimizing Kubernetes-based deployment architectures for microservices, enhancing monitoring stack with Grafana/Prometheus, ensuring consistent logs, building automation tooling, improving Docker images, partnering with DevOps, developing runbooks, and championing reliability best practices.
What you’ll do
- Design and optimize Kubernetes deployment architectures for microservices.
- Enhance monitoring stack with Grafana, Prometheus, and related tools.
- Ensure consistent, actionable logs integrated into log-indexing infrastructure.
- Write and maintain code for build automation, tooling, and operational workflows.
- Build and improve Docker images, manage workloads across Kubernetes clusters.
- Partner with DevOps to improve deployment practices and automation.
- Develop documentation for runbooks and escalation procedures.
- Champion reliability best practices and guide platform architecture evolution.
What you’ll bring
- Bachelor's in CS or related field, or equivalent experience.
- 4-5+ years programming in Java, Python, Go, or Django.
- Deep experience troubleshooting large-scale distributed systems.
- Strong interest in reliability engineering, automation, resilient systems.
- Proven communication and cross-functional teamwork.
- Strong Linux/Unix and system internals experience.
Skills
Education
Bachelor's