Verified today
Staff Site Reliability Engineer
About the role
ServiceTitan's Site Reliability & Infrastructure Engineering team owns the reliability and health of applications running on its cloud, designing signals that detect issues and building systems that keep ServiceTitan running better, faster, and cheaper as it scales. As a Staff Site Reliability Engineer, you will participate in an on-call rotation, design and maintain observability dashboards and alerting grounded in SLIs and SLOs, operate and improve the Kubernetes-based compute platform, and work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems. You will investigate and resolve production incidents, leverage AI-assisted engineering tools, define non-functional requirements, partner with product engineering teams on architecture reviews, drive adoption of reliability best practices, build automation, maintain runbooks, and contribute to CI/CD pipelines. The role is based in Bengaluru, Karnataka, India.
What you’ll do
- Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load)
- Design, build, and maintain observability dashboards and alerting grounded in SLIs and SLOs
- Operate and improve the Kubernetes-based compute platform, which runs the large majority of infrastructure
- Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems
- Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work
- Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) and build SRE skills to automate investigations and ship fixes across infrastructure and application repositories
- Define non-functional requirements — scalability, availability, performance — for new systems as they're designed
- Partner with product engineering teams to review architecture and infrastructure decisions before they ship
What you’ll bring
- 10+ years of relevant hands-on experience
- Strong, hands-on understanding of Kubernetes as a system
- Practical experience with SLIs, SLOs, and error budgets on real systems
- Solid grounding in AWS, GCP, or Azure, including networking fundamentals (subnetting, IP addressing)
- Deep experience with at least one modern observability stack (OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch)
- Strong understanding of a CI/CD system (GitHub Actions preferred; TeamCity, Azure DevOps, or GitLab CI acceptable)
- Experience building web applications on .NET stack, Python (Flask, FastAPI), or Java (Spring) deployed at scale
- Experience with AI tools (building SRE skills, root cause analysis agents deployed in production systems, etc.) is required