Verified today
Engineering Manager, Site Reliability Engineering
About the role
Service Operations Site Reliability Engineering team at athenahealth delivers highly available SaaS infrastructure, operational tooling, and observability solutions for Service Operations. The Engineering Manager leads and mentors a team of SREs, drives reliability, automation, and observability initiatives, and partners with global engineering and operations groups to reduce toil and improve service health. Chennai, on-site role with hybrid collaboration across India and the US.
What you’ll do
- Lead, coach, and develop a team of Site Reliability and Infrastructure Engineers
- Design and implement observability strategies across metrics, logs, traces, and alerts
- Manage provisioning and lifecycle of Linux systems using IaC tools
- Partner with engineering teams to ensure monitoring coverage in SaaS and Kubernetes environments
- Identify and automate operational toil through self‑service tools and documentation
- Drive Agile delivery practices including sprint planning and backlog prioritization
- Oversee incident response, post‑incident analysis, and on‑call processes
- Develop and report operational metrics for reliability and automation coverage
What you’ll bring
- 10+ years in Infrastructure, Site Reliability, or Platform Engineering
- 2+ years managing technical engineering teams
- Strong hands‑on Linux systems administration at scale
- Experience with observability platforms (metrics, logs, tracing)
- Proficiency with IaC tools such as Terraform, Puppet, Ansible
- Scripting in Python, Go, Bash or similar languages
- Experience supporting SaaS, hybrid cloud, Kubernetes/EKS environments
- Excellent communication and stakeholder management
Education
Bachelor's degree in Computer Science, Information Technology, Engineering or related field