Verified today
Senior Site Reliability Engineer
About the role
Senior Site Reliability Engineer needed to monitor, troubleshoot, and automate large-scale distributed systems at Zeta, ensuring reliability and performance across cloud and on-prem environments.
What you’ll do
- Monitor and troubleshoot system performance and incidents
- Develop and maintain alerting, runbooks, and automation scripts
- Analyze and improve system reliability and scalability
- Collaborate with cross-functional teams on architecture and ops
- Implement observability and security best practices
What you’ll bring
- 4-6 years sysadmin experience with large-scale distributed systems
- Strong cloud (AWS) and Kubernetes orchestration skills
- Proficient in Unix shell, Python, Go, and database (MySQL/PostgreSQL)
- Experience with observability tools (Prometheus, Grafana) and alerting
- Solid networking, CI/CD (Jenkins, ArgoCD), and security best practices
- BS in Computer Science or related field
Skills
Education
Bachelor's